Requirement
The challenge
- Complex and heterogeneous source systems.
- Huge volumes of data (tens of terabytes per day).
- A wide variety of data from web, social media and internal systems.
- Long ETL and reporting cycles.
- No predictive or advanced-analytics capabilities.
- High maintenance due to the legacy code base and analytics environment.
Solution
What we did
- Collected a large variety of data from different sources into AWS S3.
- Processed, cleansed and aggregated the data for analytics using Elastic MapReduce on AWS.
- Loaded the data into Amazon Redshift for ad-hoc analysis.
- Automated the complete process in Python and removed the need for proprietary tools.
Tools Amazon S3Amazon EMRAmazon RedshiftPython
Outcomes
What changed for the client
- The complete analytics environment is migrated to the cloud.
- Cost savings by eliminating the need for proprietary tools.
- Real-time analytics, avoiding time-consuming ETL and report-scheduling activities.
- 360-degree analysis of customer touch-points and customer profiles.
- Low maintenance with automated error recovery.