
Solution
Metadata-Driven Ingestion
Solution
Metadata-Driven Ingestion
Industry
CPG
Region
US
Technology
Microsoft Azure
Context
The client, a multinational manufacturer of confectionery and other consumer products, manages vast volumes of structured and unstructured data across multiple business units, geographies, and source systems. As the business scaled, so did the complexity of its data landscape — spanning ERP systems, syndicated POS feeds, trade promotion platforms, and third-party sources, each with its own format and load pattern. Their existing data ingestion pipeline was fragmented, code-driven, and effort-intensive, requiring custom development for nearly every new source, since the organization had no true metadata-driven data ingestion framework in place. It lacked the flexibility of a governed data onboarding approach, leaving teams without a reusable, standardized way to bring new sources online. This resulted in delays in data availability, operational inefficiencies, governance risks, and slower time-to-insight — a common bottleneck for CPG enterprises without a mature metadata-driven data ingestion framework or broader data governance framework.
Problem Statement
MathCo partnered with the client to build a scalable data ingestion solution — a metadata-driven ingestion framework on Azure Data Factory — enabling reusable, governed data onboarding across the enterprise:
- Developed reusable connectors to power an automated data pipeline across diverse formats (CSV, Excel, API, Delta) and load patterns, including Append, Merge, Overwrite, and SCD Type 2.
- Standardized and centralized data into a unified data lake with automated Data Quality (DQ) rules, forming the backbone of enterprise-grade data quality automation.
- Leveraged Azure DevOps to enable seamless, automated deployment of notebooks, scripts, and workflows across the enterprise data ingestion pipeline.
- Utilized Unity Catalog for data cataloging, lineage, and access control, complemented by Azure Purview for centralized data quality management and catalog governance.
Impact
- Successfully onboarded 8+ analytics products with streamlined data capture and transformation on Databricks.
- Saved 5,000+ hours in building ingestion pipelines by implementing a metadata-driven framework.
- Achieved 50% faster data ingestion for new data sources, accelerating Time-to-Insights.
- Reduced data quality issues by 40% through automated data quality checks, enhancing trust and reliability of analytics.