Lineage-Driven Data Mart Code Generation for Traceable Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in managing a data mart lifecycle is the lack of standardized practices to generate source code that effectively builds a data mart in a consumable form for analytical purposes, leading to poorly managed infrastructure and limited visibility into data lineage, which is often derived manually after deployment, resulting in tedious and error-prone processes.

Innovation Solution

A lineage-driven approach that programmatically generates source code for building, testing, and deploying data marts and pipelines using a data lineage configuration to identify data sources, relations, and processing operations, automating the process and standardizing it across different use cases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manual derivation of data lineage is used after deployment, then data lineage information can be obtained, but the process becomes tedious and error-prone

Engineering Contradiction:
Improvedata lineage visibilityVSAvoidtime for deriving data lineage
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies preliminary action by generating source code with embedded data lineage information during the development phase, rather than deriving it manually after deployment. The lineage metadata is captured upfront when the data mart is created, storing information about data sources, transformations, and dependencies in the generated code itself, eliminating the need for post-deployment lineage derivation.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If standardized source code generation practices are not implemented, then flexibility in building data marts is maintained, but infrastructure management becomes poor and scalability is limited

Engineering Contradiction:
Improveease of data mart constructionVSAvoidscalability of data mart deployment
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent applies self-service by creating an automated system that generates standardized source code for data mart construction. The system automatically produces consistent, well-structured code templates that can be reused across multiple data marts, enabling teams to build data marts systematically without requiring expert intervention for each project, thereby improving both ease of manufacture and scalability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies parameter changes by using configurable templates and parameters that can be adjusted to generate different data mart implementations from the same standardized framework. By changing parameters such as data source connections, transformation rules, and target schemas, the system can generate customized data marts while maintaining consistent structural standards, enabling scalable deployment across diverse use cases.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If expert knowledge is required for each data mart project, then code quality can be maintained, but reliance on experts increases and development time extends

Engineering Contradiction:
Improvecode qualityVSAvoidautomation of data mart generation
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The patent applies self-service by creating an automated system that generates standardized source code for data mart construction. The system automatically produces consistent, well-structured code templates that can be reused across multiple data marts, enabling teams to build data marts systematically without requiring expert intervention for each project, thereby improving both ease of manufacture and scalability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies feedback by incorporating data lineage information and metadata from actual data mart implementations back into the code generation system. This feedback loop allows the system to learn from expert-created data marts and improve its automated generation capabilities, maintaining code quality while reducing dependency on expert knowledge for each new project.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250321873A1Lineage-driven source code generation for building, testing, deploying, and maintaining data marts and data pipelines
Publication Date: 2025.10.16 CAPITAL ONE SERVICES LLC
  • US20250321873A1 patent drawing
  • US20250321873A1 patent drawing
  • US20250321873A1 patent drawing

AI summary

In some implementations, an analytics platform may process a data lineage configuration to identify one or more data sources storing one or more data sets associated with a data analytics use case and identify one or more data relations associated with the one or more data sets. The analytics platform may parse a source code template defining one or more functions to build a data mart and a data pipeline that enables the data analytics use case. The source code template may include one or more tokens to specify elements that define the one or more data sources, data sets, and data relations identified in the data lineage configuration. The analytics platform may automatically generate source code that is executable to build a data mart and a data pipeline to enable the data analytics use case based on the data lineage configuration and the source code template.