Data Lineage Mapping for Resource-Efficient Linkage Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing large volumes of data in complex systems is resource-intensive and error-prone, with challenges in tracing errors, ensuring data compliance, and optimizing resource allocation.

Innovation Solution

A data processing system performs linkage analysis on datasets, generating a processed data representation that includes data lineage information and usage metrics, and provides visualizations to orchestrate data environments, thereby reducing manual inspection and resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual inspection and processing of large volumes of data is performed, then data management can be conducted, but computing resources and time are excessively consumed

Engineering Contradiction:
Improvedata management efficiencyVSAvoidtime for data processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by proactively mapping data environments and generating data lineage information before queries are executed. This includes pre-computing relationships between datasets and storing metadata about data sources, transformations, and dependencies, enabling faster query processing and reducing the need for manual inspection during actual data operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates processed data representations that are simplified copies or abstractions of the actual complex data environments. These processed representations include aggregated metadata, data lineage summaries, and relationship mappings that allow users to analyze and manage data without directly interacting with the full complexity of the underlying datasets, thereby conserving computing resources.

Inventive Principle:
Principle #26Copying

2Reliability

If comprehensive data lineage tracking is implemented, then data compliance and error tracing are improved, but system complexity increases

Engineering Contradiction:
Improvedata compliance trackingVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments data lineage tracking into manageable components by organizing metadata into structured hierarchies. Data lineage information is divided into discrete elements such as data source metadata, transformation metadata, and dependency metadata, each stored and processed independently. This segmentation allows the system to track comprehensive data relationships without overwhelming complexity, as each segment can be managed and queried separately.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If detailed processed data representations are generated, then actionable insights are enhanced, but computing resources are consumed

Engineering Contradiction:
Improvedata insights qualityVSAvoidcomputing resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by generating processed data representations at different levels of detail based on specific query requirements. Rather than always creating fully detailed representations, the system generates only the necessary level of processing needed to answer each query, conserving computing resources while still providing actionable insights. This includes selectively aggregating data lineage information and relationship details based on the specific analytical needs.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20260030265A1Data processing system with linkage analysis
Publication Date: 2026.01.29 CAPITAL ONE SERVICES LLC
  • US20260030265A1 patent drawing
  • US20260030265A1 patent drawing
  • US20260030265A1 patent drawing

AI summary

In some implementations, a device may receive information identifying a set of data reports generated from a group of datasets. The device may request from a data source storing the group of datasets, and based on receiving the information identifying the set of data reports, source data associated with the set of data reports. The device may receive the source data associated with the set of datasets. The device may associate the source data with data lineage information identifying a set of connections between the group of datasets. The device may generate, based on associating the source data with the data lineage information, a processed data representation, wherein the processed data representation includes information identifying a set of relationships associated with the group of datasets, the set of data reports, and a set of usage metrics. The device may transmit information identifying the processed data representation.