Dataset Identification Using Composite And Alias Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data processing systems face inefficiencies in representing complex and interconnected data relationships due to duplication of datasets with multiple possible representations, leading to excessive storage utilization and inaccurate information retrieval.
Innovation Solution
A data lineage system processes unique identifiers collectively and individually to generate composite and alias identifiers, reducing duplicate graph nodes by linking new datasets to existing nodes in a data lineage graph representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a data processing system uses traditional storage methods to represent datasets with multiple identification attributes, then the system can store data in a structured format, but it leads to duplication of datasets and excessive storage utilization
Solution Approach 1:
The patent merges multiple identification attributes of a dataset into a single composite identifier by processing the attributes collectively through a hash function. This consolidation eliminates redundant storage of multiple attributes separately, directly reducing storage utilization while maintaining the ability to represent datasets with their full identification characteristics.
Solution Approach 2:
The composite identifier serves multiple functions: it uniquely identifies a dataset, enables efficient storage, and facilitates quick retrieval. By making the identifier multi-functional, the system eliminates the need for separate storage structures for different identification attributes, thereby reducing overall storage requirements while improving data representation efficiency.
2Measurement precision
If a data processing system creates separate representations for datasets with multiple identification attributes, then the system can maintain detailed information, but it results in duplicate graph nodes and inaccurate information retrieval
Solution Approach 1:
The patent combines multiple identification attributes into a single composite identifier that uniquely represents a dataset. This merging prevents the creation of duplicate graph nodes for the same dataset, as each unique combination of attributes produces a unique composite identifier. Consequently, information accuracy is maintained while graph structure complexity is reduced.
Solution Approach 2:
The patent transforms multiple identification attributes into a single parameter (composite identifier) through a hash function. This parameter change simplifies the graph structure by reducing the dimensionality of dataset representation, eliminating duplicate nodes, and improving information retrieval accuracy without losing the ability to distinguish between different datasets.
3Adaptability or versatility
If a data processing system stores multiple identification attributes individually, then the system can maintain flexibility in data access, but it increases storage requirements and processing overhead
Solution Approach 1:
The composite identifier is designed to be multi-functional, serving as a unique key for dataset identification, storage, and retrieval operations. This universal identifier maintains data access flexibility by enabling efficient queries and operations, while simultaneously reducing storage requirements by consolidating multiple attributes into a single compact representation.
Solution Approach 2:
The patent applies a hash function to transform multiple identification attributes into a single composite identifier parameter. This parameter change maintains the adaptability of data access through efficient hashing and retrieval operations, while significantly reducing storage requirements by eliminating the need to store multiple separate attributes for each dataset.
Data Source
AI summary
In some implementations, a system may receive information identifying a dataset. The system may process an identification attribute using a function that generates a first value, to generate a first identifier for the dataset. The system may search a data store storing a plurality of groupings to identify a grouping with the first identifier for the dataset. The system may extract a second identifier from the grouping with the first identifier for the dataset. The system may search a data lineage based graph representation of a plurality of datasets to identify a graph node representing the dataset. The system may update the data lineage based graph representation of the plurality of datasets to link the dataset with the at least one other dataset based on searching the data lineage based graph representation of the plurality of datasets to identify the graph node.


