Data Fusion System for Entity Attribution Across Disparate Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing large and disparate datasets to link related data across disconnected data sets poses computational challenges, making it difficult to compare and derive insights from vast amounts of data effectively.
Innovation Solution
The implementation of a data fusion system that creates a consistent object model by transforming disparate data sources into an ontology-based object model, allowing for the generation of trajectories that represent entities across multiple interactions, and enabling efficient comparison and attribution of data points despite variations in data storage and precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data from multiple disparate data sets are analyzed to link related data, then valuable insights and understanding of the data as a whole are achieved, but computational challenges and difficulty in comparing data increase
Solution Approach 1:
The patent introduces an intermediary entity (such as a location, time stamp, or common attribute) that acts as a mediator to link records from disparate data sets. Instead of directly comparing all attributes of records from different data sets, the system uses these intermediary entities as common ground for joining and analyzing data, thereby reducing computational complexity while maintaining information completeness.
Solution Approach 2:
The patent segments the data linkage process into multiple stages: first identifying common intermediary entities across data sets, then joining records based on these intermediaries, and finally performing analysis on the connected data. This segmentation breaks down the complex task of linking disparate data sets into manageable steps, reducing overall computational burden.
2Adaptability or versatility
If data sets with different storage formats and precision levels are integrated, then comprehensive data analysis is enabled, but data comparison and attribution accuracy become difficult
Solution Approach 1:
The patent transforms data from different precision levels and storage formats into a unified representation. By standardizing how data is stored and accessed (e.g., using consistent data types, normalization techniques, or abstraction layers), the system enables accurate comparison and attribution across disparate sources while maintaining adaptability to various input formats.
Solution Approach 2:
The patent creates a universal data representation framework that can handle multiple data types, precision levels, and storage formats through a single unified interface. This universal approach allows the system to accommodate diverse data sources while ensuring consistent and accurate data comparison and attribution across all inputs.
3Quantity of substance
If vast amounts of data are made available for detailed analysis, then more complicated and detailed data analyses are possible, but it becomes more difficult to compare data to other data sets
Solution Approach 1:
The patent extracts and isolates the key intermediary entities (such as common identifiers, locations, or time stamps) from large volumes of disparate data. By separating these linking elements from the bulk data, the system enables efficient comparison and joining operations without requiring direct processing of the entire data volume, thus maintaining ease of operation with large data sets.
Data Source
AI summary
Systems and methods for using disparate data sets to attribute data to an entity are disclosed. Disparate data sets can be obtained from a variety of data sources. The disclosed systems and methods can obtain a first and second data set. Trajectories can represent multiple data records in a data set associated with an entity. Trajectories from the obtained data sets can be used to associate data stored among the various data sets. The association can be based on the agreement between the trajectories. The associated data records can further be used to associate the entities related to the associated data records.


