Data Association Using Complete Lists for Exact Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data processing techniques face challenges in efficiently matching and processing large volumes of data from multiple sources, leading to increased computational and storage resources being used on irrelevant data, and there is a need for assured matching of data objects referring to the same entity without relying on statistical confidence.

Innovation Solution

A method that identifies shared attributes and values between data objects, uses a complete list to determine exact matches, and avoids processing non-matching data objects, thereby reducing overall processing time and storage requirements by ensuring only relevant data is processed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data processing techniques are used to match data objects from multiple sources, then data processing can be performed, but computational resources and storage resources are substantially increased due to processing irrelevant data

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by performing data pre-processing and validation before main processing. Complete lists are pre-computed and stored for shared attributes, so that during actual data matching, the system can quickly determine exact matches without performing extensive computations on all data objects. This preliminary preparation of complete lists eliminates the need to process irrelevant data objects during the main processing phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and processes only the relevant portion of data by identifying shared attributes between data objects and using complete lists to determine exact matches. By focusing only on data objects that have matching shared attributes with the same values, the system extracts and processes only the necessary data, avoiding the computational overhead of processing all data objects from multiple sources.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If traditional data processing techniques are used to match data objects from multiple sources, then data processing can be performed, but storage resources are substantially increased due to retaining large volumes of data

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidstorage resources
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing complete lists for shared attributes before the main processing phase. These complete lists contain all possible values for each shared attribute, allowing the system to quickly determine exact matches during data processing without needing to store and process all raw data objects. This preliminary preparation reduces the storage requirements during the main processing phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential information needed for matching by using complete lists of shared attributes. Instead of storing all data objects from multiple sources, the system extracts and stores only the complete lists of shared attribute values, which are sufficient to determine exact matches. This extraction approach significantly reduces the quantity of data that needs to be retained in storage.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If data objects are matched using statistical confidence methods, then matching can be performed, but decision-making accuracy is reduced due to reliance on probabilities rather than absolute confidence

Engineering Contradiction:
Improvematching capabilityVSAvoiddecision-making accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent uses complete lists as disposable intermediate structures that are pre-computed and then used to determine exact matches. These complete lists serve as a simple, efficient mechanism for achieving absolute confidence in matching without requiring complex statistical analysis. The system can confidently determine that two data objects refer to the same entity when they share all values in a complete list for a shared attribute, eliminating the need for probabilistic reasoning.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

4Reliability

If all data objects from multiple sources are processed, then comprehensive data analysis can be performed, but overall processing time is increased

Engineering Contradiction:
Improvedata analysis comprehensivenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts and processes only the relevant data objects by using complete lists to identify exact matches. Instead of processing all data objects from multiple sources, the system extracts only those data objects that have matching shared attributes with the same values as the first data object. This extraction approach maintains comprehensive analysis of relevant data while significantly reducing processing time by eliminating irrelevant data objects from the processing pipeline.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10997248B2Data association using complete lists
Publication Date: 2021.05.04 IGMR RES LTD
  • US10997248B2 patent drawing
  • US10997248B2 patent drawing
  • US10997248B2 patent drawing

AI summary

A method, a system and a product for performing a processing operation with respect to a first data object. The method comprises obtaining a second data object, which comprises a second set of attributes and values thereof different than the first set of attributes and values thereof of the first data object; identifying, in the first and second sets of attributes, a shared attribute, each of which having a corresponding shared value; obtaining a complete list with respect to the shared attribute; and in response to determining that a number of entries in the complete list that comprise the corresponding shared value for each of the at least one shared attribute is exactly one, processing the second data object as part of the processing operation of the first data object; and avoiding processing in the processing operation additional object; whereby reducing an overall processing time and an overall storage required for performing the processing operation.