Automated Data Matching Using Correlation Coefficients
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Companies face challenges in accurately matching and grouping data records across databases, leading to duplication and errors that affect relationships with customers and suppliers, and existing tools are inadequate for data cleansing and categorization.
Innovation Solution
A computer-implemented method that computes correlation coefficients between non-identical terms in data records to identify and combine duplicate records, using techniques such as correlation coefficient calculation, hash functions, and Boolean location vectors to efficiently classify and sort data records.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If automated data matching tools are used to group and cleanse data records, then data quality and accessibility are improved, but the complexity of the system increases and existing tools are inadequate for accurate matching of non-identical terms
Solution Approach 1:
The patent replaces traditional mechanical string-matching algorithms with a semantic field-based system that uses correlation coefficients and term associations to identify duplicate records. This substitution enables accurate matching of non-identical terms while maintaining automated operation, resolving the contradiction between reliability improvement and system complexity.
Solution Approach 2:
The patent introduces semantic fields and correlation coefficients as intermediary elements between data records. These intermediaries enable accurate matching by capturing the semantic relationship between non-identical terms, allowing the system to identify duplicates without requiring exact string matches, thus improving reliability without excessive complexity.
2Measurement precision
If traditional string-matching algorithms are used to identify duplicate records, then the system remains simple, but accurate matching of non-identical terms describing the same entity cannot be achieved
Solution Approach 1:
The patent changes the matching parameter from exact string equality to correlation coefficients based on term associations. By transforming the matching criterion from binary (identical/not identical) to a continuous correlation measure, the system achieves accurate matching of non-identical terms while using computationally efficient operations.
Solution Approach 2:
The patent performs preliminary analysis to identify associated terms and compute correlation coefficients before the actual matching process. This preliminary action creates a semantic framework that enables accurate matching of non-identical terms, reducing the complexity of the subsequent matching operation.
3Productivity
If data records are manually reviewed and corrected to eliminate duplicates, then matching accuracy is maintained, but processing time and operational costs increase substantially
Solution Approach 1:
The patent enables the system to automatically identify and group duplicate records using correlation coefficients and semantic field analysis. This self-service capability eliminates the need for manual review and correction, dramatically improving processing efficiency while maintaining high matching accuracy through automated semantic comparison.
Solution Approach 2:
The patent replaces manual data review processes with automated computational methods based on correlation coefficients. This substitution maintains matching accuracy while reducing processing time from manual operations to automated computation, significantly improving productivity.
4Extent of automation
If existing automated data mastering tools are deployed, then some level of automation is achieved, but they fail to accurately match records with non-identical terms describing the same entity
Solution Approach 1:
The patent replaces traditional automated string-matching algorithms with a semantic field-based system using correlation coefficients. This substitution maintains high automation levels while achieving accurate matching of non-identical terms by capturing semantic relationships between terms that describe the same entity.
Solution Approach 2:
The patent changes the automation approach from exact string matching to correlation-based semantic matching. This parameter change enables the automated system to accurately identify records with non-identical terms that describe the same entity, improving matching precision while maintaining full automation.
Data Source
AI summary
A computer-implemented method for processing data includes receiving an initial set of records including terms describing respective items in specified categories. Based on the initial set of records, respective term weights are calculated for at least some of the terms with respect to at least some of the categories. Each term weight indicates, for a given term and a given category, a likelihood that a record containing the given term belongs to the given category. Upon receiving a new record, not included in the initial set, respective assignment metrics are computed for two or more of the categories using the respective term weights of the particular terms in the new record with respect to the two or more of the categories. The new record is classified in one of the two or more of the categories responsively to the assignment metrics.


