Dynamic Importance Map Adjustment for Universal Data Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current master data management systems face challenges in accurately matching information across different data types, as existing algorithms are limited in their ability to handle various data types and are not dynamic enough to adapt to changes in data reliability and composition.

Innovation Solution

A method and system that generate training pairs using matching fields in records, determine similarities using an importance map with importance values, and adjust the importance map using Shapley values to improve matching accuracy across different data types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current matching algorithms are used to match information records, then matching can be performed for specific data types, but the algorithms cannot match all data types (such as people, cars, produce, dogs) with desired accuracy

Engineering Contradiction:
Improvedata type coverageVSAvoidmatching accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates a universal matching algorithm that can handle multiple data types (people, cars, produce, dogs, etc.) through a common framework. The system uses data type identification to select appropriate matching strategies and importance maps for each data type, enabling one algorithm to perform multiple matching functions that were previously required separate algorithms for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent dynamically adjusts matching parameters including importance values for different fields, similarity thresholds, and matching strategies based on the identified data type. Each data type has optimized parameters stored in importance maps that are selected and applied during the matching process, allowing the system to adapt its behavior to match the specific characteristics of each data type while maintaining a single unified algorithm.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If static importance maps are used in matching algorithms, then the matching process is simple, but the algorithms cannot adapt to changes in data reliability and composition

Engineering Contradiction:
Improveadaptability to data changesVSAvoidalgorithm complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent pre-computes and stores importance maps for different data types, containing field importance values and matching parameters determined through analysis of data characteristics and reliability. During actual matching operations, the system simply retrieves and applies the appropriate pre-computed importance map for the identified data type, avoiding complex real-time calculations while maintaining adaptability to different data types and their changing characteristics.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where matching results and data quality information are used to update and refine importance maps over time. The importance values for different fields are adjusted based on observed data reliability and matching performance, allowing the algorithm to adapt to changes in data composition and quality while maintaining a relatively simple operational structure.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20220414523A1Information Matching Using Automatically Generated Matching Algorithms
Publication Date: 2022.12.29 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220414523A1 patent drawing
  • US20220414523A1 patent drawing
  • US20220414523A1 patent drawing

AI summary

A method processes information. Training pairs are generated by a computer system using matching fields in matching pairs of records for a data type, wherein matches are present between the matching fields in the matching pairs of records. Similarities between the training pairs are determined by the computer system using an importance map with importance values for the matching fields. Shapley values are determined by the computer system using the training pairs and the similarities between the training pairs. The importance map is adjusted by the computer system using the Shapley values.