ML Label Merging for Complex Abbreviation Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data merging techniques struggle with efficiently merging data sets labeled with complex abbreviations and acronyms due to unreliable expansion methods and the lack of comprehensive data dictionaries, leading to increased time and unconscious biases in data processing and analytics.

Innovation Solution

Utilizing machine learning models to break down complex abbreviations into components, generate candidate words and phrases, and calculate similarity metrics to determine contextual similarity, enabling efficient and accurate merging of differently labeled data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual determination of data set mergeability is performed, then data scientists can evaluate each case individually, but the process becomes impractical due to large data volumes and time consumption

Engineering Contradiction:
Improveaccuracy of mergeability determinationVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical evaluation with an automated machine learning system that uses natural language processing models to expand abbreviations and calculate semantic similarity between data set labels, enabling scalable processing of large volumes of data while maintaining accurate mergeability determination

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary machine learning model that acts as a bridge between raw data labels and mergeability assessment, using abbreviation expansion and semantic similarity calculation to mediate the comparison process, thereby automating what was previously a manual task

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If simple abbreviation expansion rules are used, then the process is computationally efficient, but the expansion becomes unreliable for complex abbreviations

Engineering Contradiction:
Improveprocessing speedVSAvoidabbreviation expansion accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a dynamic approach where the system adapts its abbreviation expansion strategy based on the complexity of the abbreviation encountered, using machine learning models that can handle both simple and complex cases flexibly, thereby maintaining reliability across diverse abbreviation types without sacrificing processing efficiency

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the expansion process by using trained machine learning models that adjust their prediction behavior based on the input abbreviation structure, allowing the system to maintain high accuracy for complex abbreviations while preserving computational efficiency through optimized model inference

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive data dictionaries are created to improve label comparison, then mergeability accuracy improves, but the complexity and maintenance burden increases significantly

Engineering Contradiction:
Improvelabel comparison accuracyVSAvoiddata dictionary complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent enables the system to serve itself by training machine learning models on available data dictionaries and abbreviation mappings, allowing the models to learn expansion rules automatically from the data itself rather than requiring manually curated comprehensive dictionaries, thereby reducing maintenance burden while maintaining accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary training of machine learning models on available data dictionaries and abbreviation data before the actual mergeability assessment, allowing the models to learn expansion patterns in advance, so that during runtime the system can accurately compare labels without requiring complex real-time dictionary lookups

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12579369B2Data merging using machine learning
Publication Date: 2026.03.17 PWC PRODUCT SALES LLC
  • US12579369B2 patent drawing
  • US12579369B2 patent drawing
  • US12579369B2 patent drawing

AI summary

An exemplary system for determining whether data sets can be merged together may receive first data labeled with a first variable and second data labeled with a second variable, wherein the first variable represents a first sequence of words and the second variable represents a second sequence of words. One or more machine learning models may generate one or more first candidate words and one or more second candidate words respectively based on the first and second variables. One or more generative machine learning models may generate a predicted first sequence of words and second sequence of words respectively based on the one or more first candidate words and one or more second candidate words. Based on a similarity metric determined based on the predicted sequences and a merging condition, the system may generate a merging instruction indicating whether the first data and the second data are to be merged.