Entity Resolution Language Model Pre-Training With Contrastive Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Entity record resolution in large databases is challenging due to the lack of standardization, multiple sources, and errors in entity records, leading to inefficiencies in deduplication, linkage, and canonicalization, and requiring laborious manual rule updates.

Innovation Solution

A self-contrastive machine learning model with contrastive loss optimization is used to augment entity records and determine positive and negative contrasts, eliminating the need for manual updates and improving entity resolution efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual rule updates are used for entity resolution, then the system can handle non-standardized data, but the maintenance complexity and time consumption increase significantly

Engineering Contradiction:
Improveability to handle non-standardized entity dataVSAvoidmanual rule update maintenance
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs self-contrastive learning where the model automatically generates positive and negative contrast pairs from entity records and performs self-training without manual rule updates. The contrastive loss function enables the model to learn discriminative features autonomously, eliminating the need for continuous manual maintenance of resolution rules.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the entity resolution problem from rule-based parameter matching to learned parameter representations. By using contrastive loss optimization, the model learns optimal parameter configurations that automatically adapt to non-standardized data patterns, replacing manual rule parameter updates with automated parameter learning.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If traditional entity resolution methods are used, then processing accuracy can be maintained, but processing time and computational resources increase

Engineering Contradiction:
Improveentity resolution accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs pre-training of the language model on entity records before actual resolution tasks. This preliminary contrastive learning phase enables the model to learn discriminative patterns in advance, allowing faster and more accurate entity resolution during inference without requiring complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical rule-based entity resolution systems with a machine learning-based semantic understanding system. The contrastive loss optimization enables the model to automatically capture semantic similarities and differences, substituting manual rule processing with intelligent automated decision-making that is both faster and more accurate.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If extensive data storage is used to maintain entity records, then data completeness is improved, but storage costs and processing overhead increase

Engineering Contradiction:
Improveentity record completenessVSAvoiddata storage volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the essential discriminative features from entity records through contrastive learning, rather than storing and processing all raw data. The pre-trained language model learns to represent entities in compressed vector spaces that capture the most important distinguishing characteristics, reducing storage requirements while maintaining resolution accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12468675B2Computer-based systems configured to pre-train language models for entity resolution and methods of use thereof
Publication Date: 2025.11.11 CAPITAL ONE SERVICES LLC
  • US12468675B2 patent drawing
  • US12468675B2 patent drawing
  • US12468675B2 patent drawing

AI summary

In some embodiments, the present disclosure provides an exemplary method that may include steps of receiving a dataset of entity records, identifying, a candidate entity record of the plurality of entity records, utilizing a set of predefined rules to generate a first augmented record, and a second augmented record, utilizing, at least one contrastive loss optimization functions to train parameters of an unsupervised self-contrastive machine learning language model to distinguish between similar entity records representing a same entity and dissimilar entity records.