Multi-purpose Data Management via Contrastive Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems are limited to specific tasks and require large labeled datasets for training, making them inefficient for multi-purpose data management tasks and costly to adjust for changes in data format or sources.

Innovation Solution

A system that uses a contrastive learning technique to determine representative entity pairs in a shared space, allowing for data management tasks like integration, cleanup, and discovery without manual labeling, by training machine learning models using unlabeled data and existing transformer models like BERT.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing natural language processing systems are used for specific tasks, then task performance is improved, but system versatility and adaptability to different data management tasks deteriorate

Engineering Contradiction:
Improvetask performanceVSAvoidsystem versatility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by developing a single NLP system framework that can perform multiple data management tasks including entity matching, data integration, data cleanup, and data discovery. The system uses a unified architecture with contrastive learning that adapts to different tasks without requiring separate specialized systems, thereby achieving multi-functionality while maintaining task performance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamics through the use of contrastive learning that dynamically adapts to different data management tasks. The system can adjust its learning objectives and parameters based on the specific task requirements, allowing it to transition between different functions while maintaining optimal performance for each task.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If large labeled datasets are used for training, then model accuracy is improved, but training cost and time consumption deteriorate

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training the model using contrastive learning on unlabeled data before fine-tuning on task-specific labeled data. This pre-training phase establishes useful representations that reduce the amount of labeled data needed and accelerate subsequent training, thereby reducing both training time and resource requirements while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service through contrastive learning that automatically learns meaningful representations from unlabeled data without requiring manual annotation. The model serves itself by generating its own training signals through the contrastive objective, eliminating the need for expensive and time-consuming manual labeling while still achieving high accuracy.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If custom NLP systems are developed for each task, then task-specific performance is improved, but system complexity and maintenance cost deteriorate

Engineering Contradiction:
Improvetask-specific performanceVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent reduces system complexity by implementing a universal NLP framework that handles multiple data management tasks through a single architecture. This unified system eliminates the need to develop and maintain separate custom systems for each task, reducing overall complexity while preserving task-specific performance through adaptable learning objectives.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines multiple task-specific systems into a single unified framework by merging the entity matching, data integration, cleanup, and discovery functionalities into one system. This consolidation reduces maintenance costs and simplifies the overall system architecture while maintaining the ability to perform each specific task effectively.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240289629A1Systems and methods for multi-purpose data management
Publication Date: 2024.08.29 RECRUIT
  • US20240289629A1 patent drawing
  • US20240289629A1 patent drawing
  • US20240289629A1 patent drawing

AI summary

Disclosed embodiments relate to data management of entity pairs. Techniques can include receiving at least two sets of data and a data management task request with each including a set of entities. Techniques can determine a location of each entity in received data sets in a representative space by determining representative structure of the set of entities. Techniques can then for an entity, a set of representative entity pairs from each set of the at least two sets of data based on how close they are in the representative space. Technique can then analyze the set of representative entity pairs to identify most similar entity pairs include in a set of candidate pairs by determining closeness of location of entities in each entity pair in the representative space. Technique can then determine matched entity pairs of the candidate pairs using a first machine learning model is trained using the candidate pairs by applying labels, and utilizing the matched pairs to perform the requested data management task.