Enterprise Data Deduplication via Graph-Based Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In large enterprise databases, especially in supply chain applications, duplicate entity data entries with minor differences lead to analysis difficulties and errors in reporting, requiring efficient data processing and normalization techniques to manage distinct data sources and reduce human error.

Innovation Solution

A method involving data processing that includes receiving datasets, normalizing data fields, enriching parent data through mapping, processing with tokenization and vectorization to generate sparse vectors, and clustering based on data attributes, using AI and machine learning for accurate data normalization, de-duplication, and error removal across multiple languages and large datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data processing techniques are used to identify and eliminate duplicate entity data, then data accuracy can be maintained through human review, but processing time and labor resources increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data processing with an automated computer-based system that uses machine learning models and algorithms to identify, compare, and eliminate duplicate entity data. The system automatically processes entity data from multiple sources, applies clustering algorithms to group similar records, and uses machine learning classifiers to determine duplicates, thereby eliminating the need for manual human review while maintaining high accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data cleansing by automatically performing data normalization, deduplication, and validation without requiring manual intervention. The machine learning models continuously learn from the data and automatically improve their ability to identify duplicates, allowing the system to serve itself in maintaining data quality across enterprise applications.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated data processing techniques are deployed to eliminate duplicates, then processing efficiency improves, but accuracy decreases due to conceptually inaccurate results and high error rates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the machine learning models continuously learn from processing results and user corrections. The system incorporates feedback loops that allow it to refine its duplicate detection algorithms based on actual data patterns and user interactions, thereby improving accuracy while maintaining high processing efficiency. The feedback from data quality assessments and user validations is used to retrain and improve the models.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary data normalization and preprocessing actions before the main duplicate detection process. By pre-processing the data to standardize formats, handle missing values, and normalize entity representations beforehand, the system prepares the data in a state that enables more accurate automated duplicate identification, reducing errors in the subsequent processing stages.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If traditional data processing methods are used to handle entity data from multiple sources, then system complexity remains low, but the ability to capture and process data with different attributes under the same entity is severely limited

Engineering Contradiction:
Improvesystem complexityVSAvoiddata processing capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal data processing platform that can handle entity data from multiple diverse sources including supply chain management systems, enterprise resource planning systems, and external databases. The system uses universal machine learning models and algorithms that can process various data types and formats, adapting to different data sources and attributes while maintaining a consistent processing framework. This multi-functional approach enables the system to capture and process diverse entity attributes under unified entity identification.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of manufacture

If existing data processing techniques are applied to supply chain data, then general data cleansing can be achieved, but accurate results are difficult to obtain due to the distinct nature of supply chain data attributes

Engineering Contradiction:
Improvedata processing easeVSAvoiddata cleansing accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies local quality by implementing specialized processing rules and machine learning models tailored to specific supply chain data attributes. Rather than using a one-size-fits-all approach, the system identifies and applies specific processing techniques for different data types such as supplier names, material descriptions, and transaction data. This localized processing approach maintains ease of implementation while significantly improving the accuracy of data cleansing for supply chain-specific data characteristics.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11675839B2Data processing in enterprise application
Publication Date: 2023.06.13 NB VENTURES INC DBA GEP
  • US11675839B2 patent drawing
  • US11675839B2 patent drawing
  • US11675839B2 patent drawing

AI summary

The present invention relates to data processing system and method in supply chain application. The data processing system includes clustering of received supply chain data after normalization, tokenization and vectorization through graph-based analysis.