Enterprise Data Deduplication via Graph-Based Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large enterprise databases, especially in supply chain applications, duplicate entity data entries with minor differences lead to analysis difficulties and errors in reporting, requiring efficient data processing and normalization techniques to manage distinct data sources and reduce human error.
Innovation Solution
A method involving data processing that includes receiving datasets, normalizing data fields, enriching parent data through mapping, processing with tokenization and vectorization to generate sparse vectors, and clustering based on data attributes, using AI and machine learning for accurate data normalization, de-duplication, and error removal across multiple languages and large datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data processing techniques are used to identify and eliminate duplicate entity data, then data accuracy can be maintained through human review, but processing time and labor resources increase significantly
Solution Approach 1:
The patent replaces manual mechanical data processing with an automated computer-based system that uses machine learning models and algorithms to identify, compare, and eliminate duplicate entity data. The system automatically processes entity data from multiple sources, applies clustering algorithms to group similar records, and uses machine learning classifiers to determine duplicates, thereby eliminating the need for manual human review while maintaining high accuracy.
Solution Approach 2:
The system enables self-service data cleansing by automatically performing data normalization, deduplication, and validation without requiring manual intervention. The machine learning models continuously learn from the data and automatically improve their ability to identify duplicates, allowing the system to serve itself in maintaining data quality across enterprise applications.
2Productivity
If automated data processing techniques are deployed to eliminate duplicates, then processing efficiency improves, but accuracy decreases due to conceptually inaccurate results and high error rates
Solution Approach 1:
The patent implements feedback mechanisms where the machine learning models continuously learn from processing results and user corrections. The system incorporates feedback loops that allow it to refine its duplicate detection algorithms based on actual data patterns and user interactions, thereby improving accuracy while maintaining high processing efficiency. The feedback from data quality assessments and user validations is used to retrain and improve the models.
Solution Approach 2:
The system performs preliminary data normalization and preprocessing actions before the main duplicate detection process. By pre-processing the data to standardize formats, handle missing values, and normalize entity representations beforehand, the system prepares the data in a state that enables more accurate automated duplicate identification, reducing errors in the subsequent processing stages.
3Device complexity
If traditional data processing methods are used to handle entity data from multiple sources, then system complexity remains low, but the ability to capture and process data with different attributes under the same entity is severely limited
Solution Approach 1:
The patent creates a universal data processing platform that can handle entity data from multiple diverse sources including supply chain management systems, enterprise resource planning systems, and external databases. The system uses universal machine learning models and algorithms that can process various data types and formats, adapting to different data sources and attributes while maintaining a consistent processing framework. This multi-functional approach enables the system to capture and process diverse entity attributes under unified entity identification.
4Ease of manufacture
If existing data processing techniques are applied to supply chain data, then general data cleansing can be achieved, but accurate results are difficult to obtain due to the distinct nature of supply chain data attributes
Solution Approach 1:
The patent applies local quality by implementing specialized processing rules and machine learning models tailored to specific supply chain data attributes. Rather than using a one-size-fits-all approach, the system identifies and applies specific processing techniques for different data types such as supplier names, material descriptions, and transaction data. This localized processing approach maintains ease of implementation while significantly improving the accuracy of data cleansing for supply chain-specific data characteristics.
Data Source
AI summary
The present invention relates to data processing system and method in supply chain application. The data processing system includes clustering of received supply chain data after normalization, tokenization and vectorization through graph-based analysis.


