Data Platform Embedding Clustering for Redundant Record Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large data platforms like OSDU face issues with redundant data, including duplicate records and versions, which slow down services and complicate security and backup management due to the need for manual and cumbersome processes to identify similar records, lacking effective machine-learning solutions.
Innovation Solution
A method and system using machine-learning techniques, specifically auto-encoders or large language models, convert data records into embeddings, apply clustering algorithms, and plot them on a graph to identify and delete similar records based on similarity thresholds, reducing redundancy and enhancing service efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data records are stored in the data platform, then data availability is improved, but platform complexity increases due to redundant data
Solution Approach 1:
The patent extracts and removes redundant, duplicate, and orphan records from the data platform. By identifying and taking out unnecessary data records that duplicate information or lack proper relationships, the system reduces platform complexity while maintaining data availability of unique, valid records.
Solution Approach 2:
The patent discards redundant and duplicate data records that no longer serve a purpose in the platform. By systematically identifying and removing these unnecessary records, the system recovers storage space and reduces complexity while preserving essential data for operational needs.
2Reliability
If data records are stored in the data platform, then data completeness is improved, but service performance deteriorates due to processing overhead
Solution Approach 1:
The patent extracts and removes redundant and duplicate records that contribute to processing overhead. By taking out these unnecessary records while retaining complete, unique data, the system reduces the workload on services without compromising data completeness.
Solution Approach 2:
The system implements self-service mechanisms where the data platform automatically identifies and removes redundant records through clustering algorithms and similarity detection. This automated process maintains data completeness while improving service performance by reducing manual intervention and processing overhead.
3Measurement precision
If manual scanning is used to identify similar records, then detection accuracy is improved, but operation time increases
Solution Approach 1:
The patent replaces manual scanning with automated machine-learning-based clustering algorithms. These algorithms automatically detect similar records by converting data into embeddings and applying clustering techniques, thereby maintaining high detection accuracy while dramatically reducing the time required for identifying redundant records.
Solution Approach 2:
The system employs self-service automated detection mechanisms that continuously identify similar and duplicate records without manual intervention. The machine-learning models automatically perform detection and classification, achieving both high accuracy and efficiency in identifying redundant data.
4Loss of substance
If clustering algorithms are applied to embeddings, then data redundancy is reduced, but computational complexity increases
Solution Approach 1:
The patent applies preliminary actions by converting data records into embeddings before applying clustering algorithms. This pre-processing step organizes data in a format that facilitates efficient clustering, reducing data redundancy while managing computational complexity through structured transformation.
Solution Approach 2:
The patent changes parameters by transforming raw data into embedding representations and adjusting clustering thresholds to optimize the balance between redundancy reduction and computational complexity. By modifying data representation and algorithm parameters, the system achieves effective redundancy reduction with manageable computational requirements.
Data Source
AI summary
A method for managing a data platform includes converting a plurality of data records in the data platform into embeddings. The method also includes applying a clustering algorithm to the embeddings to identify a first subset of the embeddings corresponding to a first subset of the data records and a second subset of the embeddings corresponding to a second subset of data records. The method also includes determining that the first subset of embeddings have a similarity with respect to one another that is within a first similarity threshold. The method also includes deleting one or more of the first subset of data records from the data platform in response to the similarity of the first subset of embeddings being within the first similarity threshold.


