Data Platform Embedding Clustering for Redundant Record Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large data platforms like OSDU face issues with redundant data, including duplicate records and versions, which slow down services and complicate security and backup management due to the need for manual and cumbersome processes to identify similar records, lacking effective machine-learning solutions.

Innovation Solution

A method and system using machine-learning techniques, specifically auto-encoders or large language models, convert data records into embeddings, apply clustering algorithms, and plot them on a graph to identify and delete similar records based on similarity thresholds, reducing redundancy and enhancing service efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data records are stored in the data platform, then data availability is improved, but platform complexity increases due to redundant data

Engineering Contradiction:
Improvedata availabilityVSAvoidplatform complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes redundant, duplicate, and orphan records from the data platform. By identifying and taking out unnecessary data records that duplicate information or lack proper relationships, the system reduces platform complexity while maintaining data availability of unique, valid records.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent discards redundant and duplicate data records that no longer serve a purpose in the platform. By systematically identifying and removing these unnecessary records, the system recovers storage space and reduces complexity while preserving essential data for operational needs.

Inventive Principle:
Principle #34Discarding and recovering

2Reliability

If data records are stored in the data platform, then data completeness is improved, but service performance deteriorates due to processing overhead

Engineering Contradiction:
Improvedata completenessVSAvoidservice performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts and removes redundant and duplicate records that contribute to processing overhead. By taking out these unnecessary records while retaining complete, unique data, the system reduces the workload on services without compromising data completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system implements self-service mechanisms where the data platform automatically identifies and removes redundant records through clustering algorithms and similarity detection. This automated process maintains data completeness while improving service performance by reducing manual intervention and processing overhead.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If manual scanning is used to identify similar records, then detection accuracy is improved, but operation time increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidoperation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual scanning with automated machine-learning-based clustering algorithms. These algorithms automatically detect similar records by converting data into embeddings and applying clustering techniques, thereby maintaining high detection accuracy while dramatically reducing the time required for identifying redundant records.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system employs self-service automated detection mechanisms that continuously identify similar and duplicate records without manual intervention. The machine-learning models automatically perform detection and classification, achieving both high accuracy and efficiency in identifying redundant data.

Inventive Principle:
Principle #25Self-service

4Loss of substance

If clustering algorithms are applied to embeddings, then data redundancy is reduced, but computational complexity increases

Engineering Contradiction:
Improvedata redundancyVSAvoidcomputational complexity
Core Design Contradiction:
Loss of substanceVSDevice complexity

Solution Approach 1:

The patent applies preliminary actions by converting data records into embeddings before applying clustering algorithms. This pre-processing step organizes data in a format that facilitates efficient clustering, reducing data redundancy while managing computational complexity through structured transformation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transforming raw data into embedding representations and adjusting clustering thresholds to optimize the balance between redundancy reduction and computational complexity. By modifying data representation and algorithm parameters, the system achieves effective redundancy reduction with manageable computational requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12572568B2Managing a data platform
Publication Date: 2026.03.10 SCHLUMBERGER TECH CORP
  • US12572568B2 patent drawing
  • US12572568B2 patent drawing
  • US12572568B2 patent drawing

AI summary

A method for managing a data platform includes converting a plurality of data records in the data platform into embeddings. The method also includes applying a clustering algorithm to the embeddings to identify a first subset of the embeddings corresponding to a first subset of the data records and a second subset of the embeddings corresponding to a second subset of data records. The method also includes determining that the first subset of embeddings have a similarity with respect to one another that is within a first similarity threshold. The method also includes deleting one or more of the first subset of data records from the data platform in response to the similarity of the first subset of embeddings being within the first similarity threshold.