Data Curation Pipeline for Aggregate Poisoned Data Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data curation methods fail to effectively detect and mitigate large-scale, statistically consistent anomalous data and poisoned data introduced by malicious parties, which can skew computer-implemented services.

Innovation Solution

Implementing data aggregation, anomaly detection at an aggregate level, and global optimization using genetic algorithms to identify and remediate poisoned data portions, ensuring reliable data provision to downstream consumers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data aggregation and anomaly detection are implemented, then the ability to detect poisoned data is improved, but the device complexity increases

Engineering Contradiction:
Improveability to detect poisoned dataVSAvoidcomplexity of detection system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the data curation process into distinct phases: data aggregation, anomaly detection, and optimization. By dividing the complex task of poisoned data identification into manageable components, the system can implement sophisticated detection mechanisms without overwhelming complexity in any single module.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary optimization process that acts as a bridge between anomaly detection and final poisoned data identification. This intermediary layer processes the output of anomaly detection and refines it through optimization, allowing the system to maintain high detection capability while managing overall complexity through modular architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If global optimization using genetic algorithms is used to identify poisoned data portions, then the precision of poisoned data identification is improved, but the loss of time increases

Engineering Contradiction:
Improveprecision of poisoned data identificationVSAvoidtime for optimization process
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary data aggregation and anomaly detection before applying the computationally intensive global optimization process. By pre-processing the data and identifying potential anomalies first, the optimization algorithm operates on a reduced and more focused dataset, improving precision while reducing the overall time required.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies global optimization specifically to portions of data identified as potentially poisoned through preliminary anomaly detection, rather than applying it to the entire dataset. This partial application of the optimization process maintains high precision for critical data portions while significantly reducing the time cost compared to exhaustive optimization of all data.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12423280B2System and method for identifying poisoned data during data curation
Publication Date: 2025.09.23 DELL PROD LP
  • US12423280B2 patent drawing
  • US12423280B2 patent drawing
  • US12423280B2 patent drawing

AI summary

Methods and systems for curating data from data sources are disclosed. Data may be curated from various data sources before being stored in a repository and/or supplied to downstream consumers. The downstream consumers may rely on the trustworthiness of the curated data to provide desired computer-implemented services. During the data curation process, collected data may undergo quality control processes such as anomaly detection that may identify anomalies in the data. The identified anomalies may indicate the presence of poisoned data that, if provided to downstream consumers, may negatively impact the computer-implemented services facilitated by the downstream consumers. When poisoned data is detected among the data, portions of the data affected by the poisoned data (e.g., the poisoned portions) may be identified using an optimization process. The poisoned data may be used to identify and initiate performance of an action set that may reduce the impact of the poisoned data.