Data Curation Using Source Characteristics to Detect Poisoned Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data curation methods fail to effectively detect and mitigate large-scale, statistically consistent anomalous data and poisoned data designed to evade detection, posing a significant threat to downstream consumers of computer-implemented services.

Innovation Solution

Implementing data aggregation, anomaly detection at an aggregate level, and optimization processes using fitness analysis functions to identify poisoned data and malicious data sources, followed by remedial actions to manage their impact.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional anomaly detection methods are used on individual data points, then small-scale anomalies can be detected, but large-scale statistically consistent poisoned data cannot be detected

Engineering Contradiction:
Improveanomaly detection capabilityVSAvoiddetection coverage scope
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transitions from detecting anomalies at the individual data point level to detecting anomalies at the aggregate dataset level. By analyzing statistical properties of aggregated data from multiple sources, the system can identify large-scale poisoned data that maintains statistical consistency within individual sources but shows anomalies when aggregated across sources.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines data from multiple data sources and performs anomaly detection on the aggregated dataset. By merging data across sources and analyzing collective statistical properties, the system achieves both detection precision for individual anomalies and coverage for large-scale poisoned data that would be invisible when analyzing sources separately.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If data aggregation is performed to detect large-scale poisoned data, then detection coverage improves, but computational complexity increases

Engineering Contradiction:
Improvedetection coverageVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts and analyzes only the essential statistical properties (means, variances, distributions) of aggregated data rather than processing complete raw datasets. This extraction of key statistical features enables detection of large-scale poisoned data while significantly reducing computational complexity compared to analyzing all raw data points.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameters being analyzed from individual data point features to aggregate statistical parameters (mean, variance, distribution characteristics). This parameter transformation allows the system to detect large-scale anomalies with reduced computational burden by working with summarized statistical measures rather than raw data.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If statistical methods are used for anomaly detection, then detection accuracy improves, but poisoned data designed to be statistically consistent evades detection

Engineering Contradiction:
Improveanomaly detection accuracyVSAvoidevadability of poisoned data
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent detects anomalies in the aggregate statistical properties of combined data from multiple sources, rather than in individual data points. Poisoned data that is statistically consistent within its source may exhibit anomalies when aggregated with other sources, enabling detection in a different statistical dimension that bypasses the evadability of traditional methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system uses feedback from aggregate analysis to identify and mitigate poisoned data sources. By continuously monitoring aggregated statistical properties and comparing against expected patterns, the system can detect deviations caused by poisoned data and adjust detection strategies accordingly, improving accuracy against evasive attacks.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12405930B2System and method for identifying poisoned data during data curation using data source characteristics
Publication Date: 2025.09.02 DELL PROD LP
  • US12405930B2 patent drawing
  • US12405930B2 patent drawing
  • US12405930B2 patent drawing

AI summary

Methods and systems for curating data from data sources are disclosed. Data may be curated from various data sources before being supplied to downstream consumers that may rely on the trustworthiness of the curated data to facilitate desired computer-implemented services. During data curation, collected data may undergo anomaly detection to identify anomalies in the data. Data anomalies may indicate the presence of poisoned data that, if provided to downstream consumers, may negatively impact the desired computer-implemented services. When poisoned data is detected among the data, a poisoned portion of the data may be identified using an optimization process. The optimization process may consider the degree of anomalousness of the data (e.g., using statistical representations of the anomaly) and/or characteristics of the data source that supplied the anomalous data to identify the poisoned portion. Remedial actions may be identified and/or performed in order to reduce an impact of the poisoned data.