Data Curation Using Source Characteristics to Detect Poisoned Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data curation methods fail to effectively detect and mitigate large-scale, statistically consistent anomalous data and poisoned data designed to evade detection, posing a significant threat to downstream consumers of computer-implemented services.
Innovation Solution
Implementing data aggregation, anomaly detection at an aggregate level, and optimization processes using fitness analysis functions to identify poisoned data and malicious data sources, followed by remedial actions to manage their impact.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional anomaly detection methods are used on individual data points, then small-scale anomalies can be detected, but large-scale statistically consistent poisoned data cannot be detected
Solution Approach 1:
The patent transitions from detecting anomalies at the individual data point level to detecting anomalies at the aggregate dataset level. By analyzing statistical properties of aggregated data from multiple sources, the system can identify large-scale poisoned data that maintains statistical consistency within individual sources but shows anomalies when aggregated across sources.
Solution Approach 2:
The patent combines data from multiple data sources and performs anomaly detection on the aggregated dataset. By merging data across sources and analyzing collective statistical properties, the system achieves both detection precision for individual anomalies and coverage for large-scale poisoned data that would be invisible when analyzing sources separately.
2Adaptability or versatility
If data aggregation is performed to detect large-scale poisoned data, then detection coverage improves, but computational complexity increases
Solution Approach 1:
The patent extracts and analyzes only the essential statistical properties (means, variances, distributions) of aggregated data rather than processing complete raw datasets. This extraction of key statistical features enables detection of large-scale poisoned data while significantly reducing computational complexity compared to analyzing all raw data points.
Solution Approach 2:
The patent changes the parameters being analyzed from individual data point features to aggregate statistical parameters (mean, variance, distribution characteristics). This parameter transformation allows the system to detect large-scale anomalies with reduced computational burden by working with summarized statistical measures rather than raw data.
3Measurement precision
If statistical methods are used for anomaly detection, then detection accuracy improves, but poisoned data designed to be statistically consistent evades detection
Solution Approach 1:
The patent detects anomalies in the aggregate statistical properties of combined data from multiple sources, rather than in individual data points. Poisoned data that is statistically consistent within its source may exhibit anomalies when aggregated with other sources, enabling detection in a different statistical dimension that bypasses the evadability of traditional methods.
Solution Approach 2:
The system uses feedback from aggregate analysis to identify and mitigate poisoned data sources. By continuously monitoring aggregated statistical properties and comparing against expected patterns, the system can detect deviations caused by poisoned data and adjust detection strategies accordingly, improving accuracy against evasive attacks.
Data Source
AI summary
Methods and systems for curating data from data sources are disclosed. Data may be curated from various data sources before being supplied to downstream consumers that may rely on the trustworthiness of the curated data to facilitate desired computer-implemented services. During data curation, collected data may undergo anomaly detection to identify anomalies in the data. Data anomalies may indicate the presence of poisoned data that, if provided to downstream consumers, may negatively impact the desired computer-implemented services. When poisoned data is detected among the data, a poisoned portion of the data may be identified using an optimization process. The optimization process may consider the degree of anomalousness of the data (e.g., using statistical representations of the anomaly) and/or characteristics of the data source that supplied the anomalous data to identify the poisoned portion. Remedial actions may be identified and/or performed in order to reduce an impact of the poisoned data.


