Data Curation Pipeline for Aggregate Poisoned Data Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data curation methods fail to effectively detect and mitigate large-scale, statistically consistent anomalous data and poisoned data introduced by malicious parties, which can skew computer-implemented services.
Innovation Solution
Implementing data aggregation, anomaly detection at an aggregate level, and global optimization using genetic algorithms to identify and remediate poisoned data portions, ensuring reliable data provision to downstream consumers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data aggregation and anomaly detection are implemented, then the ability to detect poisoned data is improved, but the device complexity increases
Solution Approach 1:
The patent segments the data curation process into distinct phases: data aggregation, anomaly detection, and optimization. By dividing the complex task of poisoned data identification into manageable components, the system can implement sophisticated detection mechanisms without overwhelming complexity in any single module.
Solution Approach 2:
The patent introduces an intermediary optimization process that acts as a bridge between anomaly detection and final poisoned data identification. This intermediary layer processes the output of anomaly detection and refines it through optimization, allowing the system to maintain high detection capability while managing overall complexity through modular architecture.
2Measurement precision
If global optimization using genetic algorithms is used to identify poisoned data portions, then the precision of poisoned data identification is improved, but the loss of time increases
Solution Approach 1:
The patent performs preliminary data aggregation and anomaly detection before applying the computationally intensive global optimization process. By pre-processing the data and identifying potential anomalies first, the optimization algorithm operates on a reduced and more focused dataset, improving precision while reducing the overall time required.
Solution Approach 2:
The patent applies global optimization specifically to portions of data identified as potentially poisoned through preliminary anomaly detection, rather than applying it to the entire dataset. This partial application of the optimization process maintains high precision for critical data portions while significantly reducing the time cost compared to exhaustive optimization of all data.
Data Source
AI summary
Methods and systems for curating data from data sources are disclosed. Data may be curated from various data sources before being stored in a repository and/or supplied to downstream consumers. The downstream consumers may rely on the trustworthiness of the curated data to provide desired computer-implemented services. During the data curation process, collected data may undergo quality control processes such as anomaly detection that may identify anomalies in the data. The identified anomalies may indicate the presence of poisoned data that, if provided to downstream consumers, may negatively impact the computer-implemented services facilitated by the downstream consumers. When poisoned data is detected among the data, portions of the data affected by the poisoned data (e.g., the poisoned portions) may be identified using an optimization process. The poisoned data may be used to identify and initiate performance of an action set that may reduce the impact of the poisoned data.


