A defense method and system against big data poisoning attack
By collecting raw datasets in a big data environment and adding a trusted label column, removing the trusted label column, and calculating the feature value of each sample in the labeled big data dataset, this method solves the problem that existing technologies are difficult to effectively defend against big data poisoning attacks. It reduces the risk of pollution in the early stages of the data lifecycle, directly removes potential poisoning samples from the training data, reduces the pollution ratio of the model, and ensures model performance and decision accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUANENG POWER INT INC
- Filing Date
- 2026-03-24
- Publication Date
- 2026-07-03
AI Technical Summary
In the context of big data environments, existing technologies for data poisoning attacks are insufficient to address defense methods and systems against such attacks. This paper addresses a method and system for defending against big data poisoning attacks, based on existing technologies for such attacks.
By collecting the original large dataset and adding a credibility label column, removing labels with credibility values lower than the preset credibility label column, and then removing the credibility label column again, a labeled large dataset is obtained. Further removal of the credibility label column yields a high-credibility label column, and this high-credibility large dataset is obtained. The credibility values corresponding to each sample in the high-credibility large dataset are used as sample weights, and the total loss value is calculated in conjunction with the loss function of the big data model to be trained. Based on the total loss value, the parameters of the model to be trained are optimized, resulting in a big data model resistant to poisoning attacks. This resistant big data model is used to defend against big data poisoning attacks.
This approach reduces the risk of contamination in the early stages of the data lifecycle. By calculating the confidence value of each sample in the labeled large dataset and removing samples with confidence values less than a preset confidence threshold, potential poisoned samples are directly removed from the training data. This reduces the contamination rate of the training dataset, prevents the model from experiencing performance degradation and decision-making errors due to learning from poisoned data, and ultimately yields a big data model resistant to poisoning attacks.
Smart Images

Figure CN122333478A_ABST