A defense method and system against big data poisoning attack

By collecting raw datasets in a big data environment and adding a trusted label column, removing the trusted label column, and calculating the feature value of each sample in the labeled big data dataset, this method solves the problem that existing technologies are difficult to effectively defend against big data poisoning attacks. It reduces the risk of pollution in the early stages of the data lifecycle, directly removes potential poisoning samples from the training data, reduces the pollution ratio of the model, and ensures model performance and decision accuracy.

CN122333478APending Publication Date: 2026-07-03HUANENG POWER INT INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUANENG POWER INT INC
Filing Date
2026-03-24
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In the context of big data environments, existing technologies for data poisoning attacks are insufficient to address defense methods and systems against such attacks. This paper addresses a method and system for defending against big data poisoning attacks, based on existing technologies for such attacks.

Method used

By collecting the original large dataset and adding a credibility label column, removing labels with credibility values ​​lower than the preset credibility label column, and then removing the credibility label column again, a labeled large dataset is obtained. Further removal of the credibility label column yields a high-credibility label column, and this high-credibility large dataset is obtained. The credibility values ​​corresponding to each sample in the high-credibility large dataset are used as sample weights, and the total loss value is calculated in conjunction with the loss function of the big data model to be trained. Based on the total loss value, the parameters of the model to be trained are optimized, resulting in a big data model resistant to poisoning attacks. This resistant big data model is used to defend against big data poisoning attacks.

Benefits of technology

This approach reduces the risk of contamination in the early stages of the data lifecycle. By calculating the confidence value of each sample in the labeled large dataset and removing samples with confidence values ​​less than a preset confidence threshold, potential poisoned samples are directly removed from the training data. This reduces the contamination rate of the training dataset, prevents the model from experiencing performance degradation and decision-making errors due to learning from poisoned data, and ultimately yields a big data model resistant to poisoning attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333478A_ABST
    Figure CN122333478A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for defending against big data poisoning attacks, belonging to the field of information security technology. The method includes the following steps: collecting a raw big data dataset to be labeled, and adding a credibility label column to the data to obtain a labeled big data dataset; calculating the credibility value corresponding to each sample in the labeled big data dataset, and removing samples whose credibility values ​​are less than a preset credibility threshold to obtain a high-credibility big data dataset; using the credibility values ​​corresponding to each sample in the high-credibility big data dataset as sample weights, and calculating the total loss value in combination with the loss function of the big data model to be trained; optimizing the parameters of the model to be trained based on the total loss value to obtain a big data model resistant to poisoning attacks, and using the big data model resistant to poisoning attacks to defend against big data poisoning attacks. This invention can solve the problem of training data pollution caused by data poisoning in current big data systems.
Need to check novelty before this filing date? Find Prior Art