Adversarial attack based deep learning model weak label vulnerability mining method and system

CN117892317BActive Publication Date: 2026-09-22SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410108773.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2026-09-22
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

例如,对抗性攻击可以通过微小的输入扰动来欺骗深度学习模型,导致其产生错误的输出

Benefits of technology

[0033]与现有技术相比,本发明具有如下的有益效果:本发明通过变幅度对抗样本攻击检测模型中是否存在易转移标签与脆弱标签,验证了对模型的脆弱标签攻击成功率远高于其他标签,对模型的易转移标签攻击成功率远低于其他标签,实现了检测模型中存在弱标签漏洞的技术,并进一步分析了漏洞与样本分布之间的关系。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117892317B_ABST
    Figure CN117892317B_ABST
Patent Text Reader

Abstract

The application provides a kind of deep learning model weak label vulnerability mining method and system based on adversarial attack, uses arbitrary adversarial attack algorithm to mine the vulnerability label of artificial intelligence classification model, and uses other adversarial attack method to carry out target adversarial attack on the success rate change of vulnerability label and verify the vulnerability of model. Including: selecting arbitrary adversarial attack method to carry out amplitude variation adversarial sample attack, extracting easy transfer label and fragile label according to attack result, replacing other adversarial attack algorithm to attack two kinds of labels. The method innovatively proposes and verifies that the result obtained by using any kind of adversarial attack algorithm has migration when using other attack algorithms, experiments are carried out on public models, and it is verified that the method can effectively extract model vulnerability label, is a kind of whole process of mining model vulnerability mining method, has the advantages of universality, good effect, strong explainability, etc., and has practical application value in the field of artificial intelligence classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information security technology, and more specifically, to a method and system for discovering weak-labeled vulnerabilities using deep learning models based on adversarial attacks. Background Technology

[0002] Deep learning is a machine learning method based on artificial neural networks that learns feature representations of data through multi-level nonlinear transformations. With the widespread application of deep learning networks in object recognition, action recognition, video classification, and other fields, attacks on deep learning models pose a significant threat. For example, adversarial attacks can deceive deep learning models through minute input perturbations, causing them to produce incorrect outputs. Such attacks could have real-world implications; for instance, if a deep learning model used in an autonomous vehicle is subjected to an adversarial attack, it could lead to misjudging road conditions and causing traffic accidents. In attacks on artificial intelligence classification models, some labels are particularly vulnerable, easily transforming into other labels during attacks, making them more likely to succeed. This invention proposes a method and system for exploiting weak label vulnerabilities in deep learning models.

[0003] Patent document CN115481714A (application number: 202210553722.4) discloses a method and system for generating adversarial examples for a deep learning vulnerability detection model based on deep reinforcement learning. This method acquires information about the target vulnerability detection model and a set of prototype vulnerability codes, constructs adversarial code transformations, and uses a deep reinforcement learning framework to generate optimal adversarial examples for the target vulnerability detection model. This invention differs significantly from this patent in its research object and method. This invention targets weakly labeled vulnerabilities, uses variable-amplitude adversarial example attacks to extract easily transferable and vulnerable labels, and determines vulnerabilities based on the changes in the attack success rates of various adversarial methods on these two types of labels. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for discovering weak label vulnerabilities in deep learning models based on adversarial attacks.

[0005] A method for discovering weak label vulnerabilities in deep learning models based on adversarial attacks, provided by the present invention, includes:

[0006] Step S1: Obtain the model to be evaluated as the target model;

[0007] Step S2: Select any adversarial attack method to perform variable amplitude adversarial sample attacks on the target model, and record the changes in the recognition results and the migration of sample recognition labels when the attack amplitude changes;

[0008] Step S3: Calculate the confusion matrix of the target model for sample classification before and after the attack based on the recorded results;

[0009] Step S4: Extract two types of tags based on the confusion matrix: easily transferable tags and fragile tags;

[0010] Step S5: Target the easily transferable and vulnerable tags and use other adversarial attack methods to attack the target model;

[0011] Step S6: If the success rate of attacking the easily transferable label is lower than the sample recognition success rate of other labels and meets the preset conditions, and the success rate of attacking the vulnerable label is higher than the sample recognition success rate of other labels and meets the preset conditions, then the current target model is considered to have a weak label vulnerability; select another adversarial attack method and repeat steps S2 to S6.

[0012] The easily transferable tag is the tag that is used as the transfer endpoint the most times;

[0013] The vulnerable label is the label that is used as the starting point for the most transfers.

[0014] Preferably, adversarial attack methods include: gradient-based FGSM, PGD, BIM; optimization-based CW; and black-box attacks based on model opacity such as Pixel and OnePixel.

[0015] Preferably, step S2 involves: attacking adversarial samples with different attack amplitudes, recording the changes in recognition results and the migration of sample recognition labels as the attack amplitude changes, exploring the label distribution pattern in the sample space, and the uniformity of target model classification in the sample space.

[0016] Preferably, step S3 involves: calculating the confusion matrix of the classification results of the original target model and the target model after the attack; including: placing the sample with the real label i and the identification label j into the i-th row and j-th column of the confusion matrix, and recording the identification result of the label after each attack process through variable amplitude adversarial sample attack.

[0017] Preferably, step S5 involves: selecting multiple adversarial attacks from adversarial example attack methods that simultaneously support both targeted and non-targeted modes, designing a targeted attack system, including three modes: original attack, attack targeting easily transferable tags, and attack targeting vulnerable tags; the targeted attack system will output the sample recognition success rate of the model to be evaluated under each adversarial attack in the three attack modes, given by the following formula:

[0018]

[0019] A system for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks, according to the present invention, includes:

[0020] Module M1: Obtain the model to be evaluated as the target model;

[0021] Module M2: Select any adversarial attack method to perform variable amplitude adversarial sample attacks on the target model, and record the changes in the recognition results and the migration of sample recognition labels when the attack amplitude changes;

[0022] Module M3: Calculates the confusion matrix of the target model's classification of samples before and after the attack based on the recorded results;

[0023] Module M4: Extracts two types of tags based on the confusion matrix: easily transferable tags and fragile tags;

[0024] Module M5: Targets easily transferable and vulnerable tags and uses other adversarial attack methods to attack the target model;

[0025] Module M6: If the success rate of attacking easily transferable tags is lower than the success rate of identifying samples of other tags and meets the preset conditions, and the success rate of attacking vulnerable tags is higher than the success rate of identifying samples of other tags and meets the preset conditions, then the current target model is considered to have a weak tag vulnerability; select another adversarial attack method and repeatedly trigger modules M2 to M6.

[0026] The easily transferable tag is the tag that is used as the transfer endpoint the most times;

[0027] The vulnerable label is the label that is used as the starting point for the most transfers.

[0028] Preferably, adversarial attack methods include: gradient-based FGSM, PGD, BIM; optimization-based CW; and black-box attacks based on model opacity such as Pixel and OnePixel.

[0029] Preferably, module M2 employs the following methods: attacking adversarial samples with different attack amplitudes, recording the changes in recognition results and the migration of sample recognition labels as the attack amplitude changes, exploring the label distribution pattern in the sample space, and the uniformity of target model classification in the sample space.

[0030] Preferably, module M3 employs the following: calculating the confusion matrix of the classification results of the original target model and the target model after the attack; including: placing samples with the true label i and the identification label j into the i-th row and j-th column of the confusion matrix, and recording the identification result of the label after each attack process through variable amplitude adversarial sample attacks.

[0031] Preferably, module M5 employs the following approach: selecting multiple adversarial attacks from adversarial example attack methods that simultaneously support both targeted and non-targeted modes, and designing a targeted attack system, including three modes: original attack, attack targeting easily transferable tags, and attack targeting vulnerable tags; the targeted attack system will output the sample recognition success rate of the model under evaluation in each of the three attack modes for each adversarial attack, given by the following formula:

[0032]

[0033] Compared with the prior art, the present invention has the following beneficial effects: The present invention detects whether there are easily transferable labels and vulnerable labels in the model through variable amplitude adversarial sample attacks, verifies that the success rate of attacking the vulnerable labels of the model is much higher than that of other labels, and the success rate of attacking the easily transferable labels of the model is much lower than that of other labels, realizes the technology of detecting weak label vulnerabilities in the model, and further analyzes the relationship between vulnerabilities and sample distribution. Attached Figure Description

[0034] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0035] Figure 1 This is a flowchart of a method for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks.

[0036] Figure 2 This is a schematic diagram of a deep learning model weak label vulnerability discovery system based on adversarial attacks. Detailed Implementation

[0037] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0038] Example 1

[0039] This invention provides a method and system for mining weak label vulnerabilities in deep learning models based on adversarial attacks. It utilizes any adversarial attack method to perform variable-amplitude attacks on the target model, extracts easily transferable and vulnerable labels from the confusion matrix obtained from the attacks, and uses these easily transferable and vulnerable labels as targets to vary the attack success rate by changing different adversarial attack methods. This verifies whether the target model has weak label vulnerabilities.

[0040] The method for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks, taking the detection of the VGG-16 deep learning model as an example, has the following main process: Figure 1 As shown, the implementation results are as follows: Figure 2 As shown, it includes the following steps:

[0041] Step 1: Obtain the model to be evaluated as the target model; in this embodiment, the VGG-16 deep learning model is obtained as the target model.

[0042] Step 2: Select any adversarial attack algorithm to perform variable amplitude adversarial sample attacks on the target model; In this embodiment, the FGSM adversarial attack algorithm is selected to perform variable amplitude adversarial sample attacks on the target model. Different attack amplitudes are used to perform adversarial sample attacks, and the changes in the recognition results and the migration of sample recognition labels are recorded when the attack amplitude changes.

[0043] Step 3: Calculate the confusion matrix of the target model for sample classification before and after the attack based on the recorded results;

[0044] Step 4: Extract two types of tags based on the confusion matrix: easily transferable tags and vulnerable tags. Easily transferable tags are those that are easily migrated from other tags to this tag during an attack, i.e., tags that serve as the most frequent migration endpoints. Vulnerable tags are those that are easily migrated to other tags during an attack, i.e., tags that serve as the most frequent migration start points.

[0045] Step 5: Using the two types of labels as targets, attack the target model using other adversarial attack algorithms; In this embodiment, the PGD and CW adversarial attack algorithms are used to attack the target model respectively, implementing three modes: original attack, attack targeting easily transferable labels, and attack targeting vulnerable labels, and calculating the sample recognition success rate of each adversarial attack algorithm under the three attack modes.

[0046] Step 6: If the success rate of attacking easily transferable labels is significantly lower than the success rate of identifying samples with other labels, while the success rate of attacking vulnerable labels is significantly higher than the success rate of identifying samples with other labels, then the model is considered to have a weak label vulnerability. Change the adversarial attack method in Step 2, observe whether the easily transferable labels and vulnerable labels detected by different methods are consistent, and repeat Step 2 to Step 6 to avoid testing errors.

[0047] Step 7: Input samples with an equal number of different labels, extract the sample vectors from the last convolutional layer and fully connected layer of the model, calculate the distance between the sample vectors, and perform PCA to reduce the dimensionality to a two-dimensional plane for visualization, and observe the vector distance between samples with different labels.

[0048] Step 8: Further analyze the relationship between model vulnerabilities and sample distribution using the vector distance from Step 7, such as whether weakly labeled samples are too close to other samples, and whether transferred-label samples are too far from other samples.

[0049] Specifically, step 2 involves selecting any algorithm from an adversarial example attack algorithm library, including gradient-based algorithms such as FGSM, PGD, and BIM; optimization-based algorithms such as CW; and black-box attacks based on model opacity such as Pixel and OnePixel. The library includes adversarial example attacks for each category. Attacks are performed using different attack magnitudes, and the changes in recognition results and the migration of sample recognition labels are recorded as the attack magnitude changes. This process explores the label distribution patterns in the sample space and the uniformity of model classification in the sample space.

[0050] Specifically, step 3 involves: calculating the confusion matrix between the original model and the model after the attack, i.e., placing the sample with the true label i and the identification label j into the confusion matrix i row and j column, and recording the identification result of the label after each attack process through variable amplitude adversarial sample attack.

[0051] Specifically, step 4 involves: using a confusion matrix to determine the pattern in the identification process of the classification model to be evaluated where a certain type of label is more easily identified as another fixed label; using different thresholds at which the classification results of samples begin to change rapidly to indicate the characteristics of the model to be evaluated in dividing labels in the sample space; and extracting two types of vulnerability labels: easily transferable labels and vulnerable labels. Easily transferable labels refer to labels that are easy to migrate from other labels to this label during an attack, while vulnerable labels refer to labels that are easy to migrate to other labels during an attack.

[0052] Specifically, step 5 employs the following:

[0053] Multiple adversarial attack algorithms supporting both targeted and non-targeted modes were selected from a set of adversarial attack algorithms. A targeted attack system was designed, featuring three modes: original attack, attack targeting easily transferable tags, and attack targeting vulnerable tags. The targeted attack system will output the sample recognition success rate of the model under evaluation in each of the three attack modes of each adversarial attack algorithm, given by the following formula:

[0054]

[0055] Specifically, step 6 involves: randomly selecting other types of attack methods, and attacking the target model with the vulnerability labels obtained in step 4 using the three attack modes in step 5. If the success rate of attacking the easily transferable label is significantly lower than the sample recognition success rate of other labels, while the success rate of attacking the vulnerable label is significantly higher than the sample recognition success rate of other labels, then it can be verified that the known attack results of a certain adversarial sample attack algorithm have a guiding role for other attack algorithms based on the same principle. It is considered that the model to be evaluated has a universal vulnerability, that is, the model exhibits the same vulnerability label for different attack algorithms.

[0056] The present invention also provides a system for mining weak labels of deep learning models based on adversarial attacks. The system for mining weak labels of deep learning models based on adversarial attacks can be implemented by executing the process steps of the method for mining weak labels of deep learning models based on adversarial attacks. That is, those skilled in the art can understand the method for mining weak labels of deep learning models based on adversarial attacks as a preferred embodiment of the system for mining weak labels of deep learning models based on adversarial attacks.

[0057] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0058] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks, characterized in that, include: Step S1: Obtain the model to be evaluated as the target model; Step S2: Select any adversarial attack method to perform variable amplitude adversarial sample attacks on the target model, and record the changes in the recognition results and the migration of sample recognition labels when the attack amplitude changes; Step S3: Calculate the confusion matrix of the target model for sample classification before and after the attack based on the recorded results; Step S4: Extract two types of tags based on the confusion matrix: easily transferable tags and fragile tags; Step S5: Target the easily transferable and vulnerable tags and use other adversarial attack methods to attack the target model; Step S6: If the success rate of attacking the easily transferable label is lower than the sample recognition success rate of other labels and meets the preset conditions, and the success rate of attacking the vulnerable label is higher than the sample recognition success rate of other labels and meets the preset conditions, then the current target model is considered to have a weak label vulnerability; select another adversarial attack method and repeat steps S2 to S6. The easily transferable tag is the tag that is used as the transfer endpoint the most times; The vulnerable label is the label that is used as the starting point for the most transfers.

2. The method for weak label vulnerability mining in deep learning models based on adversarial attacks according to claim 1, characterized in that, Adversarial attack methods include: gradient-based FGSM, PGD, and BIM; optimization-based CW; and black-box attacks based on model opacity such as Pixel and OnePixel.

3. The method for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks according to claim 1, characterized in that, Step S2 involves: attacking adversarial samples with different attack amplitudes, recording the changes in recognition results and the migration of sample recognition labels as the attack amplitude changes, exploring the label distribution pattern in the sample space, and the uniformity of target model classification in the sample space.

4. The method for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks according to claim 1, characterized in that, Step S3 involves: calculating the confusion matrix of the classification results of the original target model and the target model after the attack; including: placing the sample with the real label i and the identification label j into the i-th row and j-th column of the confusion matrix, and recording the identification result of the label after each attack process through variable amplitude adversarial sample attack.

5. The method for weak label vulnerability mining in deep learning models based on adversarial attacks according to claim 1, characterized in that, Step S5 involves selecting multiple adversarial attacks from adversarial example attack methods that simultaneously support both targeted and non-targeted modes, and designing a targeted attack system, including three modes: original attack, attack targeting easily transferable tags, and attack targeting vulnerable tags. The targeted attack system will output the sample recognition success rate of the model to be evaluated under each of the three attack modes for each adversarial attack, given by the following formula:

6. A system for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks, characterized in that, include: Module M1: Obtain the model to be evaluated as the target model; Module M2: Select any adversarial attack method to perform variable amplitude adversarial sample attacks on the target model, and record the changes in the recognition results and the migration of sample recognition labels when the attack amplitude changes; Module M3: Calculates the confusion matrix of the target model's classification of samples before and after the attack based on the recorded results; Module M4: Extracts two types of tags based on the confusion matrix: easily transferable tags and fragile tags; Module M5: Targets easily transferable and vulnerable tags and uses other adversarial attack methods to attack the target model; Module M6: If the success rate of attacking easily transferable tags is lower than the success rate of identifying samples of other tags and meets the preset conditions, and the success rate of attacking vulnerable tags is higher than the success rate of identifying samples of other tags and meets the preset conditions, then the current target model is considered to have a weak tag vulnerability; select another adversarial attack method and repeatedly trigger modules M2 to M6. The easily transferable tag is the tag that is used as the transfer endpoint the most times; The vulnerable label is the label that is used as the starting point for the most transfers.

7. The deep learning model weak label vulnerability mining system based on adversarial attacks according to claim 6, characterized in that, Adversarial attack methods include: gradient-based FGSM, PGD, and BIM; optimization-based CW; and black-box attacks based on model opacity such as Pixel and OnePixel.

8. The system for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks according to claim 6, characterized in that, The module M2 employs the following methods: attacking adversarial samples with different attack amplitudes, recording the changes in recognition results and the migration of sample recognition labels as the attack amplitude changes, exploring the label distribution patterns in the sample space, and the uniformity of the target model classification in the sample space.

9. The deep learning model weak label vulnerability mining system based on adversarial attacks according to claim 6, characterized in that, The module M3 employs the following: calculating the confusion matrix of the classification results of the original target model and the target model after the attack; including: placing samples with the true label i and the identification label j into the i-th row and j-th column of the confusion matrix, and recording the identification result of the label after each attack process through variable amplitude adversarial sample attacks.

10. The system for discovering weakly labeled vulnerabilities in deep learning models based on adversarial attacks according to claim 6, characterized in that, Module M5 employs the following approach: selecting multiple adversarial attacks from adversarial example attack methods that simultaneously support both targeted and non-targeted modes, and designing a targeted attack system, including three modes: original attack, attack targeting easily transferable tags, and attack targeting vulnerable tags. The targeted attack system will output the sample recognition success rate of the model under evaluation in each of the three attack modes for each adversarial attack, given by the following formula:

Citation Information

Patent Citations

  • Deep learning vulnerability detection model adversarial sample generation method and system based on deep reinforcement learning

    CN115481714A

  • Method and system for enhancing anti-attack capability of model based on adversarial samples

    CN111046394A

  • Method and system for generating adversarial sample

    CN113822442A