A Defense Method for Improving the Robustness of Integrated Models

By learning each other's "vulnerability" between sub-models of deep learning models and combining feature layer mixing methods, the problem of deep learning models identifying errors in the face of noise attacks is solved, and effective defense against white box and black box attacks and maintaining clean sample recognition rates are achieved.

CN113935496BActive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111302450.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-06-13
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

When faced with slight tampering of the input images, existing deep learning models are easily interfered by noise attacks, resulting in identification errors, and existing defense methods cannot effectively defend without affecting the recognition rate of clean samples.

Method used

By extracting non-robust feature samples of all submodels on each training sample, and letting submodels learn each other's "frailty" by training each other's non-robust feature samples generated by each other, combining the feature layer mixing method, all submodels alternately train them to improve the robustness of the integrated model.

Benefits of technology

Effective defense against white and black box attacks basically does not affect the recognition rate of clean samples, and at the same time improves the overall robustness of the integrated model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113935496B_ABST
    Figure CN113935496B_ABST
Patent Text Reader

Abstract

The present invention discloses a robustness improvement defense method for an integrated model, including the following steps: S1: Extract non-robust feature samples of all sub-models on each training sample; S2: Select an untrained sub-model, and input the non-robust feature samples extracted by other sub-models into this sub-model for training respectively; S3: By combining the feature layer mixing method, mix the output values of the non-robust feature samples at the t-th intermediate feature layer of the sub-model being trained in different proportions into an intermediate layer output feature_map; S4: Continue to input the mixed feature_map into the sub-model being trained for forward propagation, and calculate the cross-entropy to update the parameters of this sub-model; S5: Train all sub-models in the integrated model respectively through the above steps S1 to S5 until all sub-models reach the maximum number of training rounds, then the final sub-models are obtained. The integrated model trained by the present invention can not only effectively defend against white-box attack and black-box attack methods, but also basically does not affect the recognition rate of clean samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and more specifically, to a method for enhancing the robustness of an integrated model against attacks. Background Art

[0002] One of the characteristics of a deep neural network model is to express the correspondence between results and features through linear combinations of linear features. Therefore, by slightly tampering with a small amount of content in the input data, a huge change will occur in the extracted features, causing the artificial intelligence system to output incorrect results. This poses a huge threat to the robustness of artificial intelligence systems based on deep learning.

[0003] For current deep learning models, if an attacker slightly tampers with the content of an input image, specific image content cannot be detected or recognized by the artificial intelligence system, posing a great challenge to the security of the artificial intelligence system. The perturbation noise of the tampered image is often relatively small and not easily detected by the human visual system, but it is easy to interfere with the judgment of the artificial intelligence system. Therefore, how to effectively defend against these noise attacks has become one of the urgent problems to be solved by current deep models. However, existing defense methods either have unsatisfactory defense effects against perturbation attacks or sacrifice the recognition rate of clean samples to achieve better defense effects against perturbation attacks. Neither can achieve the expected results, and this problem remains unsolved. Summary of the Invention

[0004] The present invention provides a method for enhancing the robustness of an integrated model against attacks to solve the above-mentioned problems existing in the prior art.

[0005] To achieve the above object of the present invention, the following technical solutions are adopted:

[0006] A method for enhancing the robustness of an integrated model against attacks, the method comprising the following steps:

[0007] S1: Extract non-robust feature samples of all sub-models on each training sample;

[0008] S2: Select an untrained sub-model, and input the non-robust feature samples extracted by other sub-models into the sub-model for training respectively. The sub-models learn each other's "vulnerabilities" through the non-robust feature samples generated by each other;

[0009] S3: By combining the feature layer mixing method, the output values of the non-robust feature samples at the t-th intermediate feature layer of the sub-model being trained are mixed in different proportions to form an intermediate layer output feature_map;

[0010] S4: Input the feature_map obtained by mixing into the sub-model being trained for forward propagation, calculate the cross-entropy to update the parameters of the sub-model;

[0011] S5: Train all sub-models in the ensemble model through the above steps S1 - S5 respectively until all sub-models reach the maximum number of training epochs, then the final sub-models are obtained.

[0012] Preferably, in step S1, before extracting the non-robust feature samples of all sub-models, an initialization operation is first performed as follows: Generate a noise matrix with dimensions h×w×c based on the uniform distribution U(-ε, ε) for the original image x s to perform the initialization operation; where h, w, and c are respectively the height, width, and channel dimensions of the images in the training sample set, and ε represents the maximum pixel value of the added perturbation.

[0013] Furthermore, in step S1, use the feature extraction algorithm to extract the non-robust feature samples of the sub-model on the non-robust feature image z, including the following steps:

[0014] S101: Randomly select another target image x;

[0015] S102: In an iterative manner, approximate the output value of the non-robust feature image z at the feature layer to the output value of the target image x at the feature layer to form the final non-robust feature sample, and its calculation formula is:

[0016]

[0017] In the formula: f i l (·) represents the output value of the l-th layer of the i-th sub-model, where the content in the parentheses represents the input of the model; z i,l represents the non-robust feature sample generated by the i-th sub-model through the l-th feature layer; s.t.||.|| ∞ represents using the infinity norm to constrain the generated non-robust feature sample.

[0018] Still further, in step S3, specifically, when training the i-th sub-model, randomly select the non-robust feature image generated by the j-th sub-model, and mix it with the non-robust feature samples generated by other sub-models except the i-th sub-model and the j-th sub-model at a ratio of λ and γ respectively at the output of the intermediate feature layer of the t-th layer of the i-th sub-model to form an intermediate layer output feature_map.

[0019] Still further, for the feature layer mixing method, its calculation formula is:

[0020]

[0021] Where λ and γ are mixing coefficients. λ is a random matrix coefficient subject to the Beta distribution, and γ is defined as (1 - λ) / (N - 2), where N is the number of sub-models; both t and l are randomly selected feature layers; k is the serial number of any sub-model other than i and j; f i t (z j,l ) is the output value of the non-robust feature sample extracted by the j-th sub-model at the t-th layer of the i-th sub-model; Y i t is the mixed result feature_map of the output values of the non-robust feature samples generated by other sub-models at the intermediate feature layer of the t-th layer of this sub-model when training the i-th sub-model.

[0022] Furthermore, calculate the cross-entropy of the mixed feature_map. The cross-entropy calculation formula is:

[0023]

[0024] Where: M is the number of categories; y s is the true category label of the original image; is the sign function, which takes 1 if the true category label is equal to c, otherwise 0; p c is the probability of being judged as category c.

[0025] Furthermore, use the cross-entropy to update the parameters of the sub-model. The formula is:

[0026]

[0027] Minimizing Equation (4) enables sub-model i to learn the "vulnerability" of other sub-models by learning the non-robust feature samples generated by other sub-models.

[0028] Furthermore, in actual test deployment, for each test sample, input it into all sub-models trained through the above steps S1 to S5 for query at the same time; obtain the prediction results of all sub-models through the query, calculate the mean of the prediction results, and use this mean as the final prediction result.

[0029] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0030] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0031] The beneficial effects of the present invention are as follows:

[0032] The present invention provides a method for enhancing the robustness of an integrated model. First, non-robust features of all sub-models are extracted from all training samples. Then, the sub-models learn each other's "vulnerabilities" by learning each other's non-robust features, thereby reducing the transferability between the sub-models. Finally, combined with the feature-level mixing method, the sub-models can better learn the non-robust features of other sub-models, further increasing the differences between the sub-models. All sub-models are alternately trained. Finally, through multiple sub-models with large differences, the overall robustness of the integrated model can be better improved. The integrated model trained by the present invention can not only effectively defend against white-box attacks and black-box attack methods, but also basically does not affect the recognition rate of clean samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of the method described in Embodiment 1.

[0034] Figure 2 is an overall process example diagram of the defense method proposed in Embodiment 1.

[0035] Figure 3 is an example diagram of the non-robust feature image generation process.

[0036] Figure 4 is an example diagram of the generated result image of the non-robust feature image z.

[0037] Figure 5 is an example diagram of the random feature mixing process.

[0038] Figure 6 is an example diagram of the integrated model update process.

[0039] Figure 7 is a success rate result diagram when the defense method proposed in Embodiment 1 defends against black-box transfer attacks.

[0040] Figure 8 is a success rate result diagram when the defense method proposed in Embodiment 1 defends against white-box attacks. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0042] Embodiment 1

[0043] As Figure 1 , Figure 2 shown, a method for enhancing the robustness of an integrated model, the method includes the following steps:

[0044] S1: Use a feature extraction algorithm to extract non-robust feature samples of all sub-models from each training sample; the process example diagram of step S1 is as Figure 3 shown.

[0045] In a specific embodiment, before extracting the non-robust feature samples of all sub-models, an initialization operation is first performed as follows: Generate a noise matrix with dimensions h×w×c based on the uniform distribution U(-ε,ε) for the original image x s to perform the initialization operation; where h, w, and c are respectively the height, width, and channel dimensions of the images in the training sample set, and ε represents the maximum pixel value of the added perturbation. The original image x s is shown as an example in Figure 3 (a).

[0046] In a specific embodiment, in step S1, use a feature extraction algorithm to extract the non-robust feature samples of the sub-models on the non-robust feature image z. The effect is shown as an example in Figure 3 as follows, including the following steps:

[0047] S101: Randomly select another target image x; the example of the target image x is shown in Figure 3 (b).

[0048] S102: In an iterative manner, approximate the output value of the non-robust feature image z at the feature layer to the output value of the target image x at the feature layer to form the final non-robust feature sample. Its calculation formula is:

[0049]

[0050] In the formula: f i l (·) represents the output value of the l-th layer of the i-th sub-model, where the content in the parentheses represents the input of the model; z i,l represents the non-robust feature sample generated by the i-th sub-model through the l-th feature layer; s.t.||.|| ∞ represents using the infinity norm to constrain the generated non-robust feature sample.

[0051] Minimizing equation (1) can make the feature representation of the non-robust feature image close to the target image while keeping the non-robust feature image as similar as possible to the original image. The non-robust feature sample is essentially an adversarial sample generated by the i-th sub-model, which has the "vulnerability" information of this sub-model, that is, it contains the non-robust features of this sub-model. The result of the non-robust feature image z is shown as an example in Figure 3 (c).

[0052] S2: Select an untrained sub-model, and input the non-robust feature samples extracted by other sub-models into this sub-model for training respectively. The sub-models generate non-robust feature samples for each other through mutual training to learn each other's "vulnerability", effectively reducing the transferability between sub-models. By mutually training the non-robust feature samples between different sub-models to learn each other's "vulnerability", the transferability between sub-models is effectively reduced.

[0053] S3: As Figure 5 shown, when training non-robust feature samples, by combining the feature layer mixing method, the output values of the non-robust feature samples at the t-th intermediate feature layer of the sub-model being trained are mixed in different proportions into an intermediate layer output feature_map.

[0054] In a specific embodiment, specifically, when training the i-th sub-model, randomly select the non-robust feature images generated by the j-th sub-model, and mix them with the non-robust feature samples generated by other sub-models except the i-th and j-th sub-models in proportion λ and γ at the output of the t-th intermediate feature layer of the i-th sub-model to form an intermediate layer output feature_map.

[0055] This embodiment can reduce the training data or feature similarity between sub-models through random mixing of feature outputs, thereby further reducing the transferability between sub-models and further enhancing the difference between sub-models.

[0056] In a specific embodiment, the formula of the feature layer mixing method is as follows:

[0057]

[0058] In the formula, λ and γ are mixing coefficients. λ is a random matrix coefficient obeying the Beta distribution, and γ is defined as (1 - λ) / (N - 2), where N is the number of sub-models; t and l are randomly selected feature layers; k is the serial number of any sub-model except i and j; f i t (z j,l ) is the output value of the non-robust feature sample extracted by the j-th sub-model at the t-th layer of the i-th sub-model; Y i t is the mixed result feature_map of the output values of the non-robust feature samples generated by other sub-models at the t-th intermediate feature layer of this sub-model when training the i-th sub-model.

[0059] A more intuitive explanation is that in each iteration of training, this embodiment randomly selects a non-robust feature image generated by a sub-model as the main training sample, and the non-robust feature samples generated by the remaining sub-models are feature-mixed with the main training sample at the feature layer with different weights to obtain a mixed feature_map. This can reduce the similarity of training features between sub-models while enabling each sub-model to still learn the non-robust features of all other sub-models.

[0060] S4: Input the mixed feature_map into the sub-model being trained for forward propagation, calculate the cross-entropy to update the parameters of this sub-model; as Figure 6 shown, it is an example diagram of the integrated model update process.

[0061] Calculate the cross-entropy of the mixed feature_map, and the cross-entropy calculation formula is:

[0062]

[0063] In the formula: M is the number of categories; y s is the true category label of the original image; is the sign function, which takes 1 if the true category label is equal to c, otherwise takes 0; p c is then the probability of being judged as category c.

[0064] Equation (3) indicates that the cross-entropy formula can measure the inconsistency between the prediction result and the true result. The larger the entropy value, the less accurate the prediction, and the smaller the entropy value, the more accurate the prediction.

[0065] Furthermore, use the cross-entropy to update the parameters of the sub-model, and its formula is:

[0066]

[0067] Minimizing Equation (4) enables sub-model i to learn the "vulnerabilities" of other sub-models by learning the non-robust feature samples generated by other sub-models. That is, adversarial samples can successfully attack other sub-models but cannot easily succeed in attacking sub-model i. Combining with the feature mixing algorithm can better learn the non-robust features of other sub-models.

[0068] S5: Train all sub-models in the integrated model through the above steps S1 - S5 respectively until all sub-models reach the maximum number of training epochs, then the final sub-models are obtained.

[0069] Furthermore, in the actual test deployment, for each test sample, all the sub-models obtained by training through the above steps S1 to S5 are input simultaneously for query; the prediction results of all the sub-models are obtained through the query, the mean value of the prediction results is calculated, and this mean value is used as the final prediction result. The formula is as follows:

[0070]

[0071] where P i represents the prediction probability result of the i-th sub-model, and P ens is the final prediction result combined from all the sub-models.

[0072] The defense effect of the method described in this embodiment is as shown in Figure 7 and Figure 8 . Figure 7 shows the success rate of the method described in this embodiment in defending against black-box transfer attacks, and Figure 8 shows the success rate of the method described in this embodiment in defending against white-box attacks. The data in the first row represents the intensity of the attack perturbation, and the numbers in the first column represent the number of sub-models used. It can be seen that the method described in this embodiment can already achieve good defense against black-box transfer attacks while maintaining a high clean sample accuracy. And when defending against white-box attacks, it also has a good defense effect against low-perturbation attacks. In addition, the defense effect of the method described in this embodiment can be further enhanced as the number of sub-models increases.

[0073] Embodiment 2

[0074] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method steps implemented are as follows:

[0075] S1: Extract non-robust feature samples of all sub-models on each training sample;

[0076] S2: Select an untrained sub-model, and input the non-robust feature samples extracted by other sub-models into this sub-model for training respectively. The sub-models generate non-robust feature samples for each other through mutual training to learn each other's "vulnerabilities";

[0077] S3: Through a combined feature layer mixing method, the output values of the non-robust feature samples at the t-th layer intermediate feature layer of the sub-model being trained are mixed in different proportions into an intermediate layer output feature_map;

[0078] S4: Input the mixed feature_map into the sub-model being trained for forward propagation, and calculate the cross-entropy to update the parameters of this sub-model;

[0079] S5: Train all sub - models in the integrated model respectively through the above steps S1 - S5 until all sub - models reach the maximum number of training epochs, then the final sub - models are obtained.

[0080] Embodiment 3

[0081] A computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method steps implemented are as follows:

[0082] S1: Extract non - robust feature samples of all sub - models on each training sample;

[0083] S2: Select an untrained sub - model, and input the non - robust feature samples extracted by other sub - models into this sub - model for training respectively. The sub - models learn each other's "vulnerability" through the non - robust feature samples generated by mutual training.

[0084] S3: By combining the feature - layer mixing method, mix the output values of the non - robust feature samples at the t - th intermediate feature layer of the sub - model being trained in different proportions into an intermediate - layer output feature_map.

[0085] S4: Continue to input the mixed feature_map into the sub - model being trained for forward propagation, and calculate the cross - entropy to update the parameters of this sub - model.

[0086] S5: Train all sub - models in the integrated model respectively through the above steps S1 - S5 until all sub - models reach the maximum number of training epochs, then the final sub - models are obtained.

[0087] Obviously, the above - mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for enhancing the robustness of an integrated model, characterized in that: The method includes the following steps: S1: Extract non-robust feature samples of all sub-models on each non-robust feature image training sample; S2: Select an untrained sub-model, and input the non-robust feature samples extracted by other sub-models into this sub-model for training respectively. The sub-models generate non-robust feature samples for each other through mutual training to learn each other's "vulnerabilities"; S3: By combining the feature layer mixing method, the output values of the non-robust feature samples in the t -th intermediate feature layer of the sub-model being trained are mixed in different proportions into an intermediate layer output feature_map; Specifically, when training the i th sub-model, randomly select the non-robust feature images generated by the th sub-model, and the non-robust feature samples generated by other sub-models except the th sub-model and the th sub-model, respectively, according to the ratios and Mix them at the output of the th layer of the intermediate feature layer of the th sub-model to form an intermediate layer output feature_map; The feature layer mixing method, its calculation formula is: (1) In the formula, and are mixing coefficients, is a random matrix coefficient that follows a Beta distribution, is defined as , where N is the number of sub-models; t and l are both randomly selected feature layers; is any sub-model serial number except and ; is the output value of the non-robust feature samples extracted by the th sub-model at the th layer of the t th sub-model; is the mixing result feature_map of the output values of the non-robust feature samples generated by other sub-models at the intermediate feature layer of the i th layer of this sub-model when training the t th sub-model; S4: Continue to input the mixed feature_map into the sub-model being trained for forward propagation, and calculate the cross-entropy to update the parameters of this sub-model; S5: Train all sub-models in the integrated model respectively through the above steps S1 - S5 until all sub-models reach the maximum number of training epochs, then the final sub-models are obtained.

2. The method for enhancing the robustness of an integrated model according to claim 1, characterized in that: Step S1. Before extracting the non-robust feature samples of all sub-models, perform an initialization operation as follows: Based on the uniform distribution generate a noise matrix with a dimension of to initialize the original image by this operation; Among them h , w , c are the height, width, and channel dimensions of the training sample set images respectively, represents the maximum pixel value of the added perturbation.

3. The method for enhancing the robustness of an integrated model according to claim 2, characterized in that: Step S1, using a feature extraction algorithm to extract non-robust feature samples of the sub-model from the non-robust feature image z which includes the following steps: S101: Randomly select another target image x ; S102: In an iterative manner, approximate the output value of the non-robust feature image z at the feature layer to the target image x at the feature layer output value to form the final non-robust feature sample, and its calculation formula is: (2) In the formula: represents the output value of the i -th layer of the l -th sub-model, where the content in parentheses represents the input of the model; represents the non-robust feature sample generated by the i -th sub-model through the l feature layer; represents using the infinity norm to constrain the generated non-robust feature samples.

4. The method for enhancing the robustness of an integrated model according to claim 3, characterized in that: Calculate the cross-entropy of the mixed feature_map, and the cross-entropy calculation formula is: (3) Wherein: M is the number of categories; is the true category label of the original image; is the sign function, if the true category label is equal to c then take 1, otherwise take 0; then it is the probability of being judged as category c ​ 5. The method for enhancing the robustness of an integrated model according to claim 4, characterized in that: Use the cross-entropy to update the parameters of the sub-model, and its formula is: (4) Minimizing equation (4) enables the sub-model i to learn the "vulnerabilities" of other sub-models by learning the non-robust feature samples generated by other sub-models.

6. The method for enhancing the robustness of an integrated model according to any one of claims 1 - 5, characterized in that: In actual test deployment, for each test sample, input it into all sub-models trained through the above steps S1 - S5 for query at the same time; obtain the prediction results of all sub-models through the query, calculate the mean of the prediction results, and use this mean as the final prediction result.

7. A computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 - 5.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 - 5.

Citation Information

Patent Citations

  • Deep learning model-based adversarial training method

    CN112016686A

  • Deep learning model defense method aiming at adversarial attack and deep learning model

    CN113127857A