Adversarial Sample Generation Method for Destroying Middle Layer Features

By perturbing the intermediate layer features multiple times in the feature space and introducing momentum, the problem of inefficient adversarial sample generation in the prior art is solved, and more transferable adversarial samples are generated, which can effectively mislead the classification results of different models.

CN120124709BActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510619864.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-11
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In the prior art, the inefficiency of generating adversarial samples based on migration is inefficient, resulting in insufficient migration of adversarial samples and the inability to effectively mislead the classification results of different models.

Method used

By perturbing the intermediate layer features multiple times in the feature space along the direction of feature importance, and introducing momentum to stabilize the update direction, generating adversarial samples, and finally inverting the perturbed features into the pixel space to form an adversarial sample with stronger migration.

Benefits of technology

The generation efficiency of adversarial samples is improved, so that the generated adversarial samples can be more efficiently migrated between different models, misleading the output of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124709B_ABST
    Figure CN120124709B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of information security, and provides an adversarial sample generation method for destroying intermediate layer features. It is verified that the existing feature-level attacks essentially perturb the primary features once along the direction of feature importance in the feature space, resulting in limited transferability. Then, a pixel space to feature space attack method (P2FA) is proposed. By directly transferring the perturbation space from the pixel space to the feature space, the efficiency of destroying important features is improved, that is, perturbing the features multiple times along the direction of feature importance in the feature space, and then inverting the perturbed features to the image to generate adversarial samples with stronger transferability. The present invention can generate more transferable adversarial samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information security, and particularly relates to a method for generating adversarial samples that disrupt intermediate layer features. Background Art

[0002] Deep neural networks (DNNs) have achieved great success in various machine learning tasks. However, a large amount of work has shown that DNNs are vulnerable to adversarial samples, which add carefully designed and imperceptible perturbations to clean images to mislead DNNs. The existence of adversarial samples has raised concerns about security-sensitive applications (such as autonomous driving and face recognition). The study of adversarial samples can not only help researchers understand the principles of adversarial samples and the deficiencies of DNNs, but also improve the defense capabilities of DNNs and the robustness of images accordingly, enabling them to remain stable and accurate in various application scenarios.

[0003] Many methods have been proposed to generate adversarial samples. According to the information the attacker has about the target model, adversarial attacks can generally be divided into two categories: white-box attacks and black-box attacks. White-box attacks mean that the attacker can obtain all the knowledge of the target model (such as structure and parameters) to generate adversarial samples through gradient information. In contrast, in black-box attacks, the attacker cannot access the structure and parameters of the model, and thus cannot obtain gradient information, which makes black-box attacks more challenging and more in line with the actual situation.

[0004] Black-box attacks can be further divided into: transferability-based attacks and query-based attacks according to different attack strategies. Query-based attacks approximate gradient information through queries or use intelligent search algorithms to explore in the input space to generate adversarial samples. However, query-based attacks are infeasible in many practical scenarios, such as face recognition and autonomous driving, because a large number of queries are not allowed in these scenarios. In contrast, transferability-based black-box attacks are more realistic and flexible because they do not require any knowledge of the target model. It first attacks a local white-box proxy model and then directly transfers the obtained adversarial samples to an unknown target model. This cross-model attack ability of adversarial samples is called transferability.

[0005] Some studies have attempted to perform attacks at the intermediate layer to enhance the transferability of adversarial samples. These feature-level attacks do not directly interfere with the output layer of the proxy model, but rather achieve higher transferability by maximizing internal feature distortion. Since the most critical features are shared among different DNN models, feature-level attacks show promise in generating stronger transferable adversarial samples. Summary of the Invention

[0006] Objective of the Invention: The objective of the present invention is to address the deficiencies in the prior art and provide a method for generating adversarial examples that disrupt intermediate layer features. The method perturbs features multiple times along the direction of feature importance in the feature space and then inversely maps the perturbed features onto the image to generate more transferable adversarial examples. The present invention can improve the efficiency of feature perturbation and generate more transferable adversarial examples.

[0007] Technical Solution: A method for generating adversarial examples that disrupt intermediate layer features according to the present invention sets the classification model as , and for this classification model, the adversarial example generation method P2FA is used to perturb the intermediate layer features of the classification model along the direction with a step size of to obtain the perturbed intermediate layer features ;

[0008] Among them, and respectively represent the input original image and the corresponding true label; represents the feature map of the -th intermediate layer;

[0009] When performing the above perturbation T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into local optima. The specific process is as follows:

[0010] ;

[0011] ;

[0012] Among them, is the decay factor in the momentum, is the perturbation step size, is the update direction of the intermediate layer features at the -th iteration, and || ||2 represents the L2 norm; is the feature importance, and the iterative formula of is as follows:

[0013] ;

[0014] ;

[0015] n is the aggregation times for obtaining the feature importance, and the value range is 1 - N, represents the fitted image at the n-th iteration, represents the perturbation magnitude, represents the cross-entropy loss;

[0016] Finally, the perturbed features Invert to the pixel space to obtain the corresponding adversarial example , that is, summarize the feature inversion as the following optimization problem:

[0017] ;

[0018] wherein, represents the square of the L2 norm.

[0019] Furthermore, before perturbing the intermediate layer features of the classification model T times, first initialize the adversarial example when not iterated and the intermediate layer feature update direction , the formulas are as follows:

[0020] ; ; ;

[0021] Then, during the above T times of perturbation, the adversarial example and the intermediate layer feature update direction generated at the -th iteration are respectively denoted as , ,

[0022]

[0023]

[0024]

[0025] The value range of the above t is [0, T - 1]; furthermore, the adversarial example obtained by the T-th perturbation is obtained.

[0026] Beneficial effects: The present invention directly transfers the perturbation space from the pixel space to the feature space to improve the efficiency of destroying important features, and then obtains the corresponding adversarial example by means of feature inversion of the perturbed features. The present invention solves the problem of low efficiency of feature perturbation in traditional feature-level attacks, and enables the generated adversarial example to have stronger transferability. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic diagram of the overall framework of the present invention;

[0028] Figure 2 is the original image in the embodiment;

[0029] Figure 3 is the adversarial example image generated by adopting the technical solution of the present invention in the embodiment;

[0030] Figure 4 is a schematic diagram of the overall framework of the existing technical solution. Detailed implementation manners

[0031] The technical solutions of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0032] Verification conclusion: For the classification model , the existing feature-level attacks include the following optimization problems.

[0033] The optimization problem of FIA is:

[0034] (1)

[0035] Among them, represents the dot product operation, represents the inner product operation, represents the aggregated gradient;

[0036] (2)

[0037] Among them, represents the number of ensembles, represents satisfying the binary matrix of the distribution, represents the image obtained by performing the dot product operation on the original image and the matrix, , represents the logits output of the y-th dimension, represents the probability of performing random pixel dropout. Let , then the optimization problem of FIA can finally be rewritten as: (3)

[0038] The optimization problem of NAA is:

[0039] (4)

[0040] represents the value of the -th neuron in

[0041] (5)

[0042] Among them, represents the number of ensembles, represents the softmax output of the y-th dimension, , represents the all-black image, and are both linear functions that map to themselves, that is, , then formula (4) is rewritten as:

[0043] (6)

[0044] This is equivalent to: (7)

[0045] Let , then the optimization problem of NAA is finally rewritten as: (8)

[0046] The optimization problem of BFA is: (9)

[0047] Where , represents the fitted image, is set to 1, making the optimization problem become: (10)

[0048] Since is a constant term independent of the optimization parameters, the formula (10) is rewritten as:

[0049] (11)

[0050] Let , the optimization problem of BFA is finally rewritten as: (12)

[0051] In summary, the optimization problem of the existing feature-level attack is summarized as formula (12).

[0052] Here, formula (12) is transformed, and the transformation result is shown in formula (13):

[0053] (13)

[0054] Where represents the perturbation step size, which is a constant, so this transformation holds.

[0055] Then, formula (13) is transformed using the cosine theorem, and the result is shown in formula (14);

[0056] (14)

[0057] Because is a constant term, formula (14) is transformed into formula (15);

[0058] (15)

[0059] Specifically, maximizing means that the feature-level attack maximally destroys the intermediate-layer features of the surrogate model, while minimizing It means that feature-level attacks tend to maximize the destruction of the intermediate-layer features of the surrogate model along the direction.

[0060] In the above conclusion verification process, by analyzing the loss functions of existing feature-level attacks, it is mathematically proven that existing feature-level attacks actually expect to perturb features along the direction of feature importance in the feature space.

[0061] That is to say, the effect of perturbing multiple times in the pixel space by existing feature-level attacks (as Figure 4 shown) is the same as the effect of perturbing features only once along the direction of feature importance in the feature space. This inefficient perturbation in the pixel space limits their improvement of adversarial transferability. To solve the problem of low efficiency of existing feature perturbations, the present invention proposes an attack method P2FA from the pixel space to the feature space. Specifically, according to the feature importance, the perturbation space is directly transferred from the pixel space to the feature space to improve the efficiency of destroying important features, and then the perturbed features are used to obtain corresponding adversarial examples through the method of feature inversion.

[0062] As Figure 1 shown, for the adversarial example generation method for destroying intermediate-layer features of the present invention, the classification model is set as , and for this classification model, the adversarial example generation method P2FA is used to perturb the intermediate-layer features of the classification model along the direction with a step size of to form ; where and respectively represent the input original image and the corresponding true label; represents the feature map of the th intermediate layer;

[0063] When performing the above perturbation T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into local optima. The specific process is as follows:

[0064] ;

[0065] ;

[0066] where is the decay factor in the momentum, is the perturbation step size, is the feature importance;

[0067] The iterative formula of

[0068] is as follows:

[0069] ;

[0070] Finally, use the feature inversion algorithm to invert the perturbed feature into the pixel space to obtain the corresponding adversarial sample , that is, summarize the feature inversion as the following optimization problem:

[0071] ;

[0072] Invert the perturbed feature into the pixel space through the feature inversion algorithm to obtain the adversarial sample .

[0073] The present invention analyzes the deficiencies existing in the existing feature-level attack algorithms, focuses on solving the problem of low efficiency of perturbing features in the existing feature-level attacks, and finally generates more transferable adversarial samples.

[0074] Before performing T perturbations on the intermediate layer features of the classification model in this embodiment, first initialize the adversarial sample when not iterated and the intermediate layer feature update direction , and the formulas are as follows:

[0075] ; ; ;

[0076] Then, during the above T perturbations, the adversarial sample and the intermediate layer feature update direction generated at the th iteration are respectively denoted as , ,

[0077]

[0078]

[0079]

[0080] The value range of the above t is [0, T - 1]; and then the adversarial sample obtained after the Tth perturbation is .

[0081] The above adversarial sample generation method can be represented by the P2FA algorithm.

[0082]

[0083] To verify the technical effect of the present invention, this embodiment compares the attack success rates of the technical solution P2FA of the present invention and the existing feature-level attacks (FIA, NAA, BFA).

[0084] Table 1 Comparison of the attack success rates with existing feature-level attacks (FIA, NAA, BFA).

[0085]

[0086] Table 1 generates adversarial samples with Inception-v3 (Inc-v3), Inception-v4 (Inc-v4), Inception-ResNet-v2 (IncRes-v2), and ResNet-152 (Res-152) as source models to attack different target models respectively. Figure 2 is the original clean image input in this embodiment, Figure 3 is the corresponding adversarial sample generated using the P2FA technical solution of the present invention. It can be seen that the generated adversarial sample is almost visually indistinguishable from the original clean image. However, the generated adversarial sample can mislead the classification model.

Claims

1. An adversarial sample generation method for destroying the intermediate layer features, characterized in that, Set the classification model as , for this classification model, use the adversarial sample generation method P2FA to perturb the intermediate layer features of the classification model along the feature importance direction with a step size of to obtain the perturbed intermediate layer features ; Among them, and respectively represent the input original image and the corresponding ground truth label; represents the feature map of the th intermediate layer; When performing the above-mentioned perturbation T times, momentum is introduced to stabilize the update direction in the feature space and prevent the perturbed features from falling into local optima. The specific process is as follows: ; ; Among them, is the decay factor in momentum, is the perturbation step size, is the update direction of the intermediate layer features at the -th iteration, is the update direction of the intermediate layer features at the -th iteration, and || ||2 represents the L2 norm; is the feature importance, The iteration formula of ; ; n is the number of aggregations for obtaining feature importance, and its value range is 1 - N; represents the fitted image at the n-th iteration, represents the perturbation size, represents the cross-entropy loss; Finally, the perturbed features are inverted to the pixel space to obtain the corresponding adversarial sample , that is, the feature inversion is summarized as the following optimization problem: ; Among them, represents the square of the L2 norm.

2. The adversarial sample generation method for destroying the intermediate layer features according to claim 1, wherein, Before perturbing the intermediate layer features of the classification model T times, first initialize the adversarial sample when not iterating and the update direction of the intermediate layer features , the formula is as follows: ; ; ; Then, during the above T perturbation processes, the adversarial example and the intermediate layer feature update direction generated at the -th iteration are respectively denoted as and . ; ; ; The value range of the above-mentioned t is [0, T - 1]; thus, the adversarial example obtained by the T-th perturbation is obtained. .

Citation Information

Patent Citations

  • Multi-modal sentiment classification method based on isomorphism and heterogeneity dynamic information interaction

    CN116010595A

  • Generation method of mobility confrontation sample and black box attack method

    CN118052273A