Image recognition model backdoor attack method based on target label

Through the target label-based backdoor attack framework and shadow model to generate poisoning data, the problem of insufficient concealment and attack intensity in the image recognition model is solved, and efficient image recognition model attack is achieved, suitable for many-to-many attack scenarios.

CN120372582APending Publication Date: 2025-07-25TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410103442.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The backdoor attack methods of existing image recognition models are insufficient in terms of concealment and attack strength, especially in many-to-many attack scenarios, the attack success rate is low and the attack model needs to be accessed, which affects the actual threat.

Method used

A backdoor attack framework based on target label is adopted to generate highly concealed poisoning data by training the shadow model, and a backdoor trigger is generated using gradient descent optimization, and the characteristics of the human visual system are considered to enhance attack concealment and attack intensity.

Benefits of technology

In complex scenarios, a higher attack success rate is achieved. The attack is well concealed and does not affect the benign accuracy of the attacked model. The applicability is wide and does not require the specific information of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372582A_ABST
    Figure CN120372582A_ABST
Patent Text Reader

Abstract

A backdoor attack for deep learning is proved to form a serious threat to a real application scene. For an existing backdoor attack method in the field of image recognition, some defects still exist, and the existing backdoor attack method either needs to access an attacked model, so that the threatening of the attacked model can be obviously reduced in practice, or the existing backdoor attack method is poor in concealment. In addition, in a complex scene (many-to-many attack scene), the attack success rate of the existing method still needs to be improved. Therefore, the invention provides a backdoor attack framework based on the target label so as to improve the problem of the existing backdoor attack method. In order to solve the problem of weak attack concealment, the method provides an optimization algorithm by considering human eye visual perception characteristics to ensure the concealment of the backdoor trigger. Aiming at the problem of insufficient backdoor attack intensity, the invention innovatively provides a tag-related attack method, and in the method, backdoor samples are continuously optimized and generated aiming at a specific target tag, so that the backdoor attack capability is remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image attack, and more specifically, relates to a backdoor attack method for an image recognition model based on a target label. Background Art

[0002] In today's digital age, image recognition technology has become a core part of the field of artificial intelligence. The application scope of this technology covers various fields from autonomous driving to medical image diagnosis. With the wide application of image recognition, the demand for its security and stability has become increasingly important. In recent years, the security issues of image recognition models have attracted wide attention. In the backdoor attack on an image recognition model, in addition to directly modifying the model to embed a backdoor, a more convenient and general way is to embed a backdoor by poisoning (modifying) a part of the images in the training set. After the attacker sets the target label to be attacked and the backdoor trigger, a part of the data in the training set is added with the backdoor trigger and its label is modified to the target label, and then the poisoned training set is used to train the model to obtain a model with a backdoor. There are two ways to set the target label: (1) Many-to-one attack: that is, all samples containing the backdoor trigger are set to a specific label; (2) Many-to-many attack: that is, samples containing the backdoor trigger are set to different labels according to the different original sample labels. For the existing backdoor attack methods in the field of image recognition, although a high attack success rate can be achieved, there are still some deficiencies. They either need to access the attacked model, which may significantly reduce their threat in practice, or perform poorly in terms of concealment. In addition, in complex scenarios (such as the many-to-many attack scenario), the attack success rate of the existing methods remains to be improved.

[0003] Adversarial attack is an attack method against a deep learning model. This attack occurs in the model inference stage. By making tiny and targeted perturbations to the input samples, the model is made to produce incorrect outputs or misjudgments. This attack is based on the exploitation of model vulnerabilities. By adding carefully designed perturbations, the attacker attempts to deceive the model into producing incorrect prediction results. For different attack purposes, the attacker can conduct targeted or untargeted attacks. In the targeted attack scenario, the attack target is to make the model's prediction be the category predefined by the attacker. On the contrary, in the untargeted attack scenario, the attack target is to make the model produce incorrect prediction results. And adversarial samples are the poisoned samples produced by adversarial attacks. Summary of the Invention

[0004] In view of the defects and improvement requirements of existing backdoor attack technologies, the present invention proposes a backdoor attack framework based on target labels to address the issues of weak attack concealment and low attack intensity in existing methods. To address the problem of weak attack concealment, the present invention proposes an optimization algorithm by considering the characteristics of human visual perception to ensure the concealment of the backdoor trigger. To address the insufficient backdoor attack intensity, this paper innovatively proposes a label-related attack method. In this method, by designing an optimization method for specific target labels to generate backdoor samples, the ability of the backdoor attack is significantly enhanced.

[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0006] A method for backdoor attack on an image recognition model based on target labels, comprising the steps of:

[0007] Step 1. Train a shadow model: Since an attacker usually cannot easily obtain the internal information of the target model, and to make the attack more effective, we use a shadow model to replace the target model to generate poisoned data, that is, train the shadow model on a clean data set to make it fit the image data features.

[0008] Step 2. Generate poisoned data: After selecting the target label to be attacked, use the shadow model's understanding of the image features of this category to generate highly concealed poisoned data.

[0009] Step 3. Attack: Use the generated poisoned data to train the target model.

[0010] Step 4. Test: After the model is deployed to the production environment, the attacker uses the poisoned test data generated in Step 2 to attack the model.

[0011] Furthermore, the specific steps of Step 1 are as follows:

[0012] Step 1.1: Use the same clean data set as the target model to train a shadow model. That is, given the training data D train , we train a shadow model g by minimizing the following loss function θ : where g can be any CNN-based shadow model, represents the loss function, and here the cross-entropy loss function is used:

[0013] Furthermore, the specific steps of Step 2 are as follows:

[0014] Step 2.1: Select ρ% of the data in the training set as poisoned data and modify its original label y to the target label η(y) to be attacked.

[0015] Step 2.2: Through the shadow model, generate a highly stealthy backdoor trigger δ for the target label η(y) to be attacked. Its generation process can be represented by solving the following optimization problem through the gradient descent method:

[0016]

[0017] For the first term where is the standard cross-entropy loss function. It attempts to find a δ through gradient descent such that the poisoned data x + δ is consistent with the image features of the category where the target label η(y) is located in terms of image features.

[0018] The second term γ·||δ||2 will further limit the update range of δ and accelerate the optimization process. Specifically, using the ||δ||2 term will limit the update range of δ. In this way, the stealthiness of the generated poisoned data is further improved. In addition, the number of iterations for the noise to converge to an imperceptible level is also greatly reduced.

[0019] Considering that the human visual system is sensitive to different colors to different degrees, in order to ensure the stealthiness of the trigger. Therefore, the third term ||ΔE 00 (x + δ, x||2 is used to ensure that the backdoor trigger aligns with the features of the human visual system. Where ΔE 00 uses the latest standard formula of the International Commission on Illumination.

[0020] The beneficial effects of the present invention are as follows:

[0021] (1) Strong attack effect. The present invention achieves a higher attack success rate in complex attack scenarios (many-to-many attack scenarios).

[0022] (2) Good attack stealthiness. The poisoned images generated by the present invention are almost similar to the original images, and it is very difficult for the naked eye to detect the difference between the poisoned images and the original images.

[0023] (3) Almost does not affect the benign accuracy rate of the attacked model. It has little impact on the benign accuracy rate of the attacked model.

[0024] (4) Wide attack applicability. The present invention does not need to know the specific information of the attacked model, so it has wide attack applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic flowchart of the embodiment of the present invention.

[0026] Figure 2 It is a comparison diagram between the original image and the poisoned image.

[0027] Figure 3Show the attack effect of the present invention on different data sets. The method of the present invention is named Impart. Detailed implementation manners

[0028] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0029] Step 1: Train a shadow model: Since an attacker generally cannot obtain the internal information of the model to be attacked, and since most of the existing image recognition models are based on CNNs, there is a certain structural similarity between the models. Therefore, the present invention selects to use a shadow model to replace the model to be attacked to generate poisoned data, that is, the shadow model is trained on a clean data set to fit the image data features. The specific steps are as follows:

[0030] Step 1.1: Use the same clean data set as the model to be attacked to train a shadow model. That is, given the training data D train , we train a shadow model g by minimizing the following loss function θ : where g can be any CNN-based shadow model, represents the loss function, and here the cross-entropy loss function is used:

[0031] Step 2: Generate poisoned data: After selecting the target label to be attacked, use the shadow model to generate highly concealed poisoned data based on the understanding of the image features of the category where the target label is located. The specific steps are as follows:

[0032] Step 2.1: Select ρ% of the data in the training set as poisoned data, and modify its original label y to the target label η(y) to be attacked.

[0033] Step 2.2: Through the shadow model, generate a highly concealed backdoor trigger δ for the target label η(y) to be attacked. Its generation process can be represented by solving the following optimization problem through the gradient descent method:

[0034]

[0035] For the first term where is the standard cross-entropy loss function. It tries to find a δ through gradient descent such that the poisoned data x + δ is consistent with the image features of the category where the target label η(y) is located in terms of image features.

[0036] The second term, γ·||δ||2, will further restrict the update range of δ and accelerate the optimization process. Specifically, using the ||δ||2 term will limit the update range of δ. In this way, the stealthiness of the generated poisoning data is further improved. In addition, the number of iterations for the noise to converge to an imperceptible level is also greatly reduced.

[0037] Considering that the human visual system is more sensitive to different colors, in order to ensure the stealthiness of the trigger. Therefore, the third term ||ΔE 00 (x + δ, x||2 is used to ensure that the backdoor trigger aligns with the characteristics of the human visual system. Where ΔE 00 The latest standard formula of the International Commission on Illumination is used.

[0038] Such as Figure 2 As shown, in order to demonstrate the attack stealthiness of the present invention, a comparison chart between the original image and the poisoned image is listed. Among them, the first row is the original clean image, the second row is the poisoned image generated by the present invention, the third row is the difference image between the two, and the fourth row is the result of magnifying the difference value between the two by 10 times. It can be seen from this that the poisoned image generated by the present invention is almost identical to the original image, indicating that the attack stealthiness of the present invention is high.

[0039] Step 3. Attack: Use the generated poisoned data to train the attacked model.

[0040] Step 4. Test: After the model is deployed to the production environment, the attacker uses the poisoned test data generated in Step 2 to attack the model.

[0041] Such as Figure 3 As shown, in order to demonstrate the attack effect of the present invention, the present invention selects three different shadow models and verifies the attack effect on three datasets: In the multi-to-multi attack scenario, the attack success rate of the present invention on the CIFAR-10 and CIFAR-100 datasets is better than the baseline method, and the benign accuracy of the model is hardly reduced. At the same time, on the GTSRB dataset, the present invention achieves results comparable to the baseline method. In particular, in the GTSRB dataset, since the accuracy of PreActResNet18 on the GTSRB clean dataset is approximately 99%, the attack success rate of all previous works on this dataset can also reach approximately 99%. However, when the complexity of the dataset increases, the attack accuracy of the baseline method drops rapidly, while our method achieves an attack success rate of 92.94% on the CIFAR-100 dataset when using EfficientNetB0 as the shadow model, which is much higher than the baseline method.

[0042] The parts not detailed in the present invention belong to the well-known technologies in the art.

Claims

1. A backdoor attack method for an image recognition model based on target tags, characterized in that, It includes the following steps: Step 1, train a shadow model: Since the attacker cannot obtain the internal information of the model under attack, in order to make the attack more effective, a shadow model is used to replace the model under attack to generate poisoned data, that is, the shadow model is trained on a clean data set to fit the image data features; Step 2, generate poisoned data: After selecting the target label to be attacked, use the shadow model's understanding of the image features of this category to generate highly concealed poisoned data; Step 3, attack: Use the generated poisoned data to train the model under attack; Step 4, test: After the model is deployed to the production environment, the attacker uses the poisoned test data generated in Step 2 to attack the model.

2. The backdoor attack method for an image recognition model based on a target label according to claim 1, wherein Step 1, the specific steps are as follows: Step 1.1: Train a shadow model using the same clean dataset as the attacked model; that is, given training data D train , we train a shadow model by minimizing the following loss function where g can be any CNN-based shadow model, denotes the loss function, and here the cross-entropy loss function is used:

3. The method for backdoor attack on an image recognition model based on target tags according to claim 1, wherein Step 2, the specific steps are as follows: Step 2.1: Select ρ% of the data in the training set as poisoned data and modify its original label y to the target label η(y) to be attacked; Step 2.2: Through the shadow model, generate a highly concealed backdoor trigger δ for the target label η(y) to be attacked; Its generation process can be represented by solving the following optimization problem through the gradient descent method: For the first item where is the standard cross-entropy loss function; it attempts to find a δ through gradient descent such that the poisoned data x+δ is consistent with the image features of the class where the target label η(y) is located in terms of image features; The second term γ·||δ||2 will further limit the update range of δ and speed up the optimization process; specifically, using the ||δ||2 term will limit the update range of δ; in this way, the concealment of the generated poisoned data is further improved; in addition, the number of iterations for the noise to converge to an imperceptible level is also greatly reduced; Considering that the human visual system is more sensitive to different colors, in order to ensure the concealment of the trigger; therefore, the third item ||ΔE is used 00 (x + δ, x||2 to ensure that the backdoor trigger is aligned with the characteristics of the human visual system; where ΔE 00 Use the latest standard formula of the International Commission on Illumination