Sample generation method and device for transferable adversarial attack on autonomous driving

Through the processing and feature mixing technology of adversarial samples, adversarial samples with good migration capabilities are generated, which solves the problem of low migrationability of adversarial samples of existing autonomous driving models and improves the safety of the model.

CN118692055BActive Publication Date: 2025-05-20TIANJIN UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410805235.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-05-20
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

When facing adversarial attacks, existing autonomous driving models are limited by the low migration of adversarial samples, resulting in lower model security.

Method used

By using the interference block to process the patch location in the original sample image, the clean and adversarial features are mixed with the feature mixing block in the target recognition model, and the interference block is updated using the gradient-based optimization algorithm to generate adversarial samples with good migration capabilities.

Benefits of technology

The migrationability of adversarial samples is improved, thereby enhancing the security of the target recognition model, making the performance of the model more stable under different environments and models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692055B_ABST
    Figure CN118692055B_ABST
Patent Text Reader

Abstract

The present disclosure provides a sample generation method and device for transferable adversarial attack for autonomous driving, which can be applied to the fields of autonomous driving technology and intelligent recognition technology. The method includes: using interference blocks to process the patch positions in the original sample image to obtain an adversarial sample image, the interference blocks are images with noise, and the patch positions are obtained by inputting the original sample image into the trained self-attention mechanism model; using feature mixing blocks in the target recognition model to mix the clean features and adversarial features of the original sample image to obtain an adversarial sample image to be updated, the target recognition model is obtained by training the recognition model based on the loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; using a gradient-based optimization algorithm to update the interference blocks in the adversarial sample image to be updated to obtain the target adversarial sample image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of autonomous driving technology and intelligent recognition technology, and in particular, to a method and device for generating samples for transferable adversarial attacks in autonomous driving. Background Art

[0002] With the development of autonomous driving technology, artificial intelligence models play an important role in various computer vision tasks, such as image classification, object detection, and other tasks. In safety-sensitive scenarios such as autonomous driving, the model can identify traffic signs on the road conditions to provide accurate driving information.

[0003] In the case where the model is under adversarial attack, adversarial images (i.e., road condition images with imperceptible perturbations added) can deceive the model and lead the model to make wrong decisions. Due to the low transferability of adversarial samples during the model adversarial attack training process (i.e., the adversarial nature is different under different models or different road condition environments), the security of the model obtained by adversarial attack training is low. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a method and device for generating samples for transferable adversarial attacks in autonomous driving.

[0005] According to a first aspect of the present disclosure, there is provided a method for generating samples for transferable adversarial attacks in autonomous driving, including: processing patch positions in an original sample image using an interference block to obtain an adversarial sample image, where the interference block is an image with noise, and the patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model; mixing the clean features and adversarial features of the original sample image using a feature mixing block in a target recognition model to obtain a to-be-updated adversarial sample image, where the clean features are extracted from the original sample image, the adversarial features are extracted from the adversarial sample image, the target recognition model is obtained by training a recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; and updating the interference block in the to-be-updated adversarial sample image using a gradient-based optimization algorithm to obtain a target adversarial sample image.

[0006] According to an embodiment of the present disclosure, mixing the clean features and adversarial features of the original sample image using a feature mixing block in a target recognition model to obtain a to-be-updated adversarial sample image includes: when the feature mixing block in the target recognition model is in a working state, mixing the clean features and adversarial features of the original sample image to obtain a to-be-updated adversarial sample image, where the target recognition model includes a plurality of feature mixing blocks, and the plurality of feature mixing blocks are respectively located in different structural layers of the target recognition model.

[0007] According to an embodiment of the present disclosure, updating the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image includes: when the feature mixing block in the target recognition model is in a working state, updating the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image.

[0008] According to an embodiment of the present disclosure, using the interference block to process the patch position in the original sample image to obtain an adversarial sample image further includes: mapping the patch position to the original sample image to obtain a target mask position; performing mask processing on the target mask position in the original sample image to obtain a target mask image; superimposing the target mask image and the interference block to obtain an adversarial sample image.

[0009] According to an embodiment of the present disclosure, the above method further includes: training a class attention map corresponding to the original sample image by using a self-attention mechanism; determining a patch position that satisfies a preset threshold in the class attention map.

[0010] According to an embodiment of the present disclosure, the above method further includes: training a recognition model based on a loss function value to obtain model parameters to be adjusted of the trained recognition model; fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters; obtaining a trained target recognition model while keeping the fine-tuned model parameters unchanged.

[0011] According to an embodiment of the present disclosure, fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters includes: when multiple feature mixing blocks in the target recognition model are not in a working state, fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters.

[0012] According to an embodiment of the present disclosure, fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters further includes: redefining the loss function as:

[0013]

[0014] L f is the redefined loss function, represents the loss function value of the adversarial sample image J is the loss function before redefinition, y is the label, is the model parameter in the t-th iteration process, represents the loss function value of the original sample image x.

[0015] According to an embodiment of the present disclosure, the mixing formula of the feature mixing block is as follows:

[0016]

[0017] f t ’ is the mixed feature corresponding to the adversarial sample image to be updated during the t-th iteration is the clean feature during the t-th iteration, f t is the adversarial feature, and r is the mixing ratio

[0018] The second aspect of the present disclosure provides a sample generation device for transferable adversarial attacks for autonomous driving, including: a processing module, configured to process the patch positions in the original sample image by using an interference block to obtain an adversarial sample image, where the patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model; a mixing processing module, configured to mix the clean feature and the adversarial feature of the original sample image by using a feature mixing block in the target recognition model to obtain an adversarial sample image to be updated, where the clean feature is extracted from the original sample image, the adversarial feature is extracted from the adversarial sample image, the target recognition model is obtained by training a recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; an updating module, configured to update the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image

[0019] The third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above method

[0020] The fourth aspect of the present disclosure further provides a computer-readable storage medium, on which an executable computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented

[0021] The fifth aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented

[0022] According to the sample generation method and device for transferable adversarial attacks for autonomous driving provided by the present disclosure, by processing the patch positions in the original sample image by using an interference block to obtain an adversarial sample image, where the patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model, since different recognition models have the same attention mechanism for recognizing the original sample image, the adversarial sample image obtained by processing the patch positions by using the interference block has transferability

[0023] Since the target recognition model is trained based on the loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image using the loss function of the recognition model, the feature mixing block in the target recognition model is used to mix the clean features and adversarial features of the original sample image to obtain a more diverse adversarial sample image to be updated.

[0024] Then, the interference block in the adversarial sample image to be updated is updated using a gradient-based optimization algorithm to obtain the target adversarial sample image. Therefore, the target adversarial sample image has good transferability, thereby improving the security of the target recognition model. Brief Description of the Drawings

[0025] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0026] Figure 1 Schematically shows an application scenario diagram of a method for generating samples for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure;

[0027] Figure 2 Schematically shows a flowchart of a method for generating samples for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure;

[0028] Figure 3 Schematically shows a schematic diagram of processing the patch position in the original sample image using an interference block to obtain an adversarial sample image according to an embodiment of the present disclosure;

[0029] Figure 4 Schematically shows a schematic diagram of the process of generating samples for transferable adversarial attacks according to an embodiment of the present disclosure;

[0030] Figure 5 Schematically shows a structural block diagram of a device for generating samples for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure; and

[0031] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing a method for generating samples for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure. Detailed Embodiments

[0032] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0033] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0034] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0035] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0036] In the technical solutions of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0037] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide corresponding operation entrances for users to choose to agree or reject the results of automated decision-making; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision-making" here refers to the activity of making a decision by a person who specializes in a certain field, has specialized experience, knowledge, and skills, and has reached a certain professional level.

[0038] Embodiments of the present disclosure provide a method for generating samples for transferable adversarial attacks in autonomous driving, including: processing the patch positions in the original sample image using an interference block to obtain an adversarial sample image, where the interference block is an image with noise, and the patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model; mixing the clean features and adversarial features of the original sample image using a feature mixing block in the target recognition model to obtain a to-be-updated adversarial sample image, where the clean features are extracted from the original sample image, the adversarial features are extracted from the adversarial sample image, the target recognition model is obtained by training a recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; updating the interference block in the to-be-updated adversarial sample image using a gradient-based optimization algorithm to obtain a target adversarial sample image.

[0039] Figure 1 FIG. schematically shows an application scenario diagram of the method for generating samples for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure.

[0040] As Figure 1 shown, the application scenario 100 according to this embodiment may include a vehicle 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the vehicle 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0041] The user can use the vehicle 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the vehicle 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0042] The vehicle 101 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0043] Server 103 may be a server that provides various services, such as a background management server (for example only) that supports the websites browsed by the user using the vehicle 101. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests, etc.) to the terminal device.

[0044] It should be noted that the sample generation method for transferable adversarial attacks in autonomous driving provided in the embodiments of the present disclosure can generally be executed by the server 103. Correspondingly, the sample generation device for transferable adversarial attacks in autonomous driving provided in the embodiments of the present disclosure can generally be set in the server 103. The sample generation for transferable adversarial attacks in autonomous driving provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the vehicle 101 and / or the server 103. Correspondingly, the sample generation for transferable adversarial attacks in autonomous driving provided in the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 103 and capable of communicating with the vehicle 101 and / or the server 103.

[0045] It should be understood that Figure 1 the numbers of vehicles, networks, and servers in are merely illustrative. According to the implementation requirements, any number of vehicles, networks, and servers can be provided.

[0046] Figure 2 Schematically shows a flowchart of a sample generation method for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure.

[0047] As Figure 2 shown, the sample generation method for transferable adversarial attacks in autonomous driving in this embodiment includes operation S210 to operation S230.

[0048] In operation S210, the patch position in the original sample image is processed using an interference block to obtain an adversarial sample image.

[0049] According to an embodiment of the present disclosure, the interference block is an image with noise, and the patch position is obtained by inputting the original sample image into a trained self-attention mechanism model.

[0050] According to an embodiment of the present disclosure, the original sample image may be a clean image without noise. For example, the original sample image is a road condition image captured during autonomous driving.

[0051] According to an embodiment of the present disclosure, the attention mechanism model can be an object recognition model with an attention mechanism. For example, the object recognition model can be a hybrid model combining a neural network model and an attention mechanism. The original sample image is trained using the object recognition model with an attention mechanism. During the training process, the original sample image will learn to focus attention on different regions of the image to perform specific tasks such as classification, detection, or segmentation. After training is completed, the sensitive region of the original sample image, i.e., the patch location, is found by observing the attention weights of the object recognition model on the original sample image. For example, the patch location of the original sample image can be the location with a traffic sign, and the object recognition model is used to recognize the traffic sign in the road condition map for autonomous driving.

[0052] According to an embodiment of the present disclosure, the adversarial sample image is an image with a tiny perturbation to deceive the object recognition model and make the object recognition model make a wrong classification.

[0053] In operation S220, the clean features and adversarial features of the original sample image are mixed using the feature mixing block in the object recognition model to obtain the adversarial sample image to be updated.

[0054] According to an embodiment of the present disclosure, the clean features are extracted from the original sample image, and the adversarial features are extracted from the adversarial sample image.

[0055] According to an embodiment of the present disclosure, the feature mixing block can perform an interaction operation on the clean features and adversarial features of different channels in the object recognition model, which helps to reduce information loss and enhance the correlation between features.

[0056] According to an embodiment of the present disclosure, the object recognition model is obtained by training the recognition model based on the loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model.

[0057] According to an embodiment of the present disclosure, the loss function can be a function for measuring the difference between the recognition result of the object recognition model and the actual label. The role of the loss function value is to guide the object recognition model to adjust parameters during the training process to minimize the error between the recognition result and the true value. The loss function can be cross-entropy, mean square error, contrast loss, etc.

[0058] In operation S230, the interference block in the adversarial sample image to be updated is updated using a gradient-based optimization algorithm to obtain the target adversarial sample image.

[0059] According to an embodiment of the present disclosure, the gradient-based optimization algorithm can be gradient descent, stochastic gradient descent, mini-batch gradient descent, etc. The gradient-based optimization algorithm is used to guide the direction and step size of parameter update in the interference block.

[0060] According to an embodiment of the present disclosure, by using an interference block to process the patch positions in the original sample image, an adversarial sample image is obtained. The patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model. Since different recognition models have the same attention mechanism when recognizing the original sample image, the adversarial sample image obtained by processing the patch positions with the interference block has transferability.

[0061] Since the target recognition model is trained based on the loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model, the clean features and adversarial features of the original sample image are mixed by using the feature mixing block in the target recognition model to obtain a more diverse adversarial sample image to be updated.

[0062] Then, the interference block in the adversarial sample image to be updated is updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image. Therefore, the target adversarial sample image has good transfer ability, thereby improving the security of the target recognition model.

[0063] According to an embodiment of the present disclosure, the above method further includes: training a category attention map corresponding to the original sample image by using a self-attention mechanism; determining patch positions that meet a preset threshold in the category attention map.

[0064] According to an embodiment of the present disclosure, first, a mask M is used to set the generation method of the adversarial sample image. The generation of the adversarial sample image can be expressed as:

[0065] (1)

[0066] is the generated adversarial sample, M is a 0-1 mask used to determine the patch positions where the interference block is generated. x is the original sample image, and δ is the added interference block.

[0067] According to an embodiment of the present disclosure, a method for improving transfer ability is set by using an attention mechanism. The transferable attack ability of adversarial perturbations is improved by finding the best patch positions on the original image. Since different recognition models have similar attention mechanisms when predicting the same target image, adding perturbations to the most sensitive regions (patch positions) can generate adversarial samples with high transfer ability, which can deceive unknown black-box models.

[0068] According to an embodiment of the present disclosure, an attention map is calculated by using a self-attention mechanism. Given an original sample image (clean sample image) x and a label y corresponding to the original sample image, the category attention map S is then calculated by using the self-attention mechanismy :

[0069] (2)

[0070] where A is specifically:

[0071] (3)

[0072] is the pixel value at position (i, j) in the k-th feature map, p y is the score for class y, and relu() represents the relu function.

[0073] According to an embodiment of the present disclosure, a mask M is defined using the generated attention map. The class attention map S is calculated y , and the index satisfying the preset threshold is found in the flattened attention map. The preset threshold can be the maximum activation value. Once the index of the maximum activation value is found, this position is mapped back to the original image, thereby defining the mask M:

[0074] (4)

[0075] , , p S is the preset patch size. i is the abscissa of the pixel, j is the ordinate of the pixel, H is the initial height, W is the initial width, H s is the height of the mask, and W s is the width of the mask. s i is the abscissa of the pixel of the patch, and s j is the ordinate of the pixel of the patch.

[0076] According to an embodiment of the present disclosure, using the interference block to process the patch position in the original sample image to obtain the adversarial sample image further includes: mapping the patch position to the original sample image to obtain the target mask position; performing mask processing on the target mask position in the original sample image to obtain the target mask image; and superimposing the target mask image and the interference block to obtain the adversarial sample image.

[0077] Figure 3 Schematically shows a schematic diagram of using the interference block to process the patch position in the original sample image to obtain the adversarial sample image according to an embodiment of the present disclosure.

[0078] As Figure 3As shown, first, the transfer attack ability of adversarial perturbations is improved by finding the optimal target mask position on the original sample image. Specifically, the self-attention mechanism is used to calculate the class attention map of the original sample image, and the target mask position is determined through the generated class attention map. The target mask position in the original sample image is masked to obtain the target mask image. Then, the adversarial sample is calculated by masking the original image and the perturbation, that is, the target mask image and the interference block are superimposed to obtain the adversarial sample image.

[0079] According to an embodiment of the present disclosure, the target recognition model is fine-tuned to encourage the perturbation to concentrate on robust features. After finding the patch position for adding the perturbation, it is necessary to generate a robust adversarial perturbation, and adversarial fine-tuning is used to encourage the perturbation to concentrate on robust features.

[0080] According to an embodiment of the present disclosure, using a gradient-based optimization algorithm to fine-tune the model parameters to be adjusted to obtain the fine-tuned model parameters further includes: redefining the loss function as:

[0081] (5)

[0082] L f is the redefined loss function, represents the loss function value of the adversarial sample image , J is the loss function before redefinition, y is the label, and θ t is the model parameter in the t-th iteration process, represents the loss function value of the original sample image x.

[0083] According to an embodiment of the present disclosure, the loss formula is redefined using the adversarial sample and the original sample image.

[0084] After fine-tuning, the accuracy of the current model in recognizing clean samples will decrease, and the model is prone to overfitting to adversarial samples. In order to enable the model to recognize both adversarial samples and clean samples, the loss function is redefined.

[0085] According to an embodiment of the present disclosure, the above method further includes: training the recognition model based on the loss function value to obtain the model parameters to be adjusted of the trained recognition model; using a gradient-based optimization algorithm to fine-tune the model parameters to be adjusted to obtain the fine-tuned model parameters; obtaining the trained target recognition model while keeping the fine-tuned model parameters unchanged.

[0086] According to an embodiment of the present disclosure, the SGD (Stochastic Gradient Descent) optimizer is used to fine-tune the model parameters to be adjusted to obtain the fine-tuned model parameters. The fine-tuned model parameters can be expressed as:

[0087] (6)

[0088] θ t may be the model parameters during the t-th iteration may be the fine-tuned model parameters during the (t + 1)-th iteration. L f is the redefined loss function

[0089] According to an embodiment of the present disclosure, fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain the fine-tuned model parameters includes: when multiple feature mixing blocks in the target recognition model are not in the working state, fine-tuning the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain the fine-tuned model parameters

[0090] According to an embodiment of the present disclosure, the clean features are stored by the target recognition model for the original sample image, and all feature mixing blocks are deactivated during the fine-tuning process

[0091] According to an embodiment of the present disclosure, mixing the clean features and adversarial features of the original sample image by using the feature mixing blocks in the target recognition model to obtain the adversarial sample image to be updated includes: when the feature mixing blocks in the target recognition model are in the working state, mixing the clean features and adversarial features of the original sample image to obtain the adversarial sample image to be updated. The target recognition model includes multiple feature mixing blocks, and the multiple feature mixing blocks are respectively located in different structural layers of the target recognition model

[0092] According to an embodiment of the present disclosure, the mixing formula of the feature mixing block is as follows

[0093] (7)

[0094] f t ’ is the mixed feature corresponding to the adversarial sample image to be updated during the t-th iteration, f c is the clean feature during the t-th iteration, f t is the adversarial feature, and r is the mixing ratio

[0095] According to an embodiment of the present disclosure, a feature mixing block can be added to the middle layer of the target recognition model for mixing clean features and adversarial features

[0096] To further optimize the generation of adversarial samples, a feature mixing block is inserted into the intermediate layer of the target recognition model, and then clean features are extracted from the original sample image. Each clean feature is stored in its memory during the first inference of the feature mixing block. In each iteration, the feature mixing block is activated with probability p, that is, the feature mixing block is in a working state. When the feature mixing block is activated, the clean feature and the adversarial feature are mixed by linear interpolation with the mixing ratio r. That is, at the t-th iteration, the stored clean feature and the adversarial feature .

[0097] According to an embodiment of the present disclosure, updating the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image includes: when the feature mixing block in the target recognition model is in a working state, updating the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image.

[0098] According to an embodiment of the present disclosure, the feature mixing block is activated to obtain an updated interference block. When the fine-tuning is completed, we reactivate the feature mixing block with probability p, and the gradient-based optimization algorithm can be the Adam (adaptive momentestimation) algorithm. Then, based on the Adam algorithm, the interference block in the adversarial sample image to be updated is updated to obtain a target adversarial sample image.

[0099] (8)

[0100] δ t is the interference block in the t-th iteration process, and δ t+1 is the updated interference block, is the adversarial sample image in the (t + 1)-th iteration process. It should be noted that the relationship between the number of iterations for updating the adversarial sample image and the number of iterations for updating the interference block is not limited.

[0101] Figure 4 Schematically shows a schematic diagram of the sample generation process of the transferable adversarial attack according to an embodiment of the present disclosure.

[0102] As Figure 4As shown, the patch positions in the original sample image are processed using interference blocks to obtain adversarial sample images. The interference blocks are images with noise, and the patch positions are obtained by inputting the original sample image into the trained self-attention mechanism model. The model parameters θ of the target recognition model (i.e., the model in the figure) are updated using the original sample image and the adversarial sample image. The clean features and adversarial features of the original sample image are mixed using the feature mixing block (i.e., FMB, Feature mixing block) in the activated target recognition model to obtain the adversarial sample image to be updated. At the same time, the target recognition model can be fine-tuned using the adversarial sample image to be updated. The interference blocks in the adversarial sample image to be updated are updated using a gradient-based optimization algorithm to obtain the target adversarial sample image.

[0103] To enable the target recognition model to have the ability to recognize adversarial samples and clean samples, the target recognition model is fine-tuned using the adversarial sample image and the original sample image so that the result is close to the true value, thereby obtaining the fine-tuned model parameters. After obtaining the fine-tuned model parameters, the clean features and adversarial features are mixed through the mixing ratio r, and the perturbations in the adversarial sample image to be updated are iteratively updated for the target recognition model to obtain the target adversarial sample image. The main steps are as follows:

[0104] (1) Set the adversarial sample generation method using the mask M.

[0105] (2) Set the method to improve the transfer ability using the attention mechanism.

[0106] (3) Calculate the attention map using the attention mechanism, i.e., the patch position.

[0107] (4) Define the mask using the generated patch position.

[0108] (5) Fine-tune the target recognition model to encourage the perturbations to concentrate on the robust features.

[0109] (6) Redefine the loss function using the adversarial sample image and the original sample image.

[0110] (7) Fine-tune the target recognition model according to the SGD optimizer to obtain the updated parameters.

[0111] (8) Add a feature mixing block to the structure layer of the model to achieve the mixing process of clean features and adversarial features.

[0112] (9) Activate the feature mixing block to update the perturbations in the adversarial sample image to be updated.

[0113] When different recognition models predict the same object, they have similar attention mechanisms. Through the attention mechanism, an attention map is generated for the original sample image, and perturbations are added to the most sensitive regions of the image through the attention map. In addition, mixing clean features and adversarial features enables the model to recognize both clean samples and adversarial samples simultaneously, and the adversarial samples generated by this method have good transferability.

[0114] Figure 5 A structural block diagram of a sample generation device for transferable adversarial attacks in autonomous driving according to an embodiment of the present disclosure is schematically shown.

[0115] As Figure 5 shown, the sample generation device 500 for transferable adversarial attacks in autonomous driving according to this embodiment includes a processing module 510, a mixing processing module 520, and an updating module 530.

[0116] The processing module 510 is used to process the patch positions in the original sample image by using an interference block to obtain an adversarial sample image, and the patch positions are obtained by inputting the original sample image into a trained self-attention mechanism model. In one embodiment, the processing module 510 can be used to perform the operation S210 described above, which will not be elaborated here.

[0117] The mixing processing module 520 is used to mix and process the clean features and adversarial features of the original sample image by using a feature mixing block in the target recognition model to obtain a to-be-updated adversarial sample image. The clean features are extracted from the original sample image, the adversarial features are extracted from the adversarial sample image, the target recognition model is obtained by training a recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model. In one embodiment, the mixing processing module 520 can be used to perform the operation S220 described above, which will not be elaborated here.

[0118] The updating module 530 is used to update the interference block in the to-be-updated adversarial sample image by using a gradient-based optimization algorithm to obtain a target adversarial sample image. In one embodiment, the updating module 530 can be used to perform the operation S230 described above, which will not be elaborated here.

[0119] According to an embodiment of the present disclosure, the mixing processing module 520 includes a mixing processing sub-module. The mixing processing sub-module is used to mix and process the clean features and adversarial features of the original sample image to obtain a to-be-updated adversarial sample image when the feature mixing block in the target recognition model is in a working state. The target recognition model includes a plurality of feature mixing blocks, and the plurality of feature mixing blocks are respectively located in different structural layers of the target recognition model.

[0120] According to an embodiment of the present disclosure, the update module 530 includes an update sub-module. The update sub-module is configured to update the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image when the feature mixing block in the target recognition model is in a working state.

[0121] According to an embodiment of the present disclosure, the processing module further includes a mapping sub-module, a mask processing sub-module, and a superposition sub-module. The mapping sub-module is configured to map the patch position to the original sample image to obtain a target mask position. The mask processing sub-module is configured to perform mask processing on the target mask position in the original sample image to obtain a target mask image. The superposition sub-module is configured to superpose the target mask image and the interference block to obtain an adversarial sample image.

[0122] According to an embodiment of the present disclosure, the above device further includes a first training module and a determination module. The training module is configured to train a class attention map corresponding to the original sample image by using a self-attention mechanism. The determination module is configured to determine a patch position that meets a preset threshold in the class attention map.

[0123] According to an embodiment of the present disclosure, the above device further includes a second training module, a fine-tuning module, and an obtaining module. The second training module is configured to train the recognition model based on the loss function value to obtain the model parameters to be adjusted of the trained recognition model. The fine-tuning module is configured to fine-tune the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters. The obtaining module is configured to obtain the trained target recognition model while keeping the fine-tuned model parameters unchanged.

[0124] According to an embodiment of the present disclosure, the fine-tuning module includes a fine-tuning sub-module. The fine-tuning sub-module is configured to fine-tune the model parameters to be adjusted by using a gradient-based optimization algorithm to obtain fine-tuned model parameters when none of the multiple feature mixing blocks in the target recognition model are in a working state.

[0125] According to an embodiment of the present disclosure, the fine-tuning module further includes a redefinition sub-module. The redefinition sub-module is configured to redefine the loss function.

[0126] It should be noted that the part of the sample generation device for autonomous driving transferable adversarial attacks in the embodiments of the present disclosure corresponds to the part of the sample generation method for autonomous driving transferable adversarial attacks in the embodiments of the present disclosure. For the description of the part of the sample generation device for autonomous driving transferable adversarial attacks, please refer to the part of the sample generation method for autonomous driving transferable adversarial attacks specifically, and details will not be repeated here.

[0127] According to an embodiment of the present disclosure, any of the processing module 510, the hybrid processing module 520, and the update module 530 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the processing module 510, the hybrid processing module 520, and the update module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the processing module 510, the hybrid processing module 520, and the update module 530 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0128] Figure 6 A block diagram of an electronic device suitable for implementing a sample generation method for an autonomous driving transferable adversarial attack according to an embodiment of the present disclosure is schematically shown.

[0129] As Figure 6 shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0130] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0131] According to an embodiment of the present disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.

[0132] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0133] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0134] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method for generating samples of transferable adversarial attacks for autonomous driving provided by the embodiments of the present disclosure.

[0135] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiments of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0136] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0137] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiments of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0138] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0140] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0141] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A sample generation method for transferable adversarial attack in autonomous driving, characterized in that: The method comprises: Using an interference block to process the patch position in the original sample image to obtain an adversarial sample image, the interference block is an image with noise, and the patch position is obtained by inputting the original sample image into a trained self-attention mechanism model; Using a feature mixing block in a target recognition model, the clean features and the adversarial features of the original sample image are mixed to obtain an adversarial sample image to be updated, wherein the clean features are obtained by feature extraction from the original sample image, and the adversarial features are obtained by feature extraction from the adversarial sample image. The target recognition model is obtained by training the recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; Using a gradient-based optimization algorithm to update the interference block in the adversarial sample image to be updated, to obtain a target adversarial sample image; The step of mixing the clean features and the adversarial features of the original sample image using the feature mixing block in the target recognition model to obtain the adversarial sample image to be updated includes: When the feature mixing block in the target recognition model is in a working state, the clean features and the adversarial features of the original sample image are mixed to obtain an adversarial sample image to be updated, the target recognition model includes a plurality of the feature mixing blocks, and the plurality of the feature mixing blocks are respectively located in different structural layers of the target recognition model, and the feature mixing block is in a working state, indicating that the feature mixing block is activated with a probability p in each iteration; The step of updating the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image includes: When the feature mixing block in the target recognition model is in a working state, the interference block in the adversarial sample image to be updated is updated using the gradient-based optimization algorithm to obtain the target adversarial sample image.

2. The method according to claim 1, characterized in that The method of processing the patch position in the original sample image by using the interference block to obtain the adversarial sample image also includes: Mapping the patch position to the original sample image to obtain a target mask position; Performing mask processing on the target mask position in the original sample image to obtain a target mask image; The target mask image and the interference block are superimposed to obtain the adversarial sample image.

3. The method according to claim 2, characterized in that The method further comprises: Using the self-attention mechanism to train and obtain a category attention map corresponding to the original sample image; The patch position that meets a preset threshold is determined in the category attention map.

4. The method according to claim 1, characterized in that: The method further comprises: Training the recognition model based on the loss function value to obtain model parameters to be adjusted of the trained recognition model; Fine-tuning the model parameters to be adjusted using the gradient-based optimization algorithm to obtain fine-tuning model parameters; While keeping the fine-tuning model parameters unchanged, the trained target recognition model is obtained.

5. The method according to claim 4, characterized in that The step of fine-tuning the model parameters to be adjusted by using the gradient-based optimization algorithm to obtain the fine-tuned model parameters comprises: When the plurality of feature mixing blocks in the target recognition model are not in a working state, the gradient-based optimization algorithm is used to fine-tune the model parameters to be adjusted to obtain the fine-tuning model parameters.

6. The method according to any one of claims 4 to 5, characterized in that The step of fine-tuning the model parameters to be adjusted by using the gradient-based optimization algorithm to obtain the fine-tuned model parameters further comprises: Redefine the loss function as: L f is the redefined loss function, Represents the adversarial sample image The loss function value of , J is the loss function before redefinition, y is the label, θ t is the model parameter in the tth iteration process, Represents the loss function value of the original sample image x.

7. The method according to claim 1, characterized in that The mixing formula of the feature mixing block is as follows: f t ’ is the mixed feature corresponding to the adversarial sample image to be updated in the t-th iteration process, is the clean feature in the t-th iteration process, f t is the adversarial feature, r is the mixing ratio, and t is an integer greater than 1.

8. A sample generation device for autonomous driving transferable adversarial attack, characterized in that: The device comprises: A processing module, used to process the patch position in the original sample image using the interference block to obtain an adversarial sample image, wherein the patch position is obtained by inputting the original sample image into the trained self-attention mechanism model; A hybrid processing module, used for performing hybrid processing on the clean features and adversarial features of the original sample image using a feature hybrid block in a target recognition model to obtain an adversarial sample image to be updated, wherein the clean features are obtained by feature extraction from the original sample image, and the adversarial features are obtained by feature extraction from the adversarial sample image. The target recognition model is obtained by training a recognition model based on a loss function value, and the loss function value is obtained by processing the original sample image and the adversarial sample image based on the loss function of the recognition model; An updating module, configured to update the interference block in the adversarial sample image to be updated by using a gradient-based optimization algorithm to obtain a target adversarial sample image; The mixing processing module comprises: The mixing processing submodule is used for mixing the clean features and the adversarial features of the original sample image to obtain the adversarial sample image to be updated when the feature mixing block in the target recognition model is in a working state, wherein the target recognition model includes a plurality of the feature mixing blocks, and the plurality of the feature mixing blocks are respectively located in different structural layers of the target recognition model, and the feature mixing block is in a working state, indicating that the feature mixing block is activated with a probability p in each iteration; The update module includes: The updating submodule is used to update the interference block in the adversarial sample image to be updated by using the gradient-based optimization algorithm when the feature mixing block in the target recognition model is in a working state, so as to obtain the target adversarial sample image.

Citation Information

Patent Citations

  • Adversarial sample generation method, model training method, image recognition method and device

    CN114742170A

  • Robustness evaluation and enhancement method for deep electroencephalogram identity authentication model

    CN116541700A

  • Remote sensing image-oriented patch deployable attack resisting method, device and equipment

    CN116844052A

  • Sample generation method and device, model training method and device, image processing method and device, equipment and medium

    CN117671409A