A Physical Adversarial Coating Generation Method and Device for Vehicle Object Detectors

Through an autonomously designed differentiable end-to-end neural renderer, combining environmental feature extraction and gradient backpass modules, a realistic physical adversarial coating is generated, which solves the problem of inaccurate texture and environmental features in the existing methods, and improves the attack effect and robustness.

CN119540158BActive Publication Date: 2025-07-22HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411552641.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2024-11-01
Publication Date
2025-07-22
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

The existing physical adversarial coating method based on neural renderers cannot accurately match the vehicle surface texture and environmental characteristics, resulting in unsatisfactory attack effect.

Method used

Adopt an autonomously designed differentiable end-to-end advanced neural renderer, combining environmental feature extraction network, gradient backhaul module and random perturbation module to generate realistic physical confrontation coating.

Benefits of technology

The generated adversarial coating has better robustness and attack effect in multi-weather environments, and directly optimizes the UV map without post-processing, improving the accuracy and effect of attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540158B_ABST
    Figure CN119540158B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for generating a physical adversarial coating for a vehicle target detector, including: obtaining a mask image, a vehicle image, and shooting camera perspective data from a multi-weather vehicle dataset; using the mask image to segment the vehicle image to obtain a foreground vehicle image and a background image of the vehicle image; inputting the foreground vehicle image, predetermined vehicle model data, a vehicle UV map, and the shooting camera perspective data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective; combining the rendered vehicle image with the background image to form a complete image, inputting the image into a random perturbation module to add perturbations, and inputting the perturbed image into the target detector to obtain a detection result; using the detection result, calculating a loss through a self-designed loss function and optimizing the vehicle coating through gradient backpropagation. The vehicle adversarial coating generated by the present invention can be precisely deployed and has better robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence security, and particularly to a method and device for generating physical adversarial coatings for vehicle object detectors. Background Art

[0002] Deep neural networks have achieved very good performance in object detection tasks. However, there is an attack method against artificial intelligence systems that challenges their security - adversarial attacks. Adversarial attacks refer to the attack behaviors intentionally made by malicious attackers against the weaknesses, defects, or vulnerabilities of artificial intelligence systems to mislead, damage, or deceive artificial intelligence systems. These attacks may cause artificial intelligence systems to output incorrect decisions or results. Therefore, studying adversarial attacks is of great significance and urgent need for ensuring the security of object detection models.

[0003] There are two main existing methods for generating physical adversarial coatings based on neural renderers. One is the method based on World-align projection to optimize a 2D square texture pattern and repeatably magnify it to obtain a coating that can cover the entire vehicle. The other is the method based on UV texture projection to directly and precisely optimize 3D coatings, which uses the corresponding UV map of the vehicle in the form of UV mapping for coating projection.

[0004] However, both of these methods currently have some problems. The method based on World-align projection cannot project the texture pattern onto the car in the same projection manner during the evaluation process as during the optimization process, resulting in the generated adversarial coating not being able to accurately match the vehicle, and thus having a poor attack effect on the object detector during testing. The latter rendering method based on UV mapping usually uses a relatively backward neural renderer that cannot render complex environmental features such as light and weather, resulting in a huge difference between the generated detection pictures and the real scene, and thus an unsatisfactory camouflage generation effect. In addition, the renderer used in the existing rendering method based on UV mapping cannot guarantee the differentiable path from the UV map to an intermediate rendering tensor used during the rendering process. So the current method optimizes this intermediate rendering tensor and finally post-processes this intermediate rendering tensor into a UV map that can be deployed. However, since this post-processing involves some stretching and translation of textures, this may lead to inaccurate generated adversarial coatings, resulting in a decline in the attack effect. Summary of the Invention

[0005] To solve the technical problem that there are texture differences or environmental differences between the vehicle images obtained by the renderer and the real world in the existing physical adversarial painting generation process, the embodiments of the present invention provide a method and device for generating physical adversarial painting for a vehicle target detector based on an end-to-end advanced neural renderer with accurate texture mapping and realistic effects designed independently. The technical solutions are as follows:

[0006] On the one hand, a method for generating physical adversarial painting for a vehicle target detector based on a neural renderer with accurate texture mapping and realistic effects designed independently is provided. The method includes:

[0007] S1. Obtain a mask image, a vehicle image, and shooting camera perspective data from a multi-weather vehicle dataset;

[0008] S2. Segment the vehicle image using the mask image to obtain the foreground vehicle image and the background image of the vehicle image;

[0009] S3. Input the foreground vehicle image, predetermined vehicle model data, vehicle UV map, and the shooting camera perspective data into the independently designed differentiable end-to-end advanced neural renderer to render the vehicle image in a specific scene and perspective;

[0010] S4. Combine the rendered vehicle image with the background image to form a complete image, input the complete image into a random perturbation module to add perturbations, and input the perturbed image into the target detector to obtain a detection result;

[0011] S5. Use the detection result of the target detector to calculate the loss through the independently designed loss function and optimize the vehicle painting through gradient backpropagation.

[0012] Optionally, the differentiable end-to-end advanced neural renderer includes:

[0013] An environmental feature extraction network, including an encoding module and a decoding module. In the encoding module, the vehicle image is subjected to multiple convolution processes using corresponding depthwise separable convolution algorithms. In the decoding module, corresponding upsampling operation transposed convolution or bilinear interpolation is used to restore the encoded vehicle image to the original size to restore the spatial resolution;

[0014] A sampling module, using a UV map traversal sampling method to enable all points on the UV map to participate in the assignment process of the intermediate rendering tensor;

[0015] The gradient backpropagation module optimizes gradient backpropagation by writing corresponding backpropagation code according to the CUDA forward code for converting the UV map into the intermediate rendering tensor during the process of converting the UV map into the intermediate rendering tensor, while retaining all the information required during the gradient backpropagation process, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map, realizing end-to-end UV map optimization;

[0016] The renderer rasterizes the converted intermediate rendering tensor to render a 2D image with painting and the same viewing angle as the input camera view.

[0017] Optionally, the training of the environmental feature extraction network of the differentiable end-to-end advanced neural renderer includes:

[0018] Taking the original white vehicle image in the multi-color vehicle dataset as the input of the environmental feature extraction network, and outputting two environmental feature maps;

[0019] Taking the vehicle images of other colors in the multi-color vehicle dataset as the real images of the environmental feature extraction network;

[0020] Inputting the vehicle images, predetermined vehicle model data, vehicle UV map, and shooting camera view data in the multi-color vehicle dataset into the renderer to obtain a rendering result;

[0021] Multiplying the rendering result by the first environmental feature map and adding it to the second environmental feature map, and outputting a final rendering image similar to the real image;

[0022] Calculating the loss between the final rendering image and the real image, and using the binary cross-entropy loss function for backpropagation of the reverse gradient to optimize the environmental feature extraction network.

[0023] Optionally, the sampling module of the differentiable end-to-end advanced neural renderer is specifically used for:

[0024] Initializing the RGB values of all points on the intermediate rendering tensor to all 0;

[0025] Traversing all points on the UV map and projecting them into the three-dimensional space of the intermediate rendering tensor;

[0026] Finding 8 neighboring points around the projection position on the intermediate rendering tensor according to the projection position of each point;

[0027] Calculating weights according to the distance between the projection position of each point and each neighboring point, and splitting the RGB pixel values of each point on the UV map according to the calculated weights and adding them to the RGB pixel values of the surrounding 8 neighboring points;

[0028] After the final traversal is completed, normalization processing is performed.

[0029] Optionally, in step S3, inputting the foreground vehicle image, the predetermined vehicle model data, the vehicle UV map, and the shooting camera perspective data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective includes:

[0030] Inputting the foreground vehicle image into the environmental feature extraction network to obtain two environmental feature maps;

[0031] Inputting the vehicle UV map into the sampling module, and using the UV map traversal sampling method to enable all points on the UV map to participate in the assignment process of the intermediate rendering tensor;

[0032] Inputting the vehicle UV map into the gradient backpropagation module. In the process of converting the UV map into the intermediate rendering tensor, according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, optimize the gradient backpropagation by writing the corresponding backpropagation code, and at the same time retain all the information required during the gradient backpropagation process, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map to achieve end-to-end UV map optimization;

[0033] Inputting the predetermined vehicle model data, the converted intermediate rendering tensor, and the shooting camera perspective data into the renderer to obtain a rendering result;

[0034] Multiplying the rendering result by the first environmental feature map and adding it to the second environmental feature map, outputting to obtain the final rendering result, and using a clipping function to ensure that the final rendering result does not exceed the pixel value range of 0 - 1.

[0035] Optionally, in step S4, adding perturbations to the complete image in the random perturbation module includes:

[0036] Adding random brightness changes, rotation changes, translation changes, zoom changes, and contrast changes to the complete image as a whole to improve the diversity of the generated images containing adversarial samples, thereby enhancing the robustness of the finally generated physical adversarial coating.

[0037] Optionally, the loss function calculation method in step S5 is: using the detection result of the target detector input with the perturbed image, calculating the smooth loss and the adversarial loss to obtain the result L of the loss function:

[0038]

[0039] Wherein, H b(x) represents the target detection box; gt represents the target annotation box; IoU(H b (x), gt) represents H b the intersection over union between H o (x) and gt; H c (x) and H i,j (x) respectively represent the target presence probability score and class confidence score in the bounding box; x

[0040] On the other hand, a physical adversarial painting generation device for a vehicle target detector is provided. The device is used to implement the above-mentioned physical adversarial painting generation method for a vehicle target detector. The device includes:

[0041] An information acquisition module, configured to acquire a mask image, a vehicle image, and shooting camera perspective data from a multi-weather vehicle dataset;

[0042] An image preprocessing module, configured to segment the vehicle image by using the mask image to obtain a foreground vehicle image and a background image of the vehicle image;

[0043] A rendering module, configured to input the foreground vehicle image, predetermined vehicle model data, a vehicle UV map, and the shooting camera perspective data into a self-designed differentiable end-to-end advanced neural renderer to render a vehicle image under a specific scene and perspective;

[0044] A target detection module, configured to combine the rendered vehicle image with the background image to form a complete image, input the complete image into a random perturbation module to add perturbations, and input the perturbed image into a target detector to obtain a detection result;

[0045] A painting optimization module, configured to use the detection result of the target detector to calculate a loss through a self-designed loss function and optimize the vehicle painting through gradient backpropagation.

[0046] On the other hand, an electronic device is provided. The electronic device includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, the steps of the above-mentioned physical adversarial painting generation method for a vehicle target detector are implemented.

[0047] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the computer-readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the steps of the above-mentioned physical adversarial painting generation method for a vehicle target detector.

[0048] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: solving the technical problem that there is a huge difference between the vehicle pictures rendered during the generation of adversarial coatings in the past and the real scene, thereby generating adversarial coatings with stronger attack effects; and, the generated vehicle adversarial coatings take into account the characteristics in multiple weather environments and have better robustness; and, the generated adversarial coatings are in the form of UV maps and can be implemented more accurately; and, the optimized adversarial coatings are no longer intermediate rendering tensors, but direct UV maps, and no longer require final post-processing conversion, improving the attack effect of the generated adversarial coatings. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 is a flowchart of a method for generating physical adversarial coatings for vehicle object detectors based on a renderer provided in an embodiment of the present invention;

[0051] Figure 2 is a schematic diagram of the UV traversal sampling algorithm used in the sampling module of the advanced neural renderer provided in an embodiment of the present invention;

[0052] Figure 3 is a schematic block diagram of the process of generating physical adversarial coatings for vehicle object detectors based on a renderer provided in an embodiment of the present invention;

[0053] Figure 4 is a block diagram of a device for generating physical adversarial coatings for vehicle object detectors based on a renderer provided in an embodiment of the present invention;

[0054] Figure 5 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will describe the technical solutions in the present invention with reference to the drawings.

[0056] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two. In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0057] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0058] To solve the problem of inaccurate texture mapping of the rendered picture or environmental information generated during the existing optimization of adversarial painting, the embodiments of the present invention propose a physical adversarial painting generation framework with better robustness, stronger attack effect, and can be accurately implemented. By using a differentiable end-to-end advanced neural renderer that can accurately map textures and has a realistic effect, it can precisely fuse vehicle and environmental information to generate more realistic images, thereby achieving a better attack effect. In addition, the embodiments of the present invention also use a multi-color vehicle dataset to train the environmental feature extraction network of the differentiable end-to-end advanced neural renderer, and a multi-weather vehicle dataset to train vehicle painting to enhance its robustness in different weather conditions.

[0059] First, the embodiment of the present invention prepares a multi-color vehicle dataset with a mask image, a vehicle image, and a camera angle. The mask image is used to segment the vehicle image to obtain the foreground vehicle image and the background image of the vehicle image. Subsequently, the environmental feature extraction network is trained using the multi-color vehicle dataset, and then the existing neural renderer is upgraded using the environmental feature extraction network to achieve an advanced neural renderer with accurate texture mapping and realistic effects. Moreover, a sampling module and a gradient backpropagation module are added to the differentiable end-to-end advanced neural renderer. The sampling module uses the UV map traversal sampling method to make all points on the UV map participate in the assignment process of the intermediate rendering tensor. The gradient backpropagation module, during the process of converting the UV map into the intermediate rendering tensor, optimizes the gradient backpropagation by writing the corresponding backpropagation code according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, while retaining all the information required during the gradient backpropagation process, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map to achieve end-to-end UV map optimization. Then, the foreground vehicle image, the predetermined vehicle model data, the vehicle UV map, and the shooting camera perspective data are input into the self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective, and the rendering result is combined with the background, and then input into the random perturbation module to add perturbations. Finally, the perturbed image is input into the target detector to obtain the result, the loss function is calculated, and the adversarial painting is optimized using gradient backpropagation.

[0060] Specifically, the embodiment of the present invention provides a method for generating physical adversarial painting for a vehicle target detector based on a differentiable end-to-end advanced neural renderer. This method can be implemented by an electronic device, which can be a terminal or a server. As Figure 1 shown, the processing flow of this method can include the following steps:

[0061] Step S1: Obtain a mask image, a vehicle image, and shooting camera perspective data from the multi-weather vehicle dataset.

[0062] The multi-weather data is intended to show that in order to make the generated adversarial painting robust in a multi-weather environment, adding various weather conditions to the training data can significantly improve the robustness of the attack. Therefore, in some embodiments of the present invention, a multi-weather vehicle dataset is generated in the Unreal Engine by changing the weather parameters; thereby, a vehicle dataset with different weathers is obtained. Compared with the dataset captured in reality, generating a multi-weather vehicle dataset using the Unreal Engine has the advantages of high efficiency and low cost.

[0063] The model data of the vehicle and the vehicle's UV map are pre-prepared materials, which are the same as the vehicle model in the multi-weather vehicle dataset.

[0064] In some exemplary embodiments of the present invention, the vehicle image can be generated by Unreal Engine. The vehicle image contains road traffic information, such as pedestrians, traffic signs, etc. The mask image can set the range value of the vehicle in the vehicle image to 0 and set other background areas outside the vehicle image to 1.

[0065] In some exemplary embodiments of the present invention, the shooting camera perspective data includes: camera relative vehicle azimuth angle data, camera relative vehicle elevation angle data, and camera relative vehicle distance data.

[0066] In some exemplary embodiments of the present invention, the vehicle model file may include data information such as points, faces, normals, and texture coordinates of the vehicle model.

[0067] Step S2: Use the mask image to segment the vehicle image to obtain the foreground vehicle image and the background image of the vehicle image.

[0068] In some exemplary embodiments of the present invention, a semantic segmentation camera in Unreal Engine can be used to obtain a semantic segmentation image, and then through color region judgment processing, a vehicle mask image is output. Then use the mask image to segment the vehicle image. The segmentation method is: multiply the mask image by the vehicle image to obtain the background image; perform pixel-level binary inversion on the mask image; multiply the binary-inverted mask image by the vehicle image to obtain the foreground vehicle image.

[0069] Step S3: Input the foreground vehicle image, predetermined vehicle model data, vehicle UV map, and the shooting camera perspective data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective. In this way, a rendered image, that is, a rendered vehicle image, can be obtained.

[0070] Optionally, the differentiable end-to-end advanced neural renderer includes:

[0071] An environmental feature extraction network, including an encoding module and a decoding module. In the encoding module, the vehicle image is subjected to multiple convolutional processes using a corresponding depthwise separable convolutional algorithm. In the decoding module, a corresponding upsampling operation transposed convolution or bilinear interpolation is used to restore the encoded vehicle image to its original size to restore the spatial resolution;

[0072] A sampling module that uses a UV map traversal sampling method to enable all points on the UV map to participate in the assignment process of the intermediate rendering tensor (the existing sampling module uses an intermediate rendering tensor element traversal method, which will cause the problem that some points on the UV map cannot be optimized. The embodiment of the present invention is optimized to UV map element traversal, which can optimize all points on the UV map);

[0073] The gradient backpropagation module, during the process of converting the UV map into the intermediate rendering tensor, optimizes the gradient backpropagation by writing corresponding backpropagation code according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, while retaining all the information required during the gradient backpropagation process, enabling the gradient to be backpropagated from the intermediate rendering tensor to the vehicle UV map during the gradient backpropagation process, achieving end-to-end UV map optimization, and solving the problem that the backpropagation of the existing renderer can only be backpropagated to the intermediate rendering tensor and cannot be directly backpropagated to the UV map, and it is necessary to convert the intermediate rendering tensor into a deployable UV map through a non-differentiable conversion method at the end, resulting in texture effect loss (the existing renderer uses CUDA code for forward propagation implementation during the process of converting the UV map into the intermediate rendering tensor, but since the CUDA code does not have automatic gradient backpropagation code and the existing renderer does not include the backpropagation CUDA code for this conversion, the existing renderer cannot backpropagate the gradient to the UV map. The embodiment of the present invention supplements this part of the backpropagation code to achieve effective end-to-end optimization of the UV map);

[0074] The renderer rasterizes the converted intermediate rendering tensor to render a 2D image with painting and the same viewing angle as the input camera view.

[0075] Optionally, the training of the environmental feature extraction network of the differentiable end-to-end advanced neural renderer includes:

[0076] Taking the original white vehicle image in the multi-color vehicle dataset as the input of the environmental feature extraction network, and outputting two environmental feature maps;

[0077] Taking the vehicle images of other colors in the multi-color vehicle dataset as the real images of the environmental feature extraction network;

[0078] Inputting the vehicle images, predetermined vehicle model data, vehicle UV map, and shooting camera view data in the multi-color vehicle dataset into the renderer to obtain a rendering result;

[0079] Multiplying the rendering result by the first environmental feature map and adding it to the second environmental feature map, and outputting a final rendering image similar to the real image;

[0080] Calculating the loss between the final rendering image and the real image, performing backpropagation of the reverse gradient using the binary cross-entropy loss function, and optimizing the environmental feature extraction network using the reverse gradient propagation of the loss.

[0081] In the embodiments of the present invention, it is first necessary to train the environmental feature extraction network of the differentiable end-to-end advanced neural renderer. When training the environmental feature extraction network of the differentiable end-to-end advanced neural renderer, a multi-color vehicle training set is obtained. This training set includes multiple groups of training samples. Each group of training samples includes: a foreground vehicle image, predetermined vehicle model data, a vehicle UV map of a certain color (the UV map refers to a solid-color UV map under a specific color), shooting camera perspective data, and the corresponding rendered real image of the vehicle of the corresponding color; and the environmental feature extraction network is trained using multiple groups of training samples.

[0082] In some exemplary embodiments of the present invention, the original white vehicle image in the multi-color vehicle dataset is used as the input of the environmental feature extraction network, and two environmental feature maps are output. The white vehicle image refers to setting the vehicle painting color to white in the Unreal Engine. The environmental feature extraction network is composed of a four-layer deep encoder-decoder deep neural network, and finally two environmental feature maps of the same size as the original image are output.

[0083] In some exemplary embodiments of the present invention, the vehicle images of other colors in the multi-color vehicle dataset are used as the real images of the environmental feature extraction network. The vehicle images of other colors refer to setting the vehicle painting color to other colors in the Unreal Engine, and the specific color is the same as the color of the UV map input to the following renderer.

[0084] In some exemplary embodiments of the present invention, the vehicle image, predetermined vehicle model data, vehicle UV map, and shooting camera perspective data in the multi-color vehicle dataset are input into the renderer of the differentiable end-to-end advanced neural renderer to obtain a rendering result.

[0085] In some exemplary embodiments of the present invention, the rendering result is multiplied by the first environmental feature map and added to the second environmental feature map, so as to obtain a final rendered image similar to the real image with environmental information and accurate texture mapping.

[0086] In some exemplary embodiments of the present invention, the final rendering result and the real image are used to calculate the loss, and the binary cross-entropy loss function is used for backpropagation of the reverse gradient to optimize the environmental feature extraction network.

[0087] Optionally, the sampling module of the differentiable end-to-end advanced neural renderer is specifically used for:

[0088] Initialize the RGB values of all points on the intermediate rendering tensor to all 0;

[0089] Traverse all points on the UV map and project them into the three-dimensional space of the intermediate rendering tensor (based on patches);

[0090] Based on the projection position of each point, find 8 neighboring points around the projection position on the intermediate rendering tensor;

[0091] Perform weight calculation according to the distance between the projection position of each point and each neighboring point. According to the calculated weights, split the RGB pixel values of each point on the UV map and add them to the RGB pixel values of the 8 neighboring points around;

[0092] After the final traversal is completed, perform normalization processing.

[0093] The UV traversal sampling algorithm used by the sampling module of the differentiable end-to-end advanced neural renderer according to the embodiments of the present invention is as Figure 2 shown.

[0094] Optionally, the step S3 of inputting the foreground vehicle image, the predetermined vehicle model data, the vehicle UV map, and the shooting camera perspective data into the self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective includes:

[0095] Input the foreground vehicle image into the environmental feature extraction network to obtain two environmental feature maps;

[0096] Input the vehicle UV map into the sampling module, and use the UV map traversal sampling method to enable all points on the UV map to participate in the assignment process of the intermediate rendering tensor;

[0097] Input the vehicle UV map into the gradient backpropagation module. During the process of converting the UV map into the intermediate rendering tensor, according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, optimize the gradient backpropagation by writing the corresponding backpropagation code, and at the same time retain all the information required during the gradient backpropagation process, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map to achieve end-to-end UV map optimization;

[0098] Input the predetermined vehicle model data, the converted intermediate rendering tensor, and the shooting camera perspective data into the renderer to obtain a rendering result;

[0099] Multiply the rendering result by the first environmental feature map and add it to the second environmental feature map, output the final rendering result, and use a clipping function to ensure that the final rendering result does not exceed the pixel value range of 0 - 1.

[0100] S4. Combine the rendered vehicle image with the background image to form a complete image, input the complete image into the random perturbation module to add perturbations, and input the perturbed image into the target detector to obtain a detection result;

[0101] Optionally, adding perturbations to the complete image in S4 includes:

[0102] Adding random brightness changes, rotation changes, translation changes, zoom changes, and contrast changes to the entire complete image to improve the diversity of the generated pictures containing adversarial samples, thereby enhancing the robustness of the finally generated physical adversarial coating.

[0103] The target detector refers to the target white-box detector for the attack, which can be various detectors trained on general datasets or a target detector trained on a specific vehicle dataset.

[0104] S5. Using the detection results of the target detector, calculate the loss through a self-designed loss function and optimize the vehicle coating through gradient backpropagation.

[0105] Optionally, the loss function calculation method in S5 is as follows: Using the image after perturbation as the detection result input to the target detector, calculate the smooth loss and the adversarial loss to obtain the result L of the loss function:

[0106]

[0107] Among them, H b (x) represents the target detection box; gt represents the target annotation box; IoU(H b (x), gt) represents the intersection over union between H b (x) and gt; H o (x) and H c (x) respectively represent the target existence probability score and the class confidence score in the bounding box; x i,j is the pixel value at the position where the abscissa is i and the ordinate is j in the UV map; H and W represent the length and width of the UV map; α and β are hyperparameters that control the contribution degree of the loss.

[0108] It can be understood that adversarial attacks can be divided into digital adversarial attacks and physical adversarial attacks. The former introduces pixel-level perturbations in the model input, while the latter modifies real objects or environments, thereby indirectly affecting the model input. Since directly accessing the model input usually requires system permissions, physical attacks are generally considered to be more practical. However, physical adversarial attacks are inherently more challenging because they must be robust in complex physical environments (such as viewing angles, spatial distances, lighting, and weather). In some exemplary embodiments of the present invention, the loss function for optimizing the adversarial coating consists of two parts, namely the adversarial loss and the smooth loss;

[0109] Among them, the adversarial loss function designed in the embodiments of the present invention is calculated using the output of the target detector:

[0110] H d (x) = IoU(H b (x), gt) * H c (x) * H o (x)

[0111] L 对抗 = -log(1 - max(H a (x)))

[0112] Wherein, H b (x) represents the target detection box; gt represents the target annotation box; IoU(H b (x), gt) represents the intersection over union between H b (x) and gt; H o (x) and H c (x) respectively represent the target presence probability score and the class confidence score of the bounding box; H d (x) represents the detection score, which is the product of the target presence probability, the class confidence, and the intersection over union. In some embodiments, the highest H d (x) among all detection boxes can be selected and used to calculate L 对抗 through the negative logarithm loss. By minimizing L 对抗 , the painted vehicle can be misclassified or not detected by the target detector.

[0113] Additionally, a smooth loss can be introduced into the loss function to increase the texture consistency. The smooth loss function of the rendered image is calculated by the following formula (since the optimization variable of the existing method is the intermediate rendering tensor and it cannot calculate the smooth loss, the smooth loss is calculated on the rendered result image. In the embodiments of the present invention, since the end-to-end optimization of the UV map is achieved, the smooth loss can be calculated on the UV map, which is a more direct way to calculate the smooth loss and the texture is smoother):

[0114]

[0115] Wherein, x i,j is the pixel value at the position with abscissa i and ordinate j in the UV map; H and W represent the length and width of the UV map;

[0116] Finally, the loss function can be summarized as:

[0117] L 最终 = αL 对抗 (x) + βL 光滑

[0118] Wherein, α and β are hyperparameters that control the contribution degree of the loss.

[0119] After sorting, the following formula can be obtained to obtain the result L of the loss function:

[0120] L = -α * log(1 - max(Iou(H b (x), gt) * H c (x) * H o (x)))

[0121]

[0122] Figure 3 It is a schematic block diagram of the physical adversarial painting generation process for a vehicle target detector based on a differentiable end-to-end advanced neural renderer shown according to an exemplary embodiment.

[0123] In the embodiment of the present invention, the foreground vehicle image, the predetermined vehicle model data, the vehicle UV map, and the shooting camera view data are input into the differentiable end-to-end advanced neural renderer for rendering. The rendering process of the optimized differentiable end-to-end advanced neural renderer alleviates the limitations of the existing attack methods based on backward renderers. For example, in the prior art, the World-align based method cannot accurately apply the trained adversarial painting to the vehicle surface. This may weaken its attack performance in the real world and the simulation world. Therefore, the embodiment of the present invention avoids this problem by using a neural renderer based on UV mapping. However, the existing UV mapping based methods encounter difficulties in rendering the environmental features of the vehicle surface, resulting in unrealistic rendered images. Therefore, the embodiment of the present invention introduces an environmental feature extraction network, enabling the neural renderer to combine the environmental features and the output of the renderer to obtain a realistic and accurate camouflaged vehicle image. And the existing UV mapping based methods usually cannot directly transfer the gradient to the UV map. Therefore, the embodiment of the present invention introduces a gradient backpropagation module to achieve the differentiable path in this part, and utilizes an advanced sampling module to optimize the algorithm for sampling from the UV map, so that more points can be optimized, thus realizing the true end-to-end optimization of the UV map and eliminating the inaccurate texture operation of converting the optimized intermediate rendering tensor into the UV map again in the past, thereby improving the attack effect of the generated adversarial painting. In addition, the embodiment of the present invention adds a random perturbation module, which can combine the rendered vehicle image with the background image to form a complete image, input the complete image into the random perturbation module to add perturbations, and input the perturbed image into the target detector to obtain the detection result, so as to improve the diversity of the generated images containing adversarial samples, thereby enhancing the robustness of the finally generated physical adversarial painting.

[0124] Figure 4It is a block diagram of a renderer-based physical adversarial painting generation device for a vehicle target detector shown according to an exemplary embodiment. This device can be used to implement the above-mentioned renderer-based physical adversarial painting generation method for a vehicle target detector. Referring to Figure 4 As shown in ,

[0125] , this device includes: an information acquisition module 201, an image preprocessing module 202, a rendering module 203, a target detection module 204, and a painting optimization module 205. Among them:

[0125] The information acquisition module 201 is configured to obtain a mask image, a vehicle image, and shooting camera perspective data from a multi-weather vehicle dataset;

[0126] The image preprocessing module 202 is configured to segment the vehicle image using the mask image to obtain a foreground vehicle image and a background image of the vehicle image;

[0127] The rendering module 203 is configured to input the foreground vehicle image, predetermined vehicle model data, a vehicle UV map, and the shooting camera perspective data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and perspective;

[0128] The target detection module 204 is configured to form a complete image by combining the rendered vehicle image with the background image, input the complete image into a random perturbation module to add perturbations, and input the perturbed image into a target detector to obtain a detection result;

[0129] The painting optimization module 205 is configured to use the detection result of the target detector to calculate a loss through a self-designed loss function and optimize the vehicle painting through gradient backpropagation.

[0130] For the sake of convenience of description, Figure 4 only the main components of this device are shown. The device of this embodiment can be used to execute Figure 1 the technical solutions of the method embodiments shown. Its implementation principle and technical effects are similar and will not be elaborated here.

[0131] In an exemplary embodiment, the present invention also provides an electronic device, and the electronic device includes:

[0132] a processor;

[0133] a memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the above-mentioned renderer-based physical adversarial painting generation method for a vehicle target detector are implemented.

[0134] Figure 5 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention, as shown in Figure 5 Figure 5As shown, the electronic device 300 may include a processor 3001 and a memory 3002. Optionally, the electronic device 300 may further include a transceiver 3003. Among them, the processor 3001 is connected to the memory 3002 and the transceiver 3003, for example, through a communication bus. Computer-readable instructions are stored on the memory 3002, and when the computer-readable instructions are executed by the processor 3001, the steps of the above-mentioned renderer-based physical adversarial painting generation method for vehicle target detectors are implemented.

[0135] In a specific implementation, as an embodiment, the processor 3001 may include one or more CPUs, for example Figure 5 the CPU0 and CPU1 shown in

[0136] In a specific implementation, as an embodiment, the electronic device 300 may also include multiple processors, for example Figure 5 the processor 3001 and the processor 3004 shown in

[0137] Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0138] Among them, the memory 3002 is used to store the software program for implementing the solution of the present invention and is controlled by the processor 3001 for execution. The specific implementation manner may refer to the above method embodiment and will not be elaborated here.

[0139] The transceiver 3003 is used to communicate with a network device or with a terminal device.

[0140] Optionally, the transceiver 3003 may be integrated with the processor 3001 or may exist independently and be coupled to the processor 3001 through the interface circuit of the electronic device 300. The embodiments of the present invention do not make specific limitations on this.

[0141] It should be noted that Figure 5 the structure of the electronic device 300 shown in

[0142] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the steps of the above-described renderer-based physical adversarial painting generation method for a vehicle target detector. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0143] It should be noted that, in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an ……" does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.

[0144] When "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. are mentioned in the specification, it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific feature, structure or characteristic. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0145] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0146] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or its similar expression refers to any combination of these items, including any combination of single (item) or plural items. For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.

[0147] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0148] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0149] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0151] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks and other various media that can store program codes.

[0152] The present invention covers any alternatives, modifications, equivalent methods, and solutions that are made within the spirit and scope of the present invention. For the purpose of enabling the public to thoroughly understand the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention even without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail in order to avoid unnecessary confusion to the essence of the present invention.

[0153] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a physical adversarial coating for a vehicle target detector, characterized in that, The method includes: S1. Obtain a mask image, a vehicle image, and camera view data of the shooting camera from a multi-weather vehicle dataset; S2. Use the mask image to segment the vehicle image to obtain a foreground vehicle image and a background image of the vehicle image; S3. Input the foreground vehicle image, predetermined vehicle model data, a vehicle UV map, and the camera view data of the shooting camera into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and view; S4. Combine the rendered vehicle image with the background image to form a complete image, input the complete image into a random perturbation module to add perturbations, and input the perturbed image into a target detector to obtain a detection result; S5. Use the detection result of the target detector to calculate a loss through a self-designed loss function and optimize the vehicle painting through gradient backpropagation; The differentiable end-to-end advanced neural renderer includes: An environmental feature extraction network, including an encoding module and a decoding module. In the encoding module, the vehicle image is subjected to multiple convolutional processes using a corresponding depthwise separable convolutional algorithm. In the decoding module, a corresponding upsampling operation transposed convolution or bilinear interpolation is used to restore the encoded vehicle image to its original size to restore the spatial resolution; A sampling module that uses a UV map traversal sampling method to enable all points on the UV map to participate in the assignment process of the intermediate rendering tensor; A gradient backpropagation module. During the process of converting the UV map into the intermediate rendering tensor, according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, corresponding backpropagation code is written to optimize the gradient backpropagation, and at the same time, all information required during the gradient backpropagation process is retained, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map to achieve end-to-end UV map optimization; A renderer that rasterizes the converted intermediate rendering tensor to render a 2D image with painting and the same view as the input camera view; The sampling module of the differentiable end-to-end advanced neural renderer is specifically used for: Initialize the RGB values of all points on the intermediate rendering tensor to all 0; Traverse all points on the UV map and project them into the three-dimensional space of the intermediate rendering tensor; According to the projection position of each point, find 8 neighboring points around the projection position on the intermediate rendering tensor; Perform weight calculation based on the distance between the projection position of each point and each neighboring point. According to the calculated weights, split the RGB pixel values of each point on the UV map and add them to the RGB pixel values of the surrounding 8 neighboring points; Finally, after the traversal is completed, perform normalization processing.

2. The method according to claim 1, wherein The training of the environmental feature extraction network of the differentiable end-to-end advanced neural renderer includes: Use the original white vehicle image in the multi-color vehicle dataset as the input of the environmental feature extraction network, and output two environmental feature maps; Use the vehicle images of other colors in the multi-color vehicle dataset as the real images of the environmental feature extraction network; Input the vehicle images, predetermined vehicle model data, vehicle UV maps, and camera view data in the multi-color vehicle dataset into the renderer to obtain a rendering result; Multiply the rendering result by the first environmental feature map and add it to the second environmental feature map, and output the final rendered image similar to the real image; Calculate the loss between the final rendered image and the real image, and use the binary cross-entropy loss function for backpropagation of the reverse gradient to optimize the environmental feature extraction network.

3. The method according to claim 1, wherein In step S3, input the foreground vehicle image, predetermined vehicle model data, vehicle UV map, and the camera view data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and view, including: Input the foreground vehicle image into the environmental feature extraction network to obtain two environmental feature maps; Input the vehicle UV map into the sampling module, and use the UV map traversal sampling method to make all points on the UV map participate in the assignment process of the intermediate rendering tensor; Input the vehicle UV map into the gradient-backpropagation module. During the process of converting the UV map into the intermediate rendering tensor, according to the CUDA forward code for converting the UV map into the intermediate rendering tensor, optimize the gradient backpropagation by writing the corresponding backpropagation code, and at the same time retain all the information required during the gradient backpropagation process, so that during the gradient backpropagation process, the gradient can be backpropagated from the intermediate rendering tensor to the vehicle UV map to achieve end-to-end UV map optimization; Input the predetermined vehicle model data, the converted intermediate rendering tensor, and the camera view data into the renderer to obtain a rendering result; Multiply the rendering result by the first environmental feature map and add it to the second environmental feature map, output the final rendering result, and use a clipping function to ensure that the final rendering result does not exceed the pixel value range of 0-1.

4. The method according to claim 1, wherein In step S4, adding perturbations to the complete image by inputting it into the random perturbation module includes: Add random brightness changes, rotation changes, translation changes, zoom changes, and contrast changes to the complete image as a whole to improve the diversity of the generated images containing adversarial samples, thereby enhancing the robustness of the finally generated physical adversarial coating.

5. The method according to claim 1, wherein In step S5, the loss function calculation method is: use the image after perturbation to input the detection result of the target detector, calculate the smooth loss and the adversarial loss, and obtain the result L of the loss function; Among them, H b (x) represents the target detection box; gt represents the target annotation box; IoU(H b (x), gt) represents the intersection over union between H b (x) and gt; H o (x) and H c (x) represent the target existence probability score and the class confidence score in the bounding box respectively; x i,j is the pixel value at the position where the abscissa is i and the ordinate is j in the UV map; H and W represent the length and width of the UV map; α and β are hyperparameters that control the contribution degree of the loss.

6. A physical adversarial coating generation device for a vehicle target detector, the device being used to implement the method according to any one of claims 1 to 5, characterized in that, The device includes: An information acquisition module for acquiring a mask image, a vehicle image, and camera view data from the multi-weather vehicle dataset; An image preprocessing module for segmenting the vehicle image using the mask image to obtain the foreground vehicle image and the background image of the vehicle image; A rendering module for inputting the foreground vehicle image, predetermined vehicle model data, vehicle UV map, and the camera view data into a self-designed differentiable end-to-end advanced neural renderer to render the vehicle image under a specific scene and view; A target detection module, configured to combine the rendered vehicle image with the background image to form a complete image, input the complete image into a random perturbation module to add perturbations, and input the image after perturbations into a target detector to obtain detection results; A painting optimization module, configured to utilize the detection results of the target detector, calculate losses through a self-designed loss function, and optimize the vehicle painting through gradient backpropagation.

7. An electronic device, characterized in that, The electronic device includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 5.