Three-dimensional object local confrontation attack method based on maximum aggregation area sparse strategy

Through the maximum aggregated area sparseness (MARS) strategy and neural rendering algorithm, important decision areas of 3D targets are located and texture modifications are carried out, which solves the problem of difficulty in ensuring both visual concealment and attack effects in the existing technology, and achieves efficient and highly migratory local 3D attack effects.

CN120198353APending Publication Date: 2025-06-24CHINA PETROCHEMICAL CORP +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117712.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

It is difficult to ensure both visual concealment and attack effect in existing 3D object detection scenarios, especially in complex environments, and existing methods are difficult to achieve dual attacks between manual recognition and machine recognition, and the perspective versatility is poor.

Method used

A local 3D attack framework driven by the maximum aggregation region sparseness (MARS) strategy is proposed. Through optimization regularization of 3D Mesh panels, including maximizing the aggregation degree regularity of the adversarial patch aggregation and ensuring the minimum sparsity regularity of the adversarial patch region. Combining neural rendering algorithms and data expansion technology, it adapts to the characteristics of deep neural networks, locates important decision areas of 3D targets and performs texture modifications.

Benefits of technology

It achieves the simultaneously guaranteeing visual concealment and efficient attack effect in 3D target detection scenarios, has good attack performance and migration, and can effectively attack under multiple angles and multiple distances, surpassing the full-body optimization attack effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198353A_ABST
    Figure CN120198353A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional object local confrontation attack method based on a maximum aggregation area sparse strategy, and in a physical 3D environment, the existing target detection-oriented confrontation attack (3D-AE) mainly faces the following challenges: in order to achieve a maximum attack effect, relatively large and dispersed confrontation patches are often selected, and the confrontation effect is poor. Therefore, the confrontation patch is too obvious, and the visual concealment is reduced. A better strategy is how to utilize as small as possible, and the attack effect of the aggregated adversarial patches is maximized. In order to maximize aggregation of adversarial patches, an aggregation degree regular term is designed to constrain a shielding aggregation matrix obtained based on a surface element adjacency relation; in order to ensure minimization of an adversarial patch area, sparse regularization is designed, so that shielding weights tend to be distributed in a U shape, and extreme values are limited. Expanded data of visual angle changes are obtained through neural rendering, and a universal important decision-making area under the target multi-angle condition is positioned through suppression model detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic information technology, and in particular, to a method for three-dimensional object local adversarial attack based on a maximum aggregation region sparse strategy. Background Art

[0002] In recent years, with the rapid development of deep learning in the field of computer vision, convolutional neural networks have been increasingly widely used in 3D object detection and recognition. Current work has proven that the recognition ability of neural networks is limited and is easily confused by datasets with tiny perturbations, and then outputs incorrect predictions. This kind of phenomenon is called adversarial attack, and such samples with added perturbations are called adversarial examples. Adversarial examples can deceive neural networks including many security systems, posing threats to the security and effectiveness in the actual application process. Therefore, the research on adversarial examples is extremely urgent.

[0003] Existing work related to adversarial examples mainly focuses on the research of traditional two-dimensional image adversarial attacks, which modify some pixels of the image in the digital world to confuse the neural network and has entered a mature stage. In actual applications, we need to consider how to convert the adversarial attacks in the digital world into adversarial camouflage in the physical world. Currently, the solutions can be divided into two categories. The first is to print out and place or paste 2D adversarial patches obtained from 2D images on real-world target objects, mostly based on gradient optimization methods such as FGSM, C&W, particle swarm optimization (PSO), and reinforcement learning (RL) optimization. The second is the adversarial attack based on 3D objects, which is achieved by modifying the shape or texture of 3D objects in reality. Both of these two methods have been successfully applied in the real world, indicating that 3D object adversarial attacks pose a certain threat to deep neural networks. Among them, the adversarial attacks based on 2D images are difficult to adapt to the complex changes in the real world, and their attack effects in the 3D world are generally inferior to those of the adversarial attacks based on 3D objects.

[0004] However, the current adversarial attacks directly starting from 3D objects are still in their infancy. Currently, the methods for 3D adversarial attacks mainly include placing adversarial objects with certain shapes and textures with adversarial effects, which are difficult to achieve differentiable optimization processes; creating certain optical effects through optical control, which are difficult to handle complex environmental light sources; modifying the attributes of the target itself, which are difficult to ensure both visual concealment and effective attacks in complex environments. These methods have been applied in various security scenarios such as face recognition and vehicle recognition, but their application processes have strong limitations, and it is difficult to achieve dual attacks of manual recognition and machine recognition. The 3D attacks with obvious visual effects are universal in terms of viewing angles, while the 3D attacks with unobvious visual effects are only effective at some angles. The lack of research in this field may bring serious security risks. Summary of the Invention

[0005] The objective of the present invention is to address the deficiencies of the prior art. To solve the problem of simultaneously ensuring visual concealment and attack effectiveness in 3D target detection scenarios, a local 3D attack framework driven by the Maximum Aggregation Region Sparsity (MARS) strategy is proposed to achieve efficient attacks while ensuring visual concealment. An optimized regularization for 3D Mesh patches is designed, including an aggregation degree regularization that maximizes the aggregation of adversarial patches and a sparsity regularization that ensures the minimization of the adversarial patch region, which can well adapt to the characteristics of deep neural networks and find the important decision regions of 3D targets.

[0006] The present invention first analyzes the main problems of physical attacks in terms of optimization feasibility, namely the discontinuity problem of parameter transfer between the real object surface and the two-dimensional image. Considering the color mechanism of optical images, a fixed patch of the same size as the target object is set, and the neural rendering algorithm is used to ensure the differentiability of the optimization process parameters. Patch weights are introduced to adjust the transparency of each part of the target surface, thereby conducting attacks. To adapt to the complex environmental changes in the physical world, image data of the same target in different environments is further increased through data augmentation. In addition, to balance attack efficiency and visual concealment, under regularization norms such as the aggregation degree and sparsity coefficient, the confidence of the target is reduced, and the shape and position of the patches on the target surface are learned, which enables us to locate the general important decision regions applicable under multi-angle and multi-distance shooting conditions of the item. Texture modification within this region can achieve general local camouflage.

[0007] The technical solution of the present invention is as follows:

[0008] (1) Framework

[0009] The core of the local adversarial attack on 3D objects in the aggregation region is to find the important decision regions on the target surface. Based on the imaging mechanism of optical remote sensing images and aiming at the optimization difficulties of important decision regions, the framework is divided into the following parts.

[0010] Part 1: To ensure the feasibility of the optimization process, the patch is set as a fixed perturbation to cover the original texture information, and the neural network algorithm is used to introduce grid occlusion weights to ensure the differentiability of the optimization process parameters. The target 3D model (M, T[texture_param, mask_param]) and the shooting parameters θc (shooting angle, shooting distance, light parameters, etc.) are input into the renderer R.

[0011] Part 2: To achieve efficient attacks at multiple angles and distances, the renderer outputs a large number of target images O = R(M, T; θc) under camera parameters at multiple angles and distances. The original image X is input into the target segmenter S for segmentation to obtain the real background G, and then it is combined with the rendered image O to obtain the real image I = O + G.

[0012] Part 3. To ensure the naturalness and concealment of the visual effect and achieve the balance of the number and area of regions, under the regularization norms such as aggregation degree and sparsity coefficient, the target confidence loss is reduced, and the shape and position of the target surface patches are learned to locate the important decision regions on the target surface, so as to achieve excellent attack effects by only changing a small part of the regions.

[0013] (2) Optimization

[0014] Here, the neural rendering of the renderer used by the maximum aggregation region sparsity strategy to ensure the optimization feasibility is first elaborated, then the specific design of the optimization regularization and the setting of the loss function are described, and finally the MARS algorithm process is described.

[0015] 1) Neural rendering

[0016] The process of generating an image from a 3D world is called rendering, which is located at the boundary between the 3D world and the 2D image and is crucial in computer vision. In this application, neural rendering is used to convert 3D objects into a large number of 2D images with environmental information, which can then be input into a detector for detection and gradient backpropagation. Specifically, it is the process of converting 3Dmodel(M,T[texture,mask]) into 2D images, where M is the model mesh, T is the model texture including two parts: texture and mask, and the calculation formula of texture is: T = backgroud_color × mask_weight + original_texture × (1 - mask_weight).

[0017] Neural rendering regards the rasterization process as a process that can pass gradients back, establishes a deep relationship between 3D objects and 2D images, and thus allows optimization. A polygon mesh is used as the 3D format, and a small number of parameters are used to represent the 3D shape. Here, a rendering neural network is trained, and an approximate gradient unique to neural network rendering is proposed to transfer the gradient to texture, lighting, camera, and object shape.

[0018] The rendering pipeline is to convert the vertices {V_O_I} in the object space into the vertices {V_S_I} in the screen space. The rasterization process samples vertices and patches to generate an image, and renders the color for each patch. The difficulty in this part is that sudden changes in color will cause the gradient to be 0 and cannot be backpropagated, so linear interpolation is used to replace the gradual changes between pixels.

[0019]

[0020] The gradient at x0 (x0 ∈ [a, b]) is calculated by the following formula:

[0021]

[0022] When it means the pixel is too bright. Therefore, to optimize the loss function, we need to darken the pixel. Thus, we need Similarly, when There is no case where both have the same positive or negative sign. Otherwise, we can only reverse or not optimize the loss function.

[0023] Here, it is divided into two cases: whether the target pixel is inside or outside the patch.

[0024] If the target pixel P j is inside the patch, then define the above partial derivative as 0, and use the color inside the patch for forward propagation to avoid color leakage.

[0025] If P j is inside the patch, the color coefficients will change after linear interpolation. Therefore, first calculate the derivatives on the left and right sides of x0, and make their sum be the gradient at x0.

[0026]

[0027]

[0028]

[0029] If there are multiple overlapping patches, only draw the frontmost patch on each pixel. During the backpropagation process, check whether the intersection point is drawn. If it does not overlap with the patch, do not calculate the gradient.

[0030] Each patch has its own size s t *s t *s t texture image. Use the centroid coordinate system to determine the coordinates in the texture space corresponding to the position p on the triangle {v1, v2, v3}, and then sample from the texture image using bilinear interpolation.

[0031] Lighting can be directly applied to the network. This model only considers the ambient light l a and the directional light l d , and the color of the pixel after being affected by the lighting is as follows:

[0032]

[0033] n d is the unit vector of the directional light; n j is the normal vector of this patch; I jIs the original color.

[0034] 2) Regularization for Attack Patch Aggregation

[0035] The purpose of this regularization is to ensure that the number of covered patches is within a certain range during the training process, accelerating the optimization rate. To obtain the distribution of important regions, simply considering the remaining losses is likely to result in a continuously monotonic gradient, which is insufficient to achieve the balance between the area of the region and the adversarial effect. Therefore, the element of the sparse coefficient needs to be considered during the training process.

[0036] We observe that there are many patches with medium weights during training. They occupy a large amount of spatial area, but most of them contribute limitedly to the effectiveness of the attack. This application calls it "mask uniform distribution". It achieves the attack through a set of patches with different transparencies, violating the original intention of using fixed perturbations. It does not cover the texture features of the voxels, and the found regions do not have decision-making and transferability.

[0037] Therefore, mse_loss (Mean Square Error) is selected to restrict the weight distribution, making it tend to 0 or 1 during the training process, so as to avoid the influence of some combinations of light and dark patterns on the detection results.

[0038]

[0039] Here, I is a matrix of all 1s, M is an array of mask weights, m is the number of voxels in the network, and H(M) is the matrix form of the array for calculation.

[0040] The calculation of the sparse coefficient is also divided into two parts. The above part restricts the weight value of each patch to approach 0 or 1, and the other part restricts the number of patches with weights approaching 1, so as to achieve the selection of important regions. The second part mainly uses the patch weights set above to define the sparsity, believing that the weight value of each patch is the opacity of the fixed perturbation on that patch (if it is 0, it means the perturbation does not show on that patch; if it is 1, it means the perturbation shows highly on that patch). Then the sparsity can be defined as the sum of the power functions of the patch weights. To be consistent with the mse loss, the L2 norm is selected for sparse restriction here. The final calculation formula for the sparse coefficient loss is as follows:

[0041]

[0042] 3) Final Loss Function

[0043] To ensure the attack ability of the local area, adversarial loss needs to be added. In this application, images under the corresponding mask are generated for environmental information at different angles and input into the target detector. The masking matrix is optimized using gradient backpropagation, and then the patch weight matrix with the best deception effect on the detector is obtained.

[0044] In this experiment, the following three losses are comprehensively considered in the general loss design: the bounding box regression loss that calculates the difference between the detection box and the original ground truth box; L_cls that calculates the difference between the classification and the true class; L_obj that calculates the object confidence. To achieve efficient target attacks, the three losses are reduced at a certain hyperparameter ratio during the training process.

[0045]

[0046]

[0047]

[0048] L_obj is mainly used to judge whether each cell contains a target. When there is a target in the cell, this part of the loss will consider the classification loss and the localization loss; when there is no target in the cell, only the classification loss is considered. Here, the confidence label L and the predicted confidence P are used to calculate the BCE loss (Binary Cross Entropy) to comprehensively calculate the confidence loss.

[0049] General loss function L adv is as follows:

[0050] L adv = L bbox + L obj + L cls

[0051] To sum up, the final total loss function is as follows: The ratio of the a and b hyperparameters can control the number of perturbation patches gradually approaching the required number during the training process. At the same time, the c parameter needs to occupy a certain proportion to ensure the aggressiveness of the obtained results. There is an ablation function for hyperparameter selection in the follow-up.

[0052] L = α·L agg + β·L sparse + γ·L adv

[0053] 4) Pseudocode of the MARS algorithm

[0054] To achieve the balance between visual stealth and attack effectiveness, the MARS algorithm is designed to find the important decision regions of the target. An effective loss function is designed by combining the relevant characteristics of the important decision regions with the network detection results, and its gradient is backpropagated to guide the update of the surface parameters of the target object. The search for the important decision regions is modeled as an optimization problem. Finally, texture modification operations are performed within this region. The local adversarial attack optimization algorithm based on the decision region is as follows: It is the texture optimization method adopted within the region:

[0055]

[0056] Advantages of the present invention

[0057] The present invention proposes a method for locating the important decision regions on the surface of the target, successfully realizing the local area adversarial attack that balances the visual effect and the attack effect, and highlighting the effectiveness compared with other region selection strategies. The local area search strategy combined with the texture optimization method obtains excellent attack effects, has good attack performance and transferability. Since the modified area is small, there are few non-transferable noise features learned, and the possibility of the local optimal model for the attack model is low. The experimental results of this application exceed the attack effects of full-body optimization in some cases. Description of the drawings

[0058] Figure 1 It is a schematic diagram of the data set in the specific implementation mode

[0059] Figure 2 It is a schematic diagram of the experimental result image in the specific implementation mode

[0060] Figure 3 It is a comparison chart of the experimental result data in the specific implementation mode Specific implementation mode

[0061] The present invention will be further described below in conjunction with the embodiments, but the protection scope of the present invention is not limited thereto:

[0062] As an open-source 3D target simulation and emulation platform, Carla was initially proposed to support the training, prototyping, and verification of deep learning models. As the basic data set for the work in this field, to be consistent with the field, the experimental data set in this embodiment is Carla, as Figure 1 shown. Specifically, the Carla simulator is used to generate optical remote sensing simulation images of the target vehicle's urban driving process under different perspectives, different distances, different environments, etc. A total of 15,000 images are obtained, and control point annotations are performed on them.

[0063] Six detectors in total, including the yolov3 and yolov5 series, were first trained on the Carla dataset. Different local regions were obtained using three selection strategies: empirical manual selection, random selection, and MARS. The full-body attack was used as the benchmark control group, and the attack effects on these regions were compared and analyzed.

[0064] To measure the attack ability of adversarial attacks, the attack performance of adversarial methods was quantified through AP and a custom metric, attack efficiency. In this experiment, AP refers to the average precision; the definition of attack efficiency (AE) is as follows:

[0065]

[0066] where δ AP is the difference in AP before and after the attack, and p face is the proportion of patches modified by the attack.

[0067] The experiment was divided into three parts. The first part compared the attack effects under different selection strategies, and the main purpose of this part was to prove the reliability and superiority of MARS. The second part compared the attack performances under different texture optimization methods, and the main purpose of this part was to prove the transferability of the attack method of this application. The third part adjusted some parameters and conducted ablation experiments to explore the significance of the core parameters of the method of this application.

[0068] (1) As Figure 2As shown in the figure, in the first part of the experiment, different local regions were selected using different selection strategies, and local attacks were carried out using fixed texture or optimized texture respectively. Their performances in different detection networks under the same dataset and the same training network were compared. The training network was selected as yolov3, and the detection networks were selected as yolov3, yolov5s, yolov5x, yolov5m, yolov5n, and yolov5l. MARS proposed in this application achieved the results with the highest AP drop of 0.618 and the lowest of 0.188 under the optimized network, surpassing all other selection strategies, and the difference from the results of the full-body attack was also very small. The experimental results showed that the AE of MARS stably exceeded that of the baseline and other local attack methods, reaching 2.615 in yolov3, more than doubling that of the baseline (0.9111), and at least reaching 0.608 in the yolov5 series, showing a significant gap from 0.532 of the control group. The experiment shows that the method proposed in this application comprehensively guarantees the coverage range and regional integrity, and the stability and transferability have been greatly improved compared with other methods, indicating that it is feasible to use the Local Adversarial Attacks with Maximum Aggregated Regions Sparseness Strategy to attack the detector on 3D objects.

[0069] (2) The purpose of the second part of the experiment consists of two parts: whether MARS can combine different texture modification methods for attacks, and the attack effectiveness of the model trained on the yolo detector on other detectors. In this embodiment, the Full region, Fixed (center) region, and MARS region, which performed well and were representative in the first part of the experiment, were selected as variables, and FCA and DAS were used for texture optimization. The obtained adversarial results were compared in different detection networks. To ensure the independence of the results, the yolo series was not selected as the detection network, and the selected networks mainly included mask_rcnn, cascade_rcnn, faster_rcnn, SSD, and Retinanet. The experimental results are as Figure 3As shown, FCA and DAS have their own advantages and disadvantages in different region choosing strategies. The full attack has stable effects, with FCA and DAS achieving an average AP decrease of 0.525 and 0.522 respectively; there is a significant gap between Fixed(center)+FCA and Fixed(center)+DAS, achieving an average AP decrease of 0.296 and 0.437 respectively; MARS+FCA and MARS+DAS achieve an average AP decrease of 0.407 and 0.552 respectively. The MARS attack stably outperforms the other local attacks, and its performance on DAS even exceeds that of the full attack. This shows that the important decision regions obtained by MARS conform to the model recognition decision boundary and have strong transferability in texture modification methods. By studying the attack effects on different networks, the MARS Attack, as a local attack, exceeds the other local attack on all networks. The performance of the MARS+DAS attack not only surpasses the other local region selection strategies but also exceeds the full attack effect in networks other than cascade_rcnn. This indicates that the local attack based on MARS has strong transferability on detection networks and can perform well in most networks.

[0070] (3) In the third part of the experiment, different parameter coefficients are adjusted, mainly aiming to explore the significance of the various loss parameters used in this application. In the embodiment, experiments are first conducted for different combinations of loss parameters. Since the adversarial loss ensures the basic attack effect, it is not adjusted. During training, the loss function is set as: loss_total(all), L adv (single detection loss), α·L agg +γ·L adv (with aggregation regularization) and β·L sparse +γ·L adv(with minimization regularization). The attack model is the YOLOv3 network trained for one round with this dataset. Since only the influence of different coefficients needs to be compared here, the detection models also select the YOLOv3 and YOLOv5s networks trained for one round with the Carla dataset. The experiment ensures that the other coefficients remain unchanged, and adjusts the pre-parameter of the sparse coefficient, so as to adjust the number of sparse regularization obtained by optimization, and compare the influence of the mask size on the attack efficiency. Here, the attack model uses the YOLOv3 network trained for one round with the Carla training set, and the detection models select the YOLOv3 and YOLOv5s networks trained for one round with the Carla training set. The experimental results show that as the number of masks decreases, the impact on accuracy decreases significantly, but the attack efficiency increases slightly. For the attack effect, the influence of the number of masks is relatively weak within the range of more than 2000, while the increase or decrease within the range of 0-2000 will bring a large change in accuracy. It can be speculated that for this physical target, the optimal number of core important regions for deep neural network detection is within 2000, which further proves the importance of the starting point of this research: traditional adversarial attacks always optimize in less important regions, resulting in waste of computing resources, while the method proposed in this research can well obtain a certain number of local important regions, reduce computing power while improving the visual effect, and achieve good attack effects in different detection networks under different occlusion numbers. When discussing the attack efficiency, it is found that the attack efficiency increases steadily when the number of occlusions is reduced. This further proves the importance of the research on important decision regions. During the training process, as the number of occlusions decreases, the operation rate is improved, which can be ignored in small-scale training, but is of great significance in large-scale large model training. The purpose of the research is to ensure the training speed, obtain the visual effect while achieving a significant attack effect, and the improvement of the attack efficiency gives a positive feedback to this purpose. Finally, analyzing the overall results, the method of this application achieves good adversarial results in different networks under most occlusion number settings, indicating the wide application of this solution.

[0071] The specific embodiments described in this document are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for local adversarial attack on three-dimensional objects based on the maximum clustered region sparse strategy, applied to 3D target detection scenarios, characterized in that The method comprises: 1) Set a fixed patch of the same size as the target object, set the patch as a fixed perturbation to cover the original texture information, and use the neural network algorithm to introduce the grid occlusion weight to ensure the differentiability of the optimization process parameters; 2) Data expansion: adding image data of the same target in different environments; 3) Under the regularization specification, the confidence of the target is reduced, and the shape and position of the target surface patch are learned to locate the important decision-making area of ​​the target surface.

2. The method according to claim 1, characterized in that The data expansion is specifically as follows: using a renderer to output a large number of target images under multi-angle and multi-distance camera parameters, inputting the original image into a target segmenter for segmentation to obtain a real background; and combining the real background with the target image to obtain a real image.

3. The method according to claim 1, characterized in that Train the rendering neural network and use neural rendering to convert 3D targets into a large number of 2D images with environmental information, which can then be input into the detector for detection and gradient return.

4. The method according to claim 3, characterized in that In neural rendering, linear interpolation is used to replace the gradual change in brightness I of a pixel P.

5. The method according to claim 3, characterized in that The gradient at x0 is calculated by the following formula: Among them, x0,x i ∈[a,b], I is brightness; is the gradient signal of the loss function back-propagated to the interpolation position, It is the signal for back propagation of brightness gradient I.

6. The method according to claim 5, characterized in that Calculated by the following formula: Among them, x0,x i ∈[a,b], I is brightness; is in x i The gradient at is the signal of the back propagation of the loss function for the brightness I gradient, is the gradient signal back-propagated for the interpolation position.

7. The method according to claim 5, characterized in that If there are multiple overlapping facets, only the frontmost face is drawn at each pixel. During the back propagation process, check whether the intersection is drawn. If it does not overlap with the facet, the gradient is not calculated.

8. The method according to claim 3, characterized in that The color of the pixel after being affected by light is as follows: Among them, n d is the unit vector of the directional light; n j is the normal vector of the face element; I j The original color.

9. The method according to claim 3, characterized in that When training the rendering neural network, mse_loss is used to limit the weight distribution so that it tends to be 0 or 1 during the training process, thereby preventing some combinations of dark and light patterns from affecting the detection results. The formula is as follows: Among them, I is a matrix of all 1s, M is a mask weight array, m is the number of network facets, H(M) is the matrix form of the array for calculation, and the sparse coefficient calculation formula is as follows: in, Calculated by the above formula, α is a custom hyperparameter and M is the mask weight array.

10. The method according to claim 3, characterized in that The final loss function is: L=α·L agg +β·L sparse +γ·L adv Among them, L agg is the aggregation loss function; L sparse is the sparse loss function; L adv To combat the loss function, it is calculated by the detection model, and α, β, and γ are self-set hyperparameters.