Gradient calibration confrontation texture generation method and device, equipment and storage medium
By generating texture coordinate maps and foreground masks, calculating initial gradients, and performing gradient calibration and orthogonalization, the gradient sparsity and multi-view conflict problems in physical adversarial camouflage are solved, improving the stability and effectiveness of texture optimization.
Patent Information
- Application Number
- CN202510910665.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies for physical combat camouflage suffer from gradient sparsity, multi-view gradient conflicts, and redundancy issues, resulting in insufficient effectiveness and stability of texture optimization in complex environments.
By acquiring the mesh model of the target object and multiple viewpoint parameters, a texture coordinate mapping and a foreground mask are generated. The total loss value is calculated to obtain the initial gradient. The texture image is divided into optimizable and non-optimizable regions. Gradient calibration and orthogonalization are performed to generate a fused gradient until the convergence condition is met.
It significantly improves the stability and effectiveness of physical countermeasure camouflage in complex environments, overcomes gradient sparsity, conflict and redundancy issues, and ensures the continuity and consistency of texture updates.
Smart Images

Figure CN120931797A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of physical counter-camouflage, and more particularly to a gradient-calibrated method, apparatus, device, and storage medium for generating counter-textures. Background Technology
[0002] Digital adversarial attacks involve adding subtle, invisible perturbations to the purely digital domain (such as image pixels) to cause the model to misjudge the input data (e.g., misidentifying a panda as a gibbon). These attacks do not involve physical world variables; they operate only at the digital image level. The perturbations do not need to consider real-world constraints such as changes in viewpoint, material reflection, or ambient lighting. The design and execution of the attack rely entirely on the model gradient optimization at the data level.
[0003] Physical adversarial camouflage is a technique that optimizes the surface texture of objects to deceive target detection models in real-world scenarios. Its core lies in generating adversarial texture maps, causing target objects (such as vehicles) to be missed or misjudged by the model under different shooting conditions. Its applications include, but are not limited to, confidential camouflage and privacy protection. Unlike digital adversarial attacks, physical camouflage must overcome real-world constraints such as changes in viewpoint, distance differences, and lighting interference, requiring extremely high continuity and robustness in texture optimization.
[0004] Current mainstream methods optimize textures based on differentiable rendering frameworks, which generate synthetic images by simulating the camera imaging process and then update the texture using the loss gradient of the object detection model. However, existing technologies face two major drawbacks: 1. Gradient sparsity and discontinuity: In long-distance or small-view scenes, the target object covers only a small number of pixels in the image, causing most areas of the texture image to be unable to obtain effective gradients due to lack of sampling, resulting in interruption of texture updates and local degradation. 2. Multi-view gradient conflict and redundancy: To improve the universality of physical attacks, textures need to be jointly optimized from multiple perspectives. However, in multi-view joint optimization, the gradient directions generated by different camera angles are contradictory or highly repetitive. If direct averaging fusion is used, it will lead to optimization oscillation or slow convergence.
[0005] These shortcomings severely limit the effectiveness and stability of physical countermeasure camouflage in real-world complex environments, necessitating an optimization mechanism that can simultaneously address gradient sparsity, conflict, and redundancy. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this application provides a gradient-calibrated adversarial texture generation method, apparatus, device, and storage medium to improve the stability and effectiveness of attack textures under multi-view and multi-scale conditions.
[0007] The technical solution adopted by this application to solve its technical problem is: In a first aspect, this application provides a gradient-calibrated adversarial texture generation method, the method comprising: Obtain the mesh model and multiple view parameters of the target object; For each viewpoint parameter, the mesh model is rendered based on the current viewpoint parameter, generating texture coordinate mapping and foreground mask; Based on the texture coordinate mapping, the texture image to be optimized is sampled to obtain the target region image, and the target region image is combined with the preset background image through the foreground mask to form an input simulation image; The total loss value of the input simulated image is calculated using an object detection model, and the initial gradient of the texture image to be optimized is obtained based on the total loss value. Based on the initial gradient, the texture image to be optimized is divided into optimizable regions and non-optimizable regions. Gradient calibration is then performed on the non-optimizable regions based on the optimizable regions to generate a calibration gradient for the current viewpoint. The calibration gradients from all views are aggregated, sorted in ascending order of total loss value for each view, and orthogonalized on all gradients based on the sorting results to remove redundant gradient components, generating a fused gradient. The texture image to be optimized is updated based on the fused gradient. The gradient calibration and redundancy removal steps are repeated until the convergence condition is met, and the adversarial texture image is output.
[0008] Optionally, the step of rendering the mesh model based on the current viewpoint parameters for each viewpoint parameter, and generating a texture coordinate mapping and a foreground mask, includes: Based on the current viewpoint parameters, the projection positions of the surface points of the mesh model onto the two-dimensional image are calculated using a differentiable renderer to obtain the texture coordinate mapping; Based on the aforementioned viewpoint parameters, the vertices of the 3D mesh model are projected onto the 2D image plane, and the object surface depth value corresponding to each pixel position is calculated. Compare the object surface depth values of all object surface points projected to each pixel location, and retain the surface point with the smallest object surface depth value; If the surface point belongs to the target object mesh, it is determined that the corresponding pixel position is covered by the target object, and all pixels covered by the target object are determined as the foreground mask.
[0009] Optionally, the step of calculating the total loss value of the input simulated image using the object detection model and obtaining the initial gradient of the texture image to be optimized based on the total loss value includes: The input simulated image is input into the target detection model, and the target detection model calculates the classification loss value and localization loss value of the input simulated image respectively. The total loss value is determined based on the classification loss value and the localization loss value. The gradient of the total loss value with respect to the texture image to be optimized is calculated through backpropagation, and the calculation result is determined as the initial gradient.
[0010] Optionally, the step of dividing the texture image to be optimized into optimizable and non-optimizable regions based on the initial gradient, performing gradient calibration on the non-optimizable regions based on the optimizable regions, and generating a calibration gradient for the current viewpoint includes: The initial texture image is masked, and the target optimized texture region obtained after masking is divided to obtain the optimizable and non-optimizable regions of the current viewpoint parameters. For each pixel in the non-optimizable region, search for the nearest neighbor of the corresponding pixel in the optimizable region in terms of Euclidean distance, and determine whether the distance between the pixel and the corresponding nearest neighbor satisfies a preset condition; If the preset conditions are met, the gradient of the pixel is assigned the gradient of the corresponding nearest neighbor; otherwise, the gradient of the pixel is set to zero, and after traversing and calibrating all pixels, the calibration gradient containing the calibrated gradients of all pixels is generated.
[0011] Optionally, the step of searching for the nearest neighbor of each pixel in the optimizable region in terms of Euclidean distance for each pixel in the non-optimizable region includes: A K-dimensional tree data structure is constructed based on the spatial coordinates of all pixels within the optimizable region. For each pixel in the non-optimizable region, the binary tree is recursively traversed in the K-dimensional tree data structure, and irrelevant subtree regions are excluded by comparing the distance between the pixel and the dividing hyperplane. After excluding irrelevant subtree regions, the nearest neighbor of the corresponding pixel is searched in the spatial index tree based on the Euclidean distance.
[0012] Optionally, the step of summarizing the calibration gradients of all views, sorting them in ascending order according to the total loss value of each view, and performing orthogonalization on all gradients based on the sorting result to remove redundant gradient components, and generating the fused gradient includes: Initialize an empty set as the orthogonalized gradient set; Gradient vectors to be processed are selected sequentially from the sorting results. The first gradient vector to be processed is normalized and added to the orthogonalized gradient set. Then, orthogonalization calculation is performed on each gradient vector to be processed in turn to remove redundant gradient components.
[0013] Optionally, the step of sequentially performing orthogonalization calculations on each of the gradient vectors to be processed to remove redundant gradient components includes: Multiple orthogonalization coefficients are obtained by dividing the dot product of the current gradient vector and each vector in the orthogonalized gradient set by the square of the corresponding vector's magnitude. Each of the orthogonalization coefficients is multiplied by the corresponding orthogonalized vector to obtain multiple redundant components, and the sum of all the redundant components is subtracted from the current gradient vector to obtain the orthogonalized gradient vector. The orthogonalized gradient vectors are normalized and added to the orthogonalized gradient set. This process is repeated until all the gradient vectors to be processed are traversed, and then a fused gradient is generated based on the orthogonalized gradient set.
[0014] Secondly, this application provides a gradient-calibrated adversarial texture generation apparatus, comprising: The parameter acquisition module is used to acquire the mesh model and multiple view parameters of the target object; The model rendering module is used to render the mesh model based on the current viewpoint parameters for each viewpoint parameter, generating texture coordinate mapping and foreground mask; The image synthesis module is used to sample the texture image to be optimized based on the texture coordinate mapping to obtain the target region image, and to synthesize the target region image and the preset background image into an input simulated image through the foreground mask; The loss detection module is used to calculate the total loss value of the input simulated image through the object detection model, and to obtain the initial gradient of the texture image to be optimized based on the total loss value; The gradient calibration module is used to divide the texture image to be optimized into an optimizable region and an unoptimizable region based on the initial gradient, perform gradient calibration on the unoptimizable region based on the optimizable region, and generate a calibration gradient for the current viewpoint. The redundancy removal module is used to summarize the calibration gradients of all views, sort them in ascending order according to the total loss value of each view, and perform orthogonalization on all gradients based on the sorting results to remove redundant gradient components and generate fused gradients. The texture update module is used to update the texture image to be optimized based on the fused gradient, repeatedly perform gradient calibration and redundancy removal steps until the convergence condition is met, and output the adversarial texture image.
[0015] Thirdly, this application provides an electronic device, comprising: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the electronic device to perform the methods described above.
[0016] Fourthly, this application provides a computer-readable storage medium storing a program or instructions that, when executed, implement the above-described method.
[0017] The beneficial effects of this application are: by obtaining the mesh model of the target object and multiple viewpoint parameters, a foundation is laid for subsequent accurate rendering and optimization. For each viewpoint parameter, the mesh model is rendered based on the current viewpoint parameters, generating texture coordinate mapping and a foreground mask to ensure realistic image effects can be simulated under different viewpoints. Then, the texture image to be optimized is sampled through texture coordinate mapping, combined with the foreground mask to generate the input simulation image, and the total loss value of the input simulation image is calculated to obtain the initial gradient.
[0018] Based on this, the texture image to be optimized is divided into optimizable and non-optimizable regions. Gradient calibration is performed on the non-optimizable regions to generate the calibration gradient for the current viewpoint. After the calibration gradients from multiple viewpoints are aggregated, redundant gradient components are removed by sorting and orthogonalization to generate the fused gradient.
[0019] Finally, the texture image to be optimized is updated by fusion gradient, and the process is iterated until the convergence condition is met, thereby generating the final adversarial texture image. This effectively overcomes the gradient sparsity, conflict and redundancy problems in the existing technology, and significantly improves the effectiveness and stability of physical adversarial camouflage in complex environments.
[0020] In summary, this application first utilizes the nearest gradient calibration mechanism to complete the adjacent effective gradients for invisible texture regions, ensuring the spatial continuity of texture updates; then, through a loss-priority gradient decorrelation mechanism, it sorts and orthogonalizes multi-view gradients according to attack difficulty, eliminates redundant components and retains key gradient directions, thereby solving the problems of gradient discontinuity and viewpoint conflict and multi-view signal interference in physical adversarial camouflage, and ultimately significantly improving the stability and deceptiveness of adversarial textures in complex physical environments. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the gradient-calibrated adversarial texture generation method provided in the embodiments of this application; Figure 2 This is a schematic diagram illustrating the difference in image control sampling density caused by differences in viewing angle, provided in an embodiment of this application. Figure 3This is a typical schematic diagram illustrating adversarial gradient conflict caused by differences in viewpoint in the embodiments of this application; Figure 4 This is a flowchart illustrating the physical adversarial camouflage optimization method that combines nearest gradient calibration and loss-priority gradient decorrelation provided in an embodiment of this application. Figure 5 This is a schematic diagram of the virtual structure of the gradient-calibrated adversarial texture generation device provided in this application; Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0022] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0023] The following will clearly and completely describe the concept, specific structure, and resulting technical effects of this application in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of this application. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this application can be combined interactively without contradicting each other.
[0024] Reference Figure 1 , Figure 1 This is a flowchart illustrating the gradient calibration adversarial texture generation method provided in this application embodiment. The adversarial texture is a surface texture pattern of an object generated through adversarial optimization. It serves as a technical implementation carrier for physical adversarial camouflage. Its core function is to deceive the target detection model through images generated by differentiable rendering, making the target object difficult to identify or locate in the real physical environment. Figure 1 This paper demonstrates the multiple steps involved in the gradient-calibrated adversarial texture generation method described in this paper, which are explained in detail below: In step S1, the mesh model of the target object and multiple view parameters are obtained.
[0025] Specifically, the target object refers to a physical object (such as a vehicle) that requires adversarial texture optimization, and its three-dimensional geometry is composed of a mesh model. Definition: Mesh model of the target object A topological structure is a three-dimensional structure consisting of vertices, edges, and faces, used to describe the surface geometry of a target object.
[0026] In this embodiment of the application, a viewpoint parameter is introduced to simulate the observation effect of an object in a real scene. This includes variables such as the camera's position, orientation, and focal length. Viewpoint parameters. This refers to the set of physical variables used to simulate camera capture, including the camera's spatial position, orientation angle, focal length, and imaging resolution.
[0027] For example, in one specific implementation, with a vehicle as the target object, the mesh model is usually imported from CAD design files or obtained through 3D scanning reconstruction, and its triangular mesh model is loaded from the database; the viewing angle parameters can be simulated by rotating and shooting around the object with a single camera, or generated by synchronously acquiring data from multiple fixed-position cameras, for example, including 12 sets of camera poses (horizontal azimuth angle in 30° increments, pitch angle fixed at -10°), all with a focal length of 50mm, and the imaging resolution set to 1920×1080.
[0028] In step S2, for each viewpoint parameter, the mesh model is rendered based on the current viewpoint parameter to generate texture coordinate mapping and foreground mask.
[0029] The rendering step is performed using a differentiable renderer. Implement a differentiable renderer This refers to a rendering engine that supports gradient backpropagation (such as PyTorch3D), which uses GPU acceleration to achieve real-time imaging simulation; texture coordinate mapping (UV mapping) refers to the correspondence matrix between pixels in a two-dimensional texture image and points on the surface of a three-dimensional mesh model; foreground mask refers to a binary matrix that identifies the visible area of a target object in the rendered image (usually the target area is 1 and the background is 0).
[0030] Specifically, for each viewpoint parameter Through differentiable renderer Combined with the target object mesh model This allows us to obtain the texture coordinate mapping of the object onto the image. This generates viewpoint-dependent texture sampling criteria and target contour information, providing spatial constraints for subsequent image synthesis and gradient calculation. Specifically, the texture coordinate mapping ensures that the surface texture is accurately projected into the image space, while the foreground mask isolates background interference, allowing optimization to focus on the target object region.
[0031] More specifically, in the embodiments of this application, the expression for texture coordinate mapping (UV coordinate mapping) is: .
[0032] More specifically, in the embodiments of this application, the step of rendering the mesh model based on the current viewpoint parameters for each viewpoint parameter, and generating a texture coordinate mapping and a foreground mask, includes: Based on the current viewpoint parameters, the projection positions of the surface points of the mesh model onto the two-dimensional image are calculated using a differentiable renderer to obtain the texture coordinate mapping.
[0033] In this context, the vertices of the 3D mesh model refer to the coordinates of the polygon vertices that define the geometry of the target object; the 2D image plane represents the projection plane that simulates camera imaging, and its coordinate system is aligned with the pixels of the output image; the pixel position refers to the center coordinates of the pixel grid in the image plane.
[0034] Specifically, the projection calculation is performed by the differentiable renderer, which aims to map the mesh vertices in three-dimensional space to the two-dimensional imaging plane and calculate the depth value of the nearest surface point corresponding to each pixel, thereby providing a geometric basis for depth comparison and ensuring the accuracy of subsequent visibility determination.
[0035] Furthermore, based on the viewpoint parameters, the vertices of the three-dimensional mesh model are projected onto the two-dimensional image plane, and the object surface depth value corresponding to each pixel position is calculated; Compare the object surface depth values of all object surface points projected to each pixel location, and retain the surface point with the smallest object surface depth value; Among them, the object surface depth value refers to the distance from the grid surface point to the imaging plane measured along the optical axis of the camera (i.e., the distance from the object surface to the camera); the surface point refers to the candidate grid segment projected to the same pixel position.
[0036] Specifically, depth testing determines the final visible surface point of each pixel, thus resolving the self-occlusion problem and ensuring that the foreground mask accurately reflects the surface area closest to the camera. Furthermore, parallel computing optimizations can be employed, such as simultaneously processing depth comparisons of all pixels on the GPU.
[0037] For example, in one specific implementation, for the pixel coordinate (256, 512), the depth values of the three vehicle surface points projected thereto (1.2m, 1.5m, 1.4m) are compared, and the surface point with the smallest depth (1.2m) is retained.
[0038] Furthermore, if the surface point belongs to the target object mesh, it is determined that the corresponding pixel position is covered by the target object, and all pixels covered by the target object are determined as the foreground mask.
[0039] Here, the target object mesh refers to the three-dimensional geometric surface of the object to be disguised; a pixel being covered indicates that the pixel belongs to the visible area of the target object in the image.
[0040] Specifically, the depth test results are converted into a binary mask. After traversing all pixel positions, the foreground mask is obtained, thus providing accurate foreground / background segmentation for subsequent image synthesis and avoiding background interference with texture optimization. In addition, the mesh assignment of surface points needs to be verified to eliminate interference from other objects in the environment.
[0041] For example, in one specific implementation, when the surface point reserved at pixel position (256, 512) belongs to the vehicle mesh, the position is marked as 1 in the foreground mask; if it belongs to other objects (such as streetlights), it is marked as 0.
[0042] In step S3, the texture image to be optimized is sampled based on the texture coordinate mapping to obtain the target region image, and the target region image is combined with the preset background image through the foreground mask to form the input simulation image.
[0043] Among them, the texture image to be optimized This refers to the iteratively optimized vehicle surface texture map, the initial state of which can be a standard paint job or a random noise texture; the target region image refers to an image block containing only the surface pattern of the target object, generated by sampling from the texture image to be optimized according to texture coordinate mapping; and the preset background image. The simulated background image refers to the background image that simulates the real environment (such as roads and building scenes); the input simulated image refers to the complete image that is synthesized from the target area image and the background image using a foreground mask and is used to input the target detection model.
[0044] Specifically, this is performed by a differentiable rendering system, with the aim of constructing physically believable adversarial training samples to make the optimization process approximate real imaging conditions. This ensures that gradients can be backpropagated to texture pixels through differentiable sampling, while simultaneously utilizing foreground masks. Isolate background interference to ensure that optimization focuses on the effective area of the target object.
[0045] In this embodiment of the application, the texture image to be optimized is sampled based on the texture coordinate mapping. To obtain the target region image In the steps described, the expression used for the target region image is: ; Compare the obtained target area image with the background image According to the foreground mask By synthesizing the images, the final input simulation image can be obtained and fed into the object detection model. The expression for the input simulated image is: .
[0046] Additionally, it should be noted that in one possible implementation, the background image is taken from a real-world street view dataset (such as Cityscapes); in another implementation, a background environment containing lighting variations can be dynamically synthesized using a generative adversarial network to enhance the generalization ability of adversarial textures.
[0047] In step S4, the total loss value of the input simulated image is calculated using the object detection model, and the initial gradient of the texture image to be optimized is obtained based on the total loss value.
[0048] Here, the object detection model refers to a pre-trained deep learning object detector used to identify the object category and location in the input image; the initial gradient refers to the first-order partial derivative of the total loss value with respect to the texture image to be optimized; the total loss value includes the classification loss. (Error between predicted category and true label) and localization loss (The error between the predicted bounding box and the true location) is calculated using the following loss function. The representation of: ; in, For example, cross-entropy is used to measure the classification error of the model. For localization loss, such as bounding box regression error, These are the target's true category and location label, respectively.
[0049] Specifically, the degree of failure in detection is quantified through model forward propagation, and an optimization direction signal is generated through backpropagation. In simpler terms, in adversarial optimization, the optimization objective is to maximize the aforementioned loss function. To weaken the model's performance on the input image The detection performance.
[0050] More specifically, in this embodiment of the application, the step of calculating the total loss value of the input simulated image using an object detection model, and obtaining the initial gradient of the texture image to be optimized based on the total loss value, includes: The input simulated image is input into the target detection model, and the target detection model calculates the classification loss value and localization loss value of the input simulated image respectively. The total loss value is determined based on the classification loss value and the localization loss value. The gradient of the total loss value with respect to the texture image to be optimized is calculated through backpropagation, and the calculation result is determined as the initial gradient.
[0051] Specifically, taking vehicle recognition as an example, the object detection model needs to simultaneously determine the vehicle category and locate its position using the input image. When it is necessary to optimize the vehicle texture in a simulated image, the simulated image containing the texture to be optimized is first input into the object detection model. At this time, the model will calculate the classification loss and localization loss in parallel: the classification loss reflects the degree of misclassification of vehicle type (such as sedan, truck), while the localization loss quantifies the deviation between the predicted bounding box and the actual vehicle position. For example, if the vehicle texture in the simulated image is blurred, causing the model to misidentify it as background, the classification loss will increase significantly; if texture deformation causes the vehicle outline to shift, the localization loss will also increase accordingly.
[0052] More specifically, the total loss is a combination of the classification loss and the localization loss weighted according to preset criteria. This value comprehensively reflects the impact of the current texture image on the model performance. Then, the gradient of the total loss with respect to the input simulated image is derived layer by layer along the model parameter chain using the backpropagation algorithm. It is worth noting that what needs to be optimized here is not the model parameters, but rather the texture pixel values of the input simulated image itself—that is, the texture image to be optimized. The final initial gradient vector, with its direction and size, directly indicates how to adjust the texture image (such as enhancing edge sharpness or correcting illumination reflection features) to effectively reduce the model's loss value on that image.
[0053] In step S5, the texture image to be optimized is divided into an optimizable region and an unoptimizable region based on the initial gradient. Gradient calibration is then performed on the unoptimizable region based on the optimizable region to generate a calibration gradient for the current viewpoint.
[0054] Specifically, refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the difference in image spatial sampling density caused by differences in viewpoint, provided in an embodiment of this application. Figure 2 (Top) This section shows rendered images of a vehicle at close and long distances. The size of the object varies at different distances, and the number of texture sampling points required to render the image also differs accordingly. For example... Figure 2 As shown (below), sampling is denser at close range and sparser at far range. Differences in viewing distance lead to variations in image spatial sampling density, which in turn cause changes in gradient sparsity during adversarial optimization at different distances. The sparse sampling in far-range scenes results in gradients in some texture regions being sparser than those at close range, leading to discontinuities in gradients between adjacent points and consequently, discontinuous texture updates.
[0055] Therefore, in this embodiment, the texture image is divided into optimizable regions (regions visible and sampled from the current rendering viewpoint) and non-optimizable regions (regions invisible or unsampled from the current viewpoint) based on the visible range of the current rendering viewpoint. For example, if the viewpoint only covers the right side of the vehicle, the texture on the right side belongs to the optimizable region, while the texture on the left side and other unobserved parts belong to the non-optimizable region.
[0056] Subsequently, for each pixel in the non-optimizable region, a nearest neighbor search is performed to find the pixel in the optimizable region that is closest to it in texture space. If the two pixels meet a preset condition (such as a distance threshold or texture coordinate mapping relationship), the gradient value of the nearest neighbor pixel is assigned to the corresponding pixel in the non-optimizable region; otherwise, the gradient is left as zero. This operation is implemented through the Nearest Gradient Calibration (NGC) mechanism, which aims to fill in effective gradient estimates for gradient-sparse regions. For example, unsampled texture pixels on the left side of the vehicle may be filled using the nearest neighbor gradient from the visible area on the right, thereby alleviating the problem of gradient discontinuity at long distances.
[0057] Finally, by combining the true gradient of the optimizable region with the calibrated gradient of the non-optimizable region, a complete calibration gradient for the current viewpoint is generated. This gradient field, after smoothing, can more coherently guide texture optimization, ensuring consistent texture updates under different distance conditions.
[0058] More specifically, in this embodiment of the application, the step of dividing the texture image to be optimized into optimizable regions and non-optimizable regions based on the initial gradient, performing gradient calibration on the non-optimizable regions based on the optimizable regions, and generating a calibration gradient for the current viewpoint includes: The initial texture image is masked, and the target optimized texture region obtained after masking is divided to obtain the optimizable and non-optimizable regions of the current viewpoint parameters.
[0059] Specifically, removal via masking Texture image to be optimized Remove invalid regions from the original texture (such as background or non-target object parts) to obtain the target optimized texture region. Then, based on the parameters of the current rendering viewpoint, this area is divided into optimizable regions. Non-optimizable regions Optimizable area This refers to the region visible from the current viewpoint and whose gradient can be obtained through rendering sampling (e.g., the part visible on the right side of a vehicle from a specific viewpoint, which contains actual sampling points and can obtain inverse gradients), but is not an optimizable region. This refers to regions that are not visible or have not been sampled from the current viewpoint (such as the occluded left side, whose original gradient is zero).
[0060] More specifically, in the embodiments of this application, the target optimized texture region is obtained. The expression is: .
[0061] Furthermore, for each pixel in the non-optimizable region, the nearest neighbor of the corresponding pixel in the Euclidean distance is searched in the optimizable region, and it is determined whether the distance between the pixel and the corresponding nearest neighbor satisfies a preset condition. If the preset conditions are met, the gradient of the pixel is assigned the gradient of the corresponding nearest neighbor; otherwise, the gradient of the pixel is set to zero, and after traversing and calibrating all pixels, the calibration gradient containing the calibrated gradients of all pixels is generated.
[0062] Specifically, for each pixel in the non-optimizable region, its nearest neighbor in the optimizable region is searched using Euclidean distance. Euclidean distance refers to the straight-line distance between two points in a two-dimensional plane. If this distance meets a preset condition (such as being less than or equal to a preset Euclidean distance threshold), the gradient of the pixel in the non-optimizable region is assigned the gradient value of its nearest neighbor; otherwise, its gradient is set to zero. This step is repeated to finally generate a calibration gradient containing the calibrated gradients of all pixels.
[0063] More specifically, after each round of gradient backpropagation, first retain The true gradient in , and for Each pixel in ,exist Find its nearest neighbor If the following conditions are met: ; Let the gradient at that point be:
[0064] If the condition is not met, then set the gradient at that point to zero.
[0065] Furthermore, to improve the speed of the neighbor query process, this application embodiment further proposes to accelerate the query process using the KD-Tree (K-Dimensional Tree) method. Specifically, the step of searching for the nearest neighbor of each pixel in the optimizable region in terms of Euclidean distance for each pixel in the non-optimizable region includes: A K-dimensional tree data structure is constructed based on the spatial coordinates of all pixels within the optimizable region.
[0066] Specifically, a K-dimensional tree is a spatial partitioning tree that divides space into two sub-regions by alternately selecting coordinate axes. Each node represents a partitioning hyperplane, recursively constructing subtrees. This structure can efficiently support nearest neighbor search in multidimensional space.
[0067] For each pixel in the non-optimizable region, the binary tree is recursively traversed in the K-dimensional tree data structure, and irrelevant subtree regions are excluded by comparing the distance between the pixel and the dividing hyperplane.
[0068] Specifically, starting from the root node, the distance between the target pixel and the segmentation hyperplane is compared based on the segmentation dimension and segmentation value of the current node. For example, if the current node is segmented along the x-axis and the x-coordinate of the target point is greater than the segmentation value, the right subtree is searched first; otherwise, the left subtree is searched. During the traversal, if the distance between a subtree region and the target point exceeds the currently known nearest distance, the subtree can be directly excluded to avoid invalid searches. For example, if the current nearest distance is d, and the minimum possible distance (determined by the segmentation hyperplane) between all points in a subtree and the target point is greater than d, there is no need to further traverse the subtree.
[0069] After excluding irrelevant subtree regions, the nearest neighbor of the corresponding pixel is searched in the spatial index tree based on the Euclidean distance.
[0070] Specifically, in the subtrees that have not been excluded, the above process is recursively executed until a leaf node is reached. At this point, the pixel corresponding to the leaf node is the potential nearest neighbor. The nearest neighbor and its distance are updated by calculating the Euclidean distance between the target point and the candidate points.
[0071] In step S6, the calibration gradients of all views are summarized and sorted in ascending order according to the total loss value of each view. Based on the sorting result, all gradients are orthogonalized to remove redundant gradient components and generate fused gradients.
[0072] Reference Figure 3 , Figure 3 This is a typical schematic diagram illustrating adversarial gradient conflict caused by differences in viewpoint in the embodiments of this application. Figure 3 (Top) The images show vehicle images rendered under different viewpoint parameters. It can be seen that the rendered images from different viewpoints share common regions; for example, each rendered example includes the roof. Due to the existence of these shared regions, when adversarial optimization of textures is performed at different viewpoints, the gradients obtained for each example will overlap within these shared regions. As... Figure 3 As shown (below), gradients in these overlapping regions may conflict, for example in... Figure 3 In the blue dot region (bottom left), the optimized gradients have opposite directions. Therefore, it is necessary to mitigate gradient conflicts in these regions.
[0073] Specifically, firstly, the calibrated gradients after Neural Gradient Calibration (NGC) for each viewpoint are summarized to form a gradient set corresponding to all views. Then, the gradients are sorted in ascending order based on the total loss value (composed of classification loss and localization loss) for each viewpoint. Viewpoints with smaller total loss values indicate higher adversarial difficulty for the target detection model (i.e., the model is more likely to identify the target), and therefore are given higher optimization priority. After sorting, the gradients are orthogonalized sequentially to eliminate redundant components between gradients from different views.
[0074] In this embodiment of the application, the steps of summarizing the calibration gradients of all views, sorting them in ascending order according to the total loss value of each view, and performing orthogonalization on all gradients based on the sorting result to remove redundant gradient components, and generating the fused gradient include: Initialize an empty set as the orthogonalized gradient set; Gradient vectors to be processed are selected sequentially from the sorting results. The first gradient vector to be processed is normalized and added to the orthogonalized gradient set. Then, orthogonalization calculation is performed on each gradient vector to be processed in turn to remove redundant gradient components.
[0075] Specifically, the calibration gradients (i.e., gradient vectors after the most recent gradient calibration) of all viewpoints are aggregated and sorted in ascending order according to the total loss value of each viewpoint. The smaller the total loss value of a viewpoint, the higher the adversarial difficulty of the target detection model under that viewpoint (i.e., the easier it is for the model to identify the target), so its gradient direction dominance should be preserved first.
[0076] Specifically, an empty set is first initialized as the orthogonalized gradient set to store the processed gradients. Gradient vectors to be processed are selected sequentially from the sorted gradient list. The first gradient is normalized (converted to a unit vector) and added to the orthogonalized gradient set. For each subsequent gradient vector, orthogonalization is performed to remove redundant components, and each processed gradient vector is normalized and added to the orthogonalized gradient set. Finally, all orthogonalized gradient vectors are averaged to generate the fused gradient.
[0077] More specifically, the Gram-Schmidt orthogonalization method is used. For the gradient of each viewpoint, its projection component in the current direction is subtracted from the previously processed gradients, thereby preserving independent information in the gradient that is unrelated to other viewpoints. For example, if the gradient of viewpoint A has been orthogonalized, when processing the gradient of viewpoint B, its projection in the gradient direction of viewpoint A needs to be subtracted to ensure that the two are orthogonal.
[0078] More specifically, in the embodiments of this application, the step of sequentially performing orthogonalization calculations on each of the gradient vectors to be processed to remove redundant gradient components includes: Multiple orthogonalization coefficients are obtained by dividing the dot product of the current gradient vector and each vector in the orthogonalized gradient set by the square of the corresponding vector's magnitude.
[0079] Specifically, for each gradient vector to be processed, its orthogonality coefficients with all vectors in the already orthogonalized gradient set need to be calculated. Specifically, for the current gradient vector to be processed... Calculate its relationship with each vector in the orthogonalized gradient set. The dot product of the vectors, divided by the square of the magnitude of the corresponding vector: ; The dot product measures the correlation between two vectors; a larger value indicates higher redundancy. The square of the modulus is used for normalization to ensure that the orthogonality coefficients reflect the projection ratio.
[0080] Furthermore, each of the orthogonalization coefficients is multiplied by the corresponding orthogonalized vector to obtain multiple redundant components, and the sum of all the redundant components is subtracted from the current gradient vector to obtain the orthogonalized gradient vector.
[0081] Specifically, each orthogonalization coefficient Multiply by the corresponding orthogonalized vector Redundant components are obtained. Subtract the sum of all redundant components from the current gradient vector: .
[0082] Furthermore, the orthogonalized gradient vectors are normalized and added to the orthogonalized gradient set until all the gradient vectors to be processed are traversed, and a fused gradient is generated based on the orthogonalized gradient set.
[0083] Specifically, the orthogonalized vector Normalize the gradient (by dividing it by its magnitude), add it to the orthogonalized gradient set, and after traversing all view gradients, average the vectors in the orthogonalized gradient set to obtain the fused gradient: .
[0084] Finally, the average of all orthogonalized gradients is used to generate the fused gradient. By adopting the method provided in this application, the dominant gradients of high-difficulty perspectives are retained first, and multi-perspective gradient conflicts are avoided through orthogonalization, thereby improving the consistency and efficiency of cross-perspective optimization.
[0085] In step S7, the texture image to be optimized is updated based on the fused gradient, and the gradient calibration and redundancy removal steps are repeated until the convergence condition is met, and the adversarial texture image is output.
[0086] Specifically, the fusion gradient generated in step S6 is used as an optimization signal, and the texture image is updated using the gradient descent method to optimize the texture image. The update expression is: ; in, Indicates the first Texture of round iteration; The learning rate (step size) is used to control the size of the gradient step. The loss function represents the texture image to be optimized. The gradient. The entire process continues through multiple iterations until the texture image to be optimized is reached. The convergence is achieved by finding the position that maximizes the probability of false positives or false negatives in the model, ultimately yielding the adversarial texture image.
[0087] Subsequently, the gradient calibration and redundancy removal steps are repeated, i.e., steps S4 to S7 are repeated until the convergence condition is met. The convergence condition includes, but is not limited to, the change in the total loss value being less than a preset threshold, the magnitude of the fused gradient being less than a preset threshold, or the maximum number of iterations being reached.
[0088] In summary, this application proposes a physical adversarial camouflage optimization method that combines nearest gradient calibration with loss-priority gradient decorrelation, referring to... Figure 4 , Figure 4 The flowchart of the physical adversarial camouflage optimization method provided in the embodiments of this application is shown. This method combines the nearest gradient calibration and loss-priority gradient decorrelation strategy. Figure 4 (Left side) illustrates how to use a foreground mask to fuse a non-differentiable rendered background image with a differentiable rendered foreground image to generate an input image for adversarial optimization, and then feed it into a detector to obtain adversarial gradients. Figure 4 (Middle) shows the nearest neighbor gradient calibration (NGC strategy) based on KD tree in step S5 to ensure the consistency of gradient sparsity at different observation distances. Figure 4(The right side) illustrates how the LPGD strategy is used in step S6 to mitigate gradient conflicts from different perspectives: First (top), the gradients from different perspectives are sorted according to their loss magnitude; then (bottom), the sorted gradients are orthogonalized to retain high-priority gradients while mitigating the direction of conflicting gradients. This method significantly improves the stability and adversarial effectiveness of texture optimization under multi-view and multi-distance conditions by enhancing the optimization signal at both the sampling and gradient levels. This method is not only applicable to vehicle camouflage but can also be extended to adversarial modeling and stealth tasks for other physical targets, possessing broad engineering and defense application prospects.
[0089] Reference Figure 5 , Figure 5 This is a virtual structural diagram of the gradient-calibrated adversarial texture generation apparatus provided in this application. A second aspect of this application provides a gradient-calibrated adversarial texture generation apparatus, comprising: The parameter acquisition module 100 is used to acquire the mesh model of the target object and multiple view parameters; The model rendering module 200 is used to render the mesh model based on the current viewpoint parameters for each viewpoint parameter, and generate texture coordinate mapping and foreground mask; The image synthesis module 300 is used to sample the texture image to be optimized based on the texture coordinate mapping to obtain the target region image, and to synthesize the target region image and the preset background image into an input simulated image through the foreground mask; The loss detection module 400 is used to calculate the total loss value of the input simulated image through the object detection model, and obtain the initial gradient of the texture image to be optimized based on the total loss value; The gradient calibration module 500 is used to divide the texture image to be optimized into an optimizable region and an unoptimizable region based on the initial gradient, perform gradient calibration on the unoptimizable region based on the optimizable region, and generate a calibration gradient for the current viewpoint. The redundancy removal module 600 is used to summarize the calibration gradients of all views, sort them in ascending order according to the total loss value of each view, and perform orthogonalization on all gradients based on the sorting results to remove redundant gradient components and generate fused gradients. The texture update module 700 is used to update the texture image to be optimized based on the fused gradient, repeatedly perform gradient calibration and redundancy removal steps until the convergence condition is met, and output the adversarial texture image.
[0090] The gradient-calibrated adversarial texture generation device described in this application embodiment can execute the gradient-calibrated adversarial texture generation method provided in the above embodiments. The gradient-calibrated adversarial texture generation device has the corresponding functional steps and beneficial effects of the gradient-calibrated adversarial texture generation method described in the above embodiments. For details, please refer to the embodiments of the gradient-calibrated adversarial texture generation method described above. The embodiments of this application will not be repeated here.
[0091] This application also provides an electronic device, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor and a memory, which can be connected via a bus or other means. The processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the gradient calibration adversarial texture generation method in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the gradient calibration adversarial texture generation method in the above method embodiments.
[0092] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. One or more modules are stored in the memory and, when executed by the processor, perform the gradient calibration adversarial texture generation method as described in the above method embodiments. Specific details of the above electronic device can be understood by referring to the corresponding descriptions and effects in the above method embodiments, and will not be repeated here. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it may include the processes of the embodiments of the above methods. The storage medium may be a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memory.
[0093] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification. It should be noted that the foregoing embodiments are illustrative and not restrictive of this application, and that alternative embodiments may be devised by those skilled in the art without departing from the scope of the appended claims.
Claims
1. A gradient-calibrated adversarial texture generation method, characterized in that, The method includes: Obtain the mesh model and multiple view parameters of the target object; For each viewpoint parameter, the mesh model is rendered based on the current viewpoint parameter, generating texture coordinate mapping and foreground mask; Based on the texture coordinate mapping, the texture image to be optimized is sampled to obtain the target region image, and the target region image is combined with the preset background image through the foreground mask to form an input simulation image; The total loss value of the input simulated image is calculated using an object detection model, and the initial gradient of the texture image to be optimized is obtained based on the total loss value. Based on the initial gradient, the texture image to be optimized is divided into optimizable regions and non-optimizable regions. Gradient calibration is then performed on the non-optimizable regions based on the optimizable regions to generate a calibration gradient for the current viewpoint. The calibration gradients from all views are aggregated, sorted in ascending order of total loss value for each view, and orthogonalized on all gradients based on the sorting results to remove redundant gradient components, generating a fused gradient. The texture image to be optimized is updated based on the fused gradient. The gradient calibration and redundancy removal steps are repeated until the convergence condition is met, and the adversarial texture image is output.
2. The gradient-calibrated adversarial texture generation method according to claim 1, characterized in that, The step of rendering the mesh model based on the current viewpoint parameters for each viewpoint parameter, and generating texture coordinate mapping and foreground mask, includes: Based on the current viewpoint parameters, the projection positions of the surface points of the mesh model onto the two-dimensional image are calculated using a differentiable renderer to obtain the texture coordinate mapping; Based on the aforementioned viewpoint parameters, the vertices of the 3D mesh model are projected onto the 2D image plane, and the object surface depth value corresponding to each pixel position is calculated. Compare the object surface depth values of all object surface points projected to each pixel location, and retain the surface point with the smallest object surface depth value; If the surface point belongs to the target object mesh, it is determined that the corresponding pixel position is covered by the target object, and all pixels covered by the target object are determined as the foreground mask.
3. The gradient-calibrated adversarial texture generation method according to claim 1, characterized in that, The step of calculating the total loss value of the input simulated image using the object detection model, and obtaining the initial gradient of the texture image to be optimized based on the total loss value, includes: The input simulated image is input into the target detection model, and the target detection model calculates the classification loss value and localization loss value of the input simulated image respectively. The total loss value is determined based on the classification loss value and the localization loss value. The gradient of the total loss value with respect to the texture image to be optimized is calculated through backpropagation, and the calculation result is determined as the initial gradient.
4. The gradient-calibrated adversarial texture generation method according to claim 1, characterized in that, The steps of dividing the texture image to be optimized into optimizable and non-optimizable regions based on the initial gradient, performing gradient calibration on the non-optimizable regions based on the optimizable regions, and generating a calibration gradient for the current viewpoint include: The initial texture image is masked, and the target optimized texture region obtained after masking is divided to obtain the optimizable and non-optimizable regions of the current viewpoint parameters. For each pixel in the non-optimizable region, search for the nearest neighbor of the corresponding pixel in the optimizable region in terms of Euclidean distance, and determine whether the distance between the pixel and the corresponding nearest neighbor satisfies a preset condition; If the preset conditions are met, the gradient of the pixel is assigned the gradient of the corresponding nearest neighbor; otherwise, the gradient of the pixel is set to zero, and after traversing and calibrating all pixels, the calibration gradient containing the calibrated gradients of all pixels is generated.
5. The gradient-calibrated adversarial texture generation method according to claim 4, characterized in that, The step of searching for the nearest neighbor of each pixel in the optimizable region in terms of Euclidean distance for each pixel in the non-optimizable region includes: A K-dimensional tree data structure is constructed based on the spatial coordinates of all pixels within the optimizable region. For each pixel in the non-optimizable region, the binary tree is recursively traversed in the K-dimensional tree data structure, and irrelevant subtree regions are excluded by comparing the distance between the pixel and the dividing hyperplane. After excluding irrelevant subtree regions, the nearest neighbor of the corresponding pixel is searched in the spatial index tree based on the Euclidean distance.
6. The gradient-calibrated adversarial texture generation method according to claim 1, characterized in that, The steps of summarizing the calibration gradients from all views, sorting them in ascending order by the total loss value of each view, and performing orthogonalization on all gradients based on the sorting result to remove redundant gradient components, and generating the fused gradients include: Initialize an empty set as the orthogonalized gradient set; Gradient vectors to be processed are selected sequentially from the sorting results. The first gradient vector to be processed is normalized and added to the orthogonalized gradient set. Then, orthogonalization calculation is performed on each gradient vector to be processed in turn to remove redundant gradient components.
7. The gradient-calibrated adversarial texture generation method according to claim 6, characterized in that, The step of sequentially performing orthogonalization calculations on each of the gradient vectors to be processed to remove redundant gradient components includes: Multiple orthogonalization coefficients are obtained by dividing the dot product of the current gradient vector and each vector in the orthogonalized gradient set by the square of the corresponding vector's magnitude. Each of the orthogonalization coefficients is multiplied by the corresponding orthogonalized vector to obtain multiple redundant components, and the sum of all the redundant components is subtracted from the current gradient vector to obtain the orthogonalized gradient vector. The orthogonalized gradient vectors are normalized and added to the orthogonalized gradient set. This process is repeated until all the gradient vectors to be processed are traversed, and then a fused gradient is generated based on the orthogonalized gradient set.
8. A gradient-calibrated adversarial texture generation apparatus, characterized in that, include: The parameter acquisition module is used to acquire the mesh model and multiple view parameters of the target object; The model rendering module is used to render the mesh model based on the current viewpoint parameters for each viewpoint parameter, generating texture coordinate mapping and foreground mask; The image synthesis module is used to sample the texture image to be optimized based on the texture coordinate mapping to obtain the target region image, and to synthesize the target region image and the preset background image into an input simulated image through the foreground mask; The loss detection module is used to calculate the total loss value of the input simulated image through the object detection model, and to obtain the initial gradient of the texture image to be optimized based on the total loss value; The gradient calibration module is used to divide the texture image to be optimized into an optimizable region and an unoptimizable region based on the initial gradient, perform gradient calibration on the unoptimizable region based on the optimizable region, and generate a calibration gradient for the current viewpoint. The redundancy removal module is used to summarize the calibration gradients of all views, sort them in ascending order according to the total loss value of each view, and perform orthogonalization on all gradients based on the sorting results to remove redundant gradient components and generate fused gradients. The texture update module is used to update the texture image to be optimized based on the fused gradient, repeatedly perform gradient calibration and redundancy removal steps until the convergence condition is met, and output the adversarial texture image.
9. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 7.
Citation Information
Cited By
Method and device for generating confrontation sample coating of multi-camera visual perception model
CN121482536A