Double-confrontation camouflage sample generation method
Through the dual adversarial camouflage sample generation method, the global texture optimization of neural renderers and three-dimensional models is used, combined with the joint optimization of adversarial loss, smoothing loss and color loss, the problem of adversarial samples generation and visual nature in the physical world in the existing technology is solved, and the effect of hiding the human eye and the object detector is achieved.
Patent Information
- Application Number
- CN202510088662.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to generate adversarial samples in the physical world, and unlimited adversarial attacks lead to obvious contrast between visual characteristics and background of adversarial samples, which cannot effectively deceive the human eye and target detector.
Using the dual adversarial camouflage sample generation method, the global texture optimization of neural renderers and three-dimensional models, combined with joint optimization of adversarial loss, smoothing loss and color loss, is generated to generate adversarial camouflage samples that can deceive the human eye and the target detector at the same time.
The effect of hiding the human eye and the object detector simultaneously is achieved. The generated adversarial camouflage samples can be well integrated into specific scenarios, making it difficult for the human eye to distinguish, and maintain the detection effect in the object detector.
Smart Images

Figure CN120047768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for generating dual adversarial camouflage samples. Background Art
[0002] Deep Neural Network (DNN) has made great progress in various computer vision tasks such as image classification and object detection. However, DNN models have been proven to be vulnerable to adversarial attacks, where attackers can generate adversarial samples by adding specific perturbations to images, causing the DNN model to output incorrect results. The phenomenon of adversarial samples poses a major threat to the practical applications of DNN models, such as autonomous driving or face recognition systems.
[0003] According to the amplitude constraint of the perturbation, adversarial attacks can be divided into l p (p = 2, ∞) norm perturbation adversarial attacks and unrestricted perturbation adversarial attacks. In l p (p = 2, ∞) norm perturbation adversarial attacks, the generated perturbation l p (p = 2, ∞) norm cannot exceed a predefined threshold, which can ensure that the perturbation is imperceptible. Currently, there has been a large amount of work on adversarial attacks with norm perturbations, such as the Fast Gradient Sign Method, Projected Gradient Descent Attack, and Carlini & Wagner (C&W) Attack, etc. However, due to these perturbations being relatively small, the adversarial samples generated by l p (p = 2, ∞) norm perturbation adversarial attacks cannot be printed in the real physical world, which limits further applications in real scenarios such as object detection. Therefore, in physical world attacks, more unrestricted perturbation adversarial attacks are adopted, without restricting the size or area of the generated perturbation. However, unrestricted adversarial attacks will result in obvious contrast between the visual features of the adversarial samples and the background, with poor visual naturalness and being unable to effectively deceive the human eye.
[0004] To achieve a balance between the adversarial attack performance and visual naturalness of unrestricted adversarial perturbations, some methods for image classification and object detection tasks have been proposed. In the classification task, Hosseini et al. proposed semantic adversarial samples, which changed the colors of each image in the hue-saturation-lightness space and found that this led to a significant decrease in the prediction accuracy of the DNN model. Others have used pixel coordinate optimization based on spatial transformation to make the image more realistic and more likely to be misclassified by the model. In addition to changing the color or shape of the image, some other studies have also tried to change the brightness of each image. There is also a method of generating unrestricted adversarial samples by simulating natural weather, such as fog and snow. To achieve physical attacks, some people have further combined style transfer technology with adversarial attacks to make the generated adversarial samples maintain a specific style. In the object detection task, existing methods can be divided into patch-based and disguise-based methods. The patch-based method attempts to paste an adversarial patch on the object, restricting the noise to a small local patch without perturbation constraints. A patch is usually pasted on a planar object like a stop sign, in front of an object like a person, or on the background of the image. To solve the problem that existing adversarial wearable patches are too conspicuous to the human eye, some studies have also made the optimized patches themselves maintain a certain degree of naturalness. For example, natural adversarial patches have been proposed, which use the learned image manifold of a pre-trained generative adversarial network to guide the generation of adversarial patches. In contrast, the disguise-based method is achieved by modifying the target object itself, which is more challenging due to the non-planarity of 3D objects. There are two ways to obtain a disguise. One way is to optimize the required pattern, similar to the double attention suppression attack proposed by some people, restricting the perturbation to the shape of a smile to obtain a disguise that can be pasted on the surface of a car. The other way is to optimize the disguise, expecting to obtain a disguise that can wrap the entire target surface.
[0005] Although the above work focuses on balancing the adversarial attack performance and naturalness of adversarial samples, they only focus on the naturalness of adversarial samples themselves. The technology of generating adversarial samples that can deceive both the human eye and the target detector is still challenging.
[0006] Currently, there have been some progress in the research on adversarial samples for balancing adversarial performance and naturalness. However, they often pursue the naturalness of adversarial samples themselves while ignoring the effect of deceiving the human eye and are unable to deceive both the human eye and the target monitor. Therefore, the adversarial sample attacks and disguises that can deceive both the human eye and the target detector still face challenges. Summary of the Invention
[0007] To solve some or all of the above technical problems existing in the prior art, the present invention provides a method for generating double adversarial disguise samples, and the generated samples can hide from both the human eye and the target monitor at the same time.
[0008] The technical solution of the present invention is as follows:
[0009] A dual adversarial camouflage sample generation method is provided, including:
[0010] Obtain a training data set, where the training data set contains images synthesized by rendering a target image and a scene image, and the rendering of the target image includes generation by controlling camera parameters under a neural renderer tool;
[0011] Obtain an object to be camouflaged, where the object to be camouflaged includes the surface global texture of a 3D model;
[0012] Generate a rendered target image based on the global texture of the neural renderer and the 3D model;
[0013] Calculate the mean squared error loss function between the rendered target image and the scene image;
[0014] Use the mean squared error loss function to perform gradient optimization on the global texture;
[0015] Judge whether the first preset convergence condition is satisfied; if the first preset convergence condition is satisfied, output the global texture, if the first preset convergence condition is not satisfied, optimize the global texture again;
[0016] Select some textures to be further optimized in the global texture as local textures;
[0017] Use an image segmentation network model to extract the target mask in the training data set, and segment the scene image and the target object in the picture in the training data set;
[0018] Generate a rendered target image based on the current texture of the neural renderer and the 3D model;
[0019] Synthesize the rendered target image and the segmented scene image into one image through the mask to obtain a composite picture containing the current optimized texture and the background;
[0020] Calculate the difference between the current texture and the global texture as the color loss, and use the color loss function to minimize the mean squared error between the local texture and the global texture;
[0021] Calculate the printability of the composite picture as the smooth loss, and use the smooth loss function to measure the clarity of the texture;
[0022] Use the obtained composite picture as an adversarial sample, input the adversarial sample into the target detection network, and calculate the adversarial loss according to the output result of the target detection network;
[0023] Update and optimize the texture using the color loss, smooth loss, and adversarial loss as the combined loss to obtain the optimized texture;
[0024] Determine whether the optimized texture meets the second preset convergence condition; if it meets the second preset convergence condition, output the texture that meets the second preset condition as the adversarial camouflage sample of the final texture. If it does not meet the second preset convergence condition, re-fit the picture of the second object with the object in the picture of the second training dataset that is segmented, and re-judge after the update and optimization until the judgment result meets the second preset convergence condition.
[0025] Furthermore, in the above double adversarial camouflage sample generation method, calculating the fusion degree of the scene image and the rendered target image using the mean squared error loss function includes:
[0026]
[0027] where I j represents the j-th scene image selected from the training dataset, and O i represents the i-th rendered image using the texture T g . n represents the total number of final images, including the rendered target image and the scene image. The trained global texture T g can be obtained by minimizing the MSE loss between the rendered image and the scene image.
[0028] Furthermore, in the above double adversarial camouflage sample generation method, the first preset convergence condition includes:
[0029] The number of iterative optimizations for gradient optimization of the global texture using the mean squared error loss.
[0030] Furthermore, in the above double adversarial camouflage sample generation method, the number of optimization iterations includes 100 times.
[0031] Furthermore, in the above double adversarial camouflage sample generation method, the adversarial loss function includes:
[0032]
[0033] In the formula: represents the adversarial loss function, and I adv represents the adversarial camouflage picture.
[0034] Furthermore, in the above double adversarial camouflage sample generation method, the color loss function includes:
[0035]
[0036] In the formula: Denote the color loss function as T g Denote the optimized global texture as T l Denote the optimized local texture as M l Denote the mask matrix corresponding to the local texture.
[0037] Furthermore, in the above double adversarial camouflage sample generation method, the smoothing loss function includes:
[0038]
[0039] In the formula: Denote the smoothing loss function as x i,j Denote the pixel value of the adversarial camouflage image at the position (i, j) as x i+1,j Denote the pixel value of the adversarial camouflage image at the position (i + 1, j) as x i,j+ 1 Denote the pixel value of the adversarial camouflage image at the position (i, j + 1) as x.
[0040] The main advantages of the technical solution of the present invention are as follows:
[0041] The double adversarial camouflage sample generation method of the present invention enables the generated samples to hide from the human eye and / or the target monitor simultaneously through two stages. Specifically, in the first stage, given a training data set, several background images are selected as scene images, and the global texture is trained to make the rendered image close to the scene image. After the first stage, the rendered object can be well integrated into a specific scene, making it difficult for the human eye to distinguish, thus achieving the purpose of hiding from the human eye. In the second stage, local textures are selected from the optimized global texture, and the objective function is used to detect the adversarial samples, so that the final texture is a combination of the global texture and the local texture. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0043] Figure 1 FIG. is a schematic flow chart of the effect of camouflaging a human model by using the double adversarial camouflage sample generation method provided by an embodiment of the present invention;
[0044] Figure 2 FIG. is a schematic diagram of rendering a human model by using the double adversarial camouflage sample generation method provided by an embodiment of the present invention;
[0045] Figure 3 FIG. is a schematic diagram of the principle of the double adversarial camouflage sample generation method provided by an embodiment of the present invention;
[0046] Figure 4 In the dual adversarial camouflage sample generation method according to an embodiment of the present invention, the first object is a human model, and the scene image is a schematic diagram of an adversarial camouflage sample for a snow scene;
[0047] Figure 5 In the dual adversarial camouflage sample generation method according to an embodiment of the present invention, the first object is a human model, and the scene image is a schematic diagram of an adversarial camouflage sample for a forest scene;
[0048] Figure 6 In the dual adversarial camouflage sample generation method according to an embodiment of the present invention, the first object is a human model, and the scene image is a schematic diagram of an adversarial camouflage sample for a desert scene.
[0049] Figure 7 It is a schematic flow chart of the dual adversarial camouflage sample generation method according to an embodiment of the present invention. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0051] The following combines the attached Figures 1-7 , and details the technical solutions provided by the embodiments of the present invention.
[0052] To solve the problems in the prior art, the present invention proposes a dual adversarial camouflage generation method. As Figures 1-3 shown, the present invention regards the adversarial camouflage generation as an optimization problem of the texture that can be drawn on the surface of an object, and designs a two-stage training method based on a neural renderer. In the first stage, several background images are selected as the scene images, and the global texture is trained to make the rendered image close to the scene image. After the first stage ends, the rendered object can be well integrated into a specific scene, making it difficult for the human eye to distinguish. In the second stage, local textures are selected from the optimized global texture, and three loss functions are designed to optimize the local textures. These loss functions include adversarial loss, smooth loss, and color loss. The final texture is a combination of the global texture and the local texture, which can deceive both the human vision and the target detector at the same time.
[0053] As shown in the attached Figures 1-3 and Figure 7 shown, the embodiments of the present invention provide a dual adversarial camouflage sample generation method, which includes:
[0054] Step S1: Obtain a training dataset, which includes images synthesized by rendering a target image and a scene image, where the rendered target image is generated by controlling camera parameters under a neural renderer tool;
[0055] The scene image can be any scene. In combination with actual applications, scene graphics include but are not limited to snowfields, forests, and deserts. In addition, the scene graph line can also be set as a city, etc. During application, the scene image is an image, photo, or picture taken of an actual scene.
[0056] Step S2: Obtain the object to be camouflaged, which includes the global surface texture of a 3D model;
[0057] The object to be camouflaged includes any object in the scene image, and this object includes moving objects or stationary objects. In this embodiment, it is preferably aimed at stationary objects, but it can also be aimed at moving objects. In combination with actual applications, the object to be camouflaged includes but is not limited to houses, vehicles, simple tents and other living equipment, or the human body, etc.
[0058] Step S3: Generate a rendered target image based on the neural renderer and the global texture of the 3D model;
[0059] Specifically, generating a rendered target image based on the neural renderer and the global texture of the 3D model includes overlaying the texture in the scene image onto the object to be camouflaged, so that the global texture of the object to be camouflaged is basically the same as the texture of the scene image, making the overall texture of the target object deeply integrated with the surrounding environment, obtaining a first-level adversarial camouflage sample, so as to achieve the effect of deceiving the human eye.
[0060] In order to enable the object to be camouflaged to better achieve the effect of deceiving the human eye in the scene image and not be recognized and discovered by the human eye, for stationary objects to be camouflaged, it is preferably that the texture of the obtained scene image is the texture of the object to be camouflaged at the same position in the scene image. During implementation, the texture of the scene image at the same position is overlaid on the stationary object to be camouflaged, so that the object to be camouflaged achieves the effect of confusing the human eye and not being recognized by the human eye; for moving objects to be camouflaged, the texture of the selected scene image preferably contains the same or approximately the same number of elements as the elements in the scene image, and is arranged according to the arrangement method of the elements in the scene image, so that the moving object is not easily discovered by the human eye at any position in the scene, such as the camouflage method of mountain camouflage (clothes) and mountain scene images.
[0061] Step S4: Calculate the mean square error loss function between the rendered target image and the scene image;
[0062] In the present invention, the purpose of the mean squared error loss function is to calculate the degree of fusion between the scene image and the rendered target image. Among them, the mean squared error loss function includes:
[0063]
[0064] Among them, represents the mean squared error loss function, I j represents the j-th scene image selected from the training dataset, O i represents the i-th rendered image using the texture T g The total number of final images, including the rendered target image and the scene image, can obtain the trained global texture T by minimizing the MSE loss between the rendered image and the scene image g , where i and j represent the image order in the dataset.
[0065] Step S5: Perform gradient optimization on the global texture using the mean squared error loss;
[0066] The neural renderer is a tool that converts 3D textures into 2D images. Given 3D texture parameters and setting different camera parameters, the neural renderer can generate 2D target images from different perspectives. Using the Mean Squared Error (MSE) loss function can measure the difference between the generated 2D target image and the scene image. This step aims to optimize the full-body texture parameters of the 3D model so that the MSE is as low as possible, thereby achieving the purpose of deep fusion between the target image and the scene image. The specific implementation steps of performing gradient optimization on the global texture using the mean squared error loss include steps S51 - S56:
[0067] Step S51: Determine the optimization objective: Define an objective function, which is usually based on the mean squared error loss and measures the difference between the generated 2D target image and the scene image.
[0068] Step S52: Initialize the parameters: Select a suitable parameterization method to initialize the texture parameters.
[0069] Step S53: Calculate the loss function. The mean squared error loss function is defined as follows:
[0070]
[0071] Among them, represents the mean squared error loss function, I j represents the j-th scene image selected from the training dataset, O i represents the i-th rendered image using the texture T gThe i-th rendered image, where n represents the total number of final images, including the target image and the scene image to be rendered, and the trained global texture T can be obtained by minimizing the MSE loss between the rendered image and the scene image g 。
[0072] Step S54: Calculate the gradient: Calculate the gradient of the loss function with respect to the texture parameters. Automatic differentiation tools such as TensorFlow or PyTorch can be used for gradient calculation.
[0073] Step S55: Update the parameters: Use gradient descent or other optimization algorithms to update the parameter T g 。
[0074] The update rule is as follows:
[0075] where T gnew represents the updated parameter value, T gold represents the parameter value before update, represents the gradient of the loss function L(T g ) with respect to the parameter T g , and the gradient is a vector, each component of which represents the partial derivative of the loss function L(T g ) with respect to the parameter T g i. The direction of the gradient points to the direction where the loss function grows fastest. Therefore, in gradient descent, the parameters are updated along the opposite direction of the gradient to reduce the value of the loss function, and α represents the learning rate.
[0076] Step S56: Iterative optimization: Repeat steps S53 to S55 until the stopping condition is met, such as reaching a preset number of iterations, the value of the loss function is lower than a certain threshold, or the gradient change is very small.
[0077] It should be noted that in actual operation, the above specific implementation process needs to be adjusted according to the actual application scenario and requirements. The following factors also need to be considered during the implementation: Regularization: To prevent overfitting, a regularization term may need to be added to the loss function. Efficiency: Global texture unfolding may involve a large amount of computation, and the optimization algorithm needs to be executed efficiently.
[0078] Step S6: Determine whether the first preset convergence condition is satisfied; if the first preset convergence condition is satisfied, output the global texture, and if the first preset convergence condition is not satisfied, optimize the global texture again;
[0079] In this embodiment, the first preset convergence condition includes: the number of iterative optimizations for global texture unfolding gradient optimization using the mean square error loss.
[0080] As some optional implementation methods of this embodiment, the above-mentioned optimization includes 100 times here, but it should be noted that the above-mentioned 100 times of optimization are only for exemplary purposes and are not the only limitation in the present invention. The specific number of times can be determined according to actual applications.
[0081] This setting enables the output global texture to confuse the human eye and not be discovered by the human eye.
[0082] Specifically, the adversarial loss function includes:
[0083]
[0084] Where: represents the adversarial loss function, I adv It represents the anti-camouflage picture. Represents the process of calculating the target confidence.
[0085] Step S7: Selecting a portion of the texture that needs to be further optimized from the global texture as a local texture;
[0086] Specifically, the selected local texture may be any texture or any local image in the global texture.
[0087] As an example, the object to be camouflaged is a human body, and the scene image is a forest. After performing steps S1 to S6, the global texture of the human body in the forest environment is obtained, so that the human body with forest texture is not easily recognized and discovered by the human eye in its corresponding forest environment, but is easily discovered by the target detector. Therefore, any part of the human body that is confusing the human eye, difficult or impossible to be discovered by the human eye, but can be discovered by the target detector is selected as a local image for further camouflage.
[0088] Step S8: extracting the target mask in the training data set using the image segmentation network model, and segmenting the scene image in the training data set and the target object in the picture;
[0089] Specifically, the above-mentioned scene images and target objects in the pictures segmented out of the training data set include scene images and target images. As an example, a photo of a pedestrian taken in a forest, the image area after removing the person is the scene image, and the image of the person after removing the background such as the forest is the target image.
[0090] In some optional implementations of this embodiment, the image segmentation network model includes one or more of Segment Anything Model (SAM), DINOv2, Mask2Former and Swin Transformer.
[0091] Step S9: Generate a rendered target image based on the neural renderer and the current texture of the 3D model;
[0092] The current texture includes the above-mentioned textures selected from the global texture that need to be further optimized, that is, the local texture.
[0093] Step S10: Synthesize the rendered target image and the segmented scene image into one image through a mask to obtain a composite picture containing the current optimized texture and the background;
[0094] Specifically, the composite picture is the final image obtained by superimposing the current optimized texture on the target to form a target image and further fusing it with the scene image.
[0095] Step S11: Calculate the difference between the current texture and the global texture as the color loss, and use the color loss function to minimize the mean square error between the local texture and the global texture;
[0096] Specifically, the color loss function includes:
[0097]
[0098] In the formula: represents the color loss function, T g represents the optimized global texture, T l represents the optimized local texture, M l represents the mask matrix corresponding to the local texture.
[0099] Step S12: Calculate the printability of the composite picture as the smooth loss, and use the smooth loss function to measure the clarity of the texture;
[0100] Specifically, the smooth loss function includes:
[0101]
[0102] In the formula: represents the smooth loss function, x i,j represents the pixel value of the adversarial camouflage picture at the (i, j) position, x i+1,j represents the pixel value of the adversarial camouflage picture at the (i + 1, j) position, x i,j+ 1 represents the pixel value of the adversarial camouflage picture at the (i, j + 1) position.
[0103] Step S13: Use the obtained composite picture as an adversarial sample, input the adversarial sample into the target detection network, and calculate the adversarial loss according to the output result of the target detection network;
[0104] The object detection network is any algorithm and model for object detection in the prior art, such as R-CNN, etc.
[0105] Step S14: Use the color loss, smooth loss, and adversarial loss as the combined loss to update and optimize the texture, and obtain the optimized texture.
[0106] The combined loss is the sum of the above-mentioned color loss, smooth loss, and adversarial loss.
[0107] Step S15: Determine whether the optimized texture meets the second preset convergence condition, such as reaching a certain number of iteration times. If it meets the second preset convergence condition, output the texture that meets the second preset condition as the adversarial camouflage sample of the final texture. If it does not meet the second preset convergence condition, continue to optimize the texture until the judgment result meets the second preset convergence condition.
[0108] The second preset convergence condition includes the number of iteration optimizations for unfolding gradient optimization.
[0109] The principle of the dual adversarial camouflage sample generation method of the present invention includes:
[0110] The generated samples can deceive both biological vision and machine vision in two stages. Specifically, the first stage is to optimize the global texture of the three-dimensional target model so that the overall texture of the target object is deeply integrated with the surrounding environment to obtain the first-layer adversarial camouflage sample, thus achieving the effect of deceiving the human eye. The second stage is to select local textures in the first-layer adversarial camouflage sample for optimization. The optimized texture will cover the global texture of the first stage at the corresponding positions. When optimizing, consider the combined loss function of the adversarial loss, smooth loss, and color loss to obtain the second-layer adversarial camouflage sample, so as to achieve the purpose of deceiving both the human eye and the target detector.
[0111] Through the method of the present invention, the effect of deceiving both biological vision and machine vision can be achieved. Biological vision is the human eye, and machine vision is detection equipment such as target detectors.
[0112] In some optional implementation scenarios of this embodiment, such as Figure 4 shown, in Figure 4Set the first object as a human model, set the scene image as a snow scene, and use the dual adversarial camouflage sample generation method of the present invention to hide the human model. After performing the above steps S1 - S6, the global texture of the human body in the snow environment is obtained, making the human body with snow texture not easily recognizable and discoverable by the human eye in its corresponding snow environment, but easily discoverable by the target detector during detection. Therefore, on the human body that achieves the effect of confusing the human eye, being not easily or not at all discoverable by the human eye but being discoverable by the target detector, any part is selected as a local image for further camouflage. When using the selected arbitrary part as a local image for camouflage, after the above steps S7 - S15, the current optimized texture is superimposed on the target to form a target image, and further fused with the snow image to obtain the final image. By optimizing the global texture, the overall texture of the human body is deeply fused with the surrounding snow environment, obtaining the first - stage adversarial camouflage sample, thus achieving the effect of deceiving the human eye; and by selecting and optimizing local textures in the adversarial camouflage sample, the optimized texture will cover the global texture in the corresponding position in the first stage, thus achieving the purpose of deceiving both the human eye and the target detector simultaneously. It deceives both biological vision and machine vision, that is, neither the human eye nor the target detector can recognize it.
[0113] In some alternative implementation scenarios of this embodiment, such as Figure 5 shown, in Figure 5 Set the first object as a human model, set the scene image as a forest scene, and use the dual adversarial camouflage sample generation method of the present invention to hide the human model. After performing the above steps S1 - S6, the global texture of the human body in the forest environment is obtained, making the human body with forest texture not easily recognizable and discoverable by the human eye in its corresponding forest environment, but easily discoverable by the target detector during detection. Therefore, on the human body that achieves the effect of confusing the human eye, being not easily or not at all discoverable by the human eye but being discoverable by the target detector, any part is selected as a local image for further camouflage. When using the selected arbitrary part as a local image for camouflage, after the above steps S7 - S15, the current optimized texture is superimposed on the target to form a target image, and further fused with the scene image to obtain the final image. By optimizing the global texture, the overall texture of the human body is deeply fused with the surrounding forest environment, obtaining the first - stage adversarial camouflage sample, thus achieving the effect of deceiving the human eye; and by selecting and optimizing local textures in the adversarial camouflage sample, the optimized texture will cover the global texture in the corresponding position in the first stage, thus achieving the purpose of deceiving both the human eye and the target detector simultaneously.
[0114] In some alternative implementation scenarios of this embodiment, such as Figure 6 shown, in Figure 6Set the first object as a mannequin, set the scene image as a desert scene, and use the dual adversarial camouflage sample generation method of the present invention to hide the mannequin. After performing the above steps S1 - step S6, the global texture of the human body in the desert environment is obtained, so that the human body with desert texture is not easily recognizable and discoverable by the human eye in its corresponding desert environment, but is easily discoverable by the target detector. Therefore, any part is selected from the above-mentioned human body that confuses the human eye, is not easily or cannot be discovered by the human eye, but can be discovered by the target detector as a local image for further camouflage. When using the arbitrarily selected part as a local image for camouflage, after the above steps S7 - step S15, the current optimized texture is superimposed on the target to form a target image, and further fused with the desert image to obtain the final image. By optimizing the global texture, the overall texture of the human body is deeply fused with the surrounding desert environment, obtaining the first - stage adversarial camouflage sample, thus achieving the effect of deceiving the human eye; and by selecting and optimizing the local texture in the adversarial camouflage sample, the optimized texture will cover the global texture in the corresponding position in the first stage, thus achieving the purpose of deceiving both the human eye and the target detector at the same time.
[0115] In some alternative implementation manners of this embodiment, a single scene image can be covered on the target image, or multiple scene images can be covered on the target image. When covering multiple images on the target image, the multiple scenes are pre - numbered, and then the target image is used to generate adversarial samples according to the above method according to the numbers of the multiple scene images.
[0116] In order to make the generated adversarial samples better camouflaged, when processing the images of multiple scenes using the above method, each image is operated by the above - mentioned adversarial method multiple times.
[0117] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. In addition, "front", "rear", "left", "right", "up" and "down" in this article are all referenced with the placement state shown in the drawings.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating double adversarial camouflage samples, characterized in that: include: Acquire a training data set, wherein the training data set includes an image synthesized by rendering a target image and a scene image, wherein the rendering target image includes an image generated by controlling camera parameters under a neural renderer tool; Acquire an object to be camouflaged, wherein the object to be camouflaged includes a global surface texture of a three-dimensional model; Generate a render target image based on the neural renderer and the global texture of the 3D model; Calculate the mean square error loss function between the rendered target image and the scene image; Use the mean square error loss function to perform gradient optimization on the global texture; Determine whether a first preset convergence condition is met; if the first preset convergence condition is met, output the global texture; if the first preset convergence condition is not met, re-optimize the global texture; Select some textures that need to be further optimized from the global texture as local textures; Extracting the target mask in the training data set by using the image segmentation network model, and segmenting the scene image in the training data set and the target object in the picture; generating a render target image based on the neural renderer and the current texture of the three-dimensional model; The rendered target image and the segmented scene image are synthesized into one image through a mask to obtain a composite image containing the current optimized texture and background; Calculate the difference between the current texture and the global texture as the color loss, and use the color loss function to minimize the mean square error between the local texture and the global texture; Calculate the printability of the synthesized image as a smoothness loss, and use the smoothness loss function to measure the clarity of the texture; The obtained synthetic image is used as an adversarial sample, the adversarial sample is input into the target detection network, and the adversarial loss is calculated based on the output result of the target detection network; Use color loss, smoothness loss and adversarial loss as a joint loss to update and optimize the texture to obtain the optimized texture; Determine whether the optimized texture satisfies a second preset convergence condition; if so, output the texture satisfying the second preset condition as an adversarial camouflage sample of the final texture; if not, re-fit the image of the second object with the object in the image segmented from the second training data set, and re-judge after updating and optimizing until the judgment result satisfies the second preset convergence condition.
2. The method for generating dual adversarial camouflage samples according to claim 1, characterized in that: The mean square error loss function includes: in, represents the mean square error loss function, I j represents the jth scene image selected from the training dataset, O i Represents the use of texture T g The i-th rendered image of n represents the total number of final images, including the rendered target image and the scene image. The trained global texture T can be obtained by minimizing the MSE loss between the rendered image and the scene image. g .
3. The method for generating dual adversarial camouflage samples according to claim 1, characterized in that: The first preset convergence condition includes: The number of iterations of gradient optimization for global texture using mean square error loss.
4. The method for generating dual adversarial camouflage samples according to claim 3, characterized in that: The number of optimization iterations includes 100 times.
5. The method for generating dual adversarial camouflage samples according to claim 1, characterized in that: The adversarial loss function includes: Where: represents the adversarial loss function, I adv Represents an adversarial camouflage image.
6. The method for generating dual adversarial camouflage samples according to claim 1, characterized in that: The color loss function includes: Where: represents the color loss function, T g represents the optimized global texture, T l represents the optimized local texture, M l Represents the mask matrix corresponding to the local texture.
7. The method for generating dual adversarial camouflage samples according to claim 1, characterized in that: The smoothing loss function includes: Where: represents the smooth loss function, x i,j represents the pixel value of the adversarial camouflage image at position (i, j), x i+1,j represents the pixel value of the adversarial camouflage image at position (i+1, j), x i,j+1 Represents the pixel value of the adversarial camouflage image at position (i, j+1).