Adversarial patch generation method and device for image-point cloud fusion perception model
By generating perceptual heat maps in the image-point cloud fusion perception model and determining the perception type, selecting the appropriate adversarial patch deployment location, and generating adversarial samples on the image modality, solving the problems of high cost of adversarial sample generation and difficult deployment in the prior art, and achieving efficient and low-cost adversarial sample generation.
Patent Information
- Application Number
- CN202510654957.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The prior art requires perturbing point cloud data when generating adversarial samples for image-point cloud fusion perception models, resulting in high cost and difficult deployment.
By acquiring images and corresponding point cloud data, a perceived thermal map is generated, and the perceived type of neural network model is determined based on the similarity between the perceived thermal map and the distribution of actual objects. Then, the deployment location of the adversarial patch is selected based on the perceptual type, and adversarial samples are generated on the image modality to avoid perturbation to the point cloud data.
A confrontational sample generation method that reduces costs and simplifies deployment is realized, and only generates adversarial samples on the image modality is overcome, and the problem of insufficient effectiveness of single-modal adversarial samples is improved.
Smart Images

Figure CN120182772A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence deep learning, and particularly to a method and device for generating adversarial patches for an image-point cloud fusion perception model. Background Art
[0002] Deep neural examples (DNN) models have been widely used in perception tasks such as object detection due to their excellent performance. However, a large number of studies have shown that DNN models are vulnerable to adversarial attacks, which create adversarial examples (AEs) by carefully manipulating the input of the DNN model, thereby forcing the model to make incorrect decisions. The abnormal behavior of the model may bring huge harmful consequences, especially in safety-critical fields such as autonomous driving and intelligent monitoring. This concept has further evolved into physical adversarial examples, demonstrating that adversarial example attacks can be transferred to the physical world and implemented under real environmental conditions. For example, physical adversarial examples can have an impact even when the camera captures images from different viewing distances or angles. Physical adversarial examples increase the threat level of adversarial attacks, making them a huge and practical threat to DNN models.
[0003] Current relatively advanced DNN perception models usually perform inference by combining different modality data such as images and point clouds. Recent research has shown that point cloud modality data plays a more important role in fusion perception. Therefore, for such models, the generation of adversarial examples usually focuses on the perturbation of point cloud data. For image-point cloud fusion perception models, some common strategies for implementing physical adversarial examples are to use a laser diode to perturb point cloud data, or to use a 3D printer to print an adversarial object to perturb both the image and point cloud data simultaneously. However, the perturbation methods based on devices such as laser diodes or 3D printers all require perturbing the point cloud data, which is costly and difficult to deploy. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an adversarial patch generation method, a computer device, a computer-readable storage medium, and a computer program product for an image-point cloud fusion perception model that can reduce costs and is easy to deploy.
[0005] In a first aspect, this application provides an adversarial patch generation method for an image-point cloud fusion perception model, including: Obtain a first image, randomly initialize a mask and a global perturbation according to the first image, and superimpose the global perturbation on the first image according to the mask to obtain a second image; Obtain the point cloud data corresponding to the first image, input the second image and the point cloud data into a neural network model for training, and obtain the perception heat map of the neural network model; Determine the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution; wherein, the perception types of the neural network model include a global perception model and a local perception model; Determine the deployment position of the adversarial patch according to the perception type, and model the adversarial patch based on the deployment position to obtain a third image; wherein, the deployment position of the global perception model includes the environmental background in the first image, and the deployment position of the local perception model includes the target object in the first image.
[0006] In one embodiment, inputting the second image and the point cloud data into a neural network model for training to obtain the perception heat map of the neural network model includes: Obtain the sum of the detection box confidence levels and the mask area generated by processing through the neural network model; Taking the minimization of the sum of the detection box confidence levels and the mask area as the optimization objective, and using the gradient descent algorithm to perform gradient optimization on the mask and the global perturbation to obtain an optimized mask; Use the optimized mask as the perception heat map of the neural network model.
[0007] In one embodiment, taking the minimization of the sum of the detection box confidence levels and the mask area as the optimization objective, and using the gradient descent algorithm to perform gradient optimization on the mask and the global perturbation to obtain an optimized mask includes: Input the sum of the detection box confidence levels into the mean square error function to obtain the mask loss; Input the mask area into the mean square error function to obtain the adversarial loss; Add the mask loss and the adversarial loss to obtain a sum value, take the minimization of the sum value as the optimization objective, and use the gradient descent algorithm to perform gradient optimization on the mask and the global perturbation to obtain the perception sensitive area of the neural network model; Determine the optimized mask based on the sensitive area.
[0008] In one embodiment, randomly initialize the mask according to the first image, including: Set hyperparameters γ and s, and set variables m and global perturbation p; Generate the mask according to the following formula: ; Wherein, M is the mask, [i, j] is the pixel position, and tanh is the hyperbolic tangent function.
[0009] In one embodiment, determining the perception type of the neural network model according to the similarity between the perceived heat map and the actual object distribution includes: Inputting the first image and the point cloud data into the neural network model for training, and generating an object detection frame on the first image; Setting the area enclosed by the object detection frame to 1 and the rest to 0 to obtain a binary matrix with the same size as the first image; Judging whether the perceived heat map and the binary matrix satisfy the following conditions: ; Wherein, x is the first image, M x is the perceived heat map generated for the first image, is the binary matrix, and E x is the mean value of the expressions obtained from the perceived heat maps generated by different first images, β is a hyperparameter, ⊙ is the Hadamard product operator, and ∑ is the summation operator; If the perceived heat map and the binary matrix satisfy the above conditions, it is determined that the perceived heat map is similar to the actual object distribution, and the perception type of the neural network model is determined to be a local perception model; If the perceived heat map and the binary matrix do not satisfy the above conditions, it is determined that the perceived heat map is not similar to the actual object distribution, and the perception type of the neural network model is determined to be a global perception model.
[0010] In one embodiment, modeling the adversarial patch based on the deployment position to obtain a third image includes: Determining the position parameters of the adversarial patch; According to the position parameters, projecting the adversarial patch into the pixel coordinate system of the first image according to the deployment position to obtain a pixel perturbation; Superimposing the pixel perturbation on the first image to obtain the third image.
[0011] In one embodiment, projecting the adversarial patch into the pixel coordinate system according to the deployment position to obtain a pixel perturbation, including: Obtaining the device parameters and the physical position parameters of the adversarial patch; wherein, the device parameters include the external parameters of the lidar and the internal parameters of the camera; Determining a mapping relationship according to the device parameters, substituting the physical position parameters of the adversarial patch into the mapping relationship, and solving to obtain pixel coordinates; Calculate the pixel value corresponding to the pixel coordinates using the bilinear interpolation algorithm to obtain the pixel perturbation.
[0012] In one embodiment, after superimposing the pixel perturbation on the first image to obtain the third image, the method further includes: Perform data random augmentation on the adversarial patch, and superimpose the adversarial patch after the data random augmentation on the first image to obtain a fourth image; Use multiple of the fourth images as samples and input them into the neural network model to calculate the average value of the neural network loss; Taking the minimization of the average value of the neural network loss as the optimization objective, use the gradient descent algorithm to perform gradient optimization on the pixel values of the adversarial patch, and use the optimized adversarial patch as the new adversarial patch.
[0013] In one embodiment, Performing data random augmentation on the adversarial patch includes at least one of the following: Randomly transform the size and physical position parameters of the adversarial patch; Randomly transform the contrast, saturation, and brightness of the adversarial patch; During the training process of the neural network model, superimpose pixel-level random perturbations on the adversarial patch; During the training process of the neural network model, superimpose pixel-level random perturbations on the adversarial patch, and add random noise to the injected random perturbations.
[0014] In a second aspect, the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect above are implemented.
[0015] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0016] In a fourth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0017] The above-mentioned method for generating adversarial patches for an image-point cloud fusion perception model, computer device, computer-readable storage medium, and computer program product perturb the original image and generate a corresponding perception heat map. According to the similarity between the perception heat map and the actual object distribution, the perception type of the neural network model is determined. Different adversarial patch deployment schemes are selected based on the perception type of the neural network model, and the original image is perturbed using the adversarial patches. Compared with the related technology, there is no need to perturb the point cloud data, and adversarial samples are only generated in the image modality, with simple deployment and low cost. Moreover, by making full use of the differences in the structure of the image-point cloud perception model, the problem of insufficient effectiveness of single-modal adversarial samples is overcome, which is beneficial to improving the adversarial effect of adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 FIG. is a block diagram of the hardware structure of a terminal for the method for generating adversarial patches for an image-point cloud fusion perception model in an embodiment; Figure 2 FIG. is a flowchart of the method for generating adversarial patches for an image-point cloud fusion perception model in an embodiment; Figure 3 FIG. is a schematic flowchart of generating a perception heat map for a single first image in an embodiment; Figure 4 FIG. is a flowchart of the method for determining the perception type of a neural network model in an embodiment; Figure 5 FIG. is a flowchart of modeling adversarial patches in an embodiment; Figure 6 FIG. is a schematic flowchart of deploying adversarial patches on the ground in an embodiment; Figure 7 FIG. is a flowchart of the method for generating adversarial patches for an image-point cloud fusion perception model in another embodiment; Figure 8 FIG. is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the general meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "containing", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled", etc. involved in this application do not limit to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application only distinguish similar objects and do not represent a specific order for the objects.
[0021] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 is a hardware structure block diagram of a terminal for an adversarial patch generation method for an image-point cloud fusion perception model according to an embodiment of this application. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 101 and a memory 102 for storing data. Among them, the processor 101 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA. The above terminal may further include a transmission device 103 for communication functions and an input / output device 104. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown.
[0022] The memory 102 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the adversarial patch generation method for the image-point cloud fusion perception model in this embodiment. The processor 101 executes various functional applications and data processing by running the computer program stored in the memory 102, that is, implements the above method. The memory 102 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 102 may further include a memory remotely disposed relative to the processor 101, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0023] The transmission device 103 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 103 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 103 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] In the related art, the perturbation methods based on devices such as laser diodes or 3D printers have the following specific limitations: (1) Since it is necessary to perturb the point cloud data, professional equipment support is required, the cost is relatively high, and the deployment difficulty is relatively large.
[0025] (2) These methods often only target certain specific objects (for example, making a target detection system unable to detect a certain vehicle), and the range of objects affected by the adversarial samples is relatively limited.
[0026] (3) These methods need to access the structure and parameters of the DNN model, but often only use an end-to-end method to optimize the adversarial samples, without a differential adversarial sample generation strategy for a specific model, resulting in poor effects of the adversarial samples.
[0027] Based on the analysis of the above situation, in one embodiment, Figure 2 A flowchart of an adversarial patch generation method for an image-point cloud fusion perception model is provided. This method can run on Figure 1 the terminal. For ease of understanding, Figure 3 A schematic flowchart of generating a perception heat map for a single first image in this embodiment is provided. Combining Figure 2 and Figure 3 , this method includes the following steps: Step S101: Obtain a first image, randomly initialize a mask and a global perturbation according to the first image, and superimpose the global perturbation on the first image according to the mask to obtain a second image.
[0028] The first image is the original image for which an adversarial sample is to be generated, that is, the Figure 3 unperturbed image in. The mask and the global perturbation are the same in shape as the first image. Specifically, an image, a mask, and a global perturbation can all be represented by multi-dimensional arrays in a computer. An image is a multi-dimensional array of (3, h, w), a mask is a multi-dimensional array of (1, h, w), and a global perturbation is a multi-dimensional array of (3, h, w). Where h is the height and w is the width. Since the image and the perturbation have color information and have three RGB channels, the first digit of the array is 3. The mask (Mask) and the global perturbation (Global Perturbation) are consistent with the first image in terms of the number of dimensions, spatial size, and shape structure - all three are three-dimensional arrays. Among them, the number of channels of the mask is 1, the number of channels of the image and the perturbation is 3, and their spatial dimensions (height h and width w) are the same, that is, the shapes are (1, h, w), (3, h, w), and (3, h, w) respectively, meeting the condition of element-wise operation or superposition at the same spatial position.
[0029] The channel value of the mask (abbreviated as the mask value) can be 0, 1, or a value between 0 and 1. The channel value of the mask is used to determine whether the pixel at the corresponding position will be processed or in what way it will be processed.
[0030] Perturbation refers to making small and almost imperceptible modifications to the input data to deceive a machine learning model and make it produce incorrect classification results. For example, in image recognition, changing the color or brightness of some pixels; in point cloud data, slightly moving the position of points or changing attributes such as their reflection intensity. In this embodiment, the global perturbation refers to changing or interfering with the entire first image. The perturbation methods include but are not limited to adding noise, changing brightness or contrast, affine transformation, etc. This step superimposes the global perturbation on the first image according to the mask, which can make the adversarial sample difficult to detect while retaining the key information of the image, or ensure that the adversarial effect persists no matter from which angle it is observed.
[0031] As an example, when superimposing the global perturbation on the first image according to the mask, the content of the original image will be changed at the positions where the mask value is 1, while the content of the original image will remain unchanged at the positions where the mask value is 0, thereby obtaining the second image, that is, the Figure 3 perturbed image in.
[0032] Step S102: Obtain the point cloud data corresponding to the first image, input the second image and the point cloud data into the neural network model for training, and obtain the perception heat map of the neural network model.
[0033] Point cloud data refers to a set composed of a large number of discrete points, where each point represents a certain position in three-dimensional space. Each point contains coordinate information (x, y, z). Optionally, each point can also include attributes such as color and reflectivity. Point clouds can be obtained by lidar. For example, in application scenarios such as autonomous vehicles, robot navigation, and 3D modeling, lidar is used to collect the three-dimensional structure of the environment and output it, thus obtaining point cloud data.
[0034] In this embodiment, the benign point cloud data (undisturbed point cloud data) and the second image are input into the neural network model for training to obtain the sum of detection box confidences and the mask area generated by the processing of the neural network model; taking the minimization of the sum of detection box confidences and the mask area as the optimization goal, the gradient descent algorithm is used to perform gradient optimization on the mask and the global perturbation to obtain the optimized mask; the optimized mask is used as the perception heat map of the neural network model. Among them, the neural network model has a certain perception ability after training, so the trained neural network model can be defined as an image-point cloud fusion perception model.
[0035] It should be noted that the object of perturbation in this embodiment is the first image rather than the point cloud data. When training the neural network model, the input objects are the benign point cloud data (undisturbed point cloud data) and the second image (perturbed image), which reduces the requirements for professional equipment, so it is beneficial to reduce costs and deployment difficulties.
[0036] In the process of optimizing the mask, the sum of detection box confidences f score can be input into the mean square error function MSE to obtain the mask loss L mask ; the mask area M is input into the mean square error function MSE to obtain the adversarial loss L adv ; the mask loss L mask and the adversarial loss L adv are added to obtain the sum value. Taking the minimization of the sum value as the optimization goal, the gradient descent algorithm is used to perform gradient optimization on the mask and the global perturbation to obtain the perception sensitive area of the neural network model; the optimized mask is determined based on the sensitive area. Specifically, the mask optimization problem can be formally expressed as finding the value of the following formula: ; where ,, ; λ is a hyperparameter and can take 1, and arg min is used to indicate the parameter value that makes the objective function reach the minimum value.
[0037] In this embodiment, the optimization process for the sum of the two loss functions can be regarded as a dual process: minimizing the adversarial loss means that stronger perturbations need to be superimposed on the image, while conversely, minimizing the mask loss will result in a decrease in the amplitude of the perturbations. Therefore, the optimization process finally converges to obtain the perceptually sensitive region of the neural network model, that is, a new mask M is obtained.
[0038] Step S103: Determine the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution; wherein, the perception type of the neural network model includes a global perception model and a local perception model.
[0039] Wherein, the actual object is the object included in the first image. If it is determined that the perception heat map is similar to the actual object distribution, it is determined that the neural network model is a local perception model; conversely, if it is determined that the perception heat map is not similar to the actual object distribution, it is determined that the neural network model is a global perception model.
[0040] Step S104: Determine the deployment position of the adversarial patch according to the perception type, and model the adversarial patch based on the deployment position to obtain a third image; wherein, the deployment position of the global perception model includes the environmental background in the first image, and the deployment position of the local perception model includes the target object in the first image.
[0041] For the global perception model, it is selected to paste the adversarial patch in the environmental background (such as the ground, wall). For the local perception model, it is selected to paste the adversarial patch on the target object (such as a vehicle, traffic sign). Model the above two patch deployment schemes, and use parameters to control the patch size and paste position, so as to generate a simulated patch in the digital domain. In this application, the adversarial patch is a means to implement the adversarial sample. Therefore, the adversarial sample that appears in some embodiments of this application can specifically refer to the adversarial patch.
[0042] Since this embodiment only involves image perturbation, compared with the multi-modal perturbation (both image and point cloud are perturbed) in the related art, if the traditional single-modal adversarial sample generation method is used, the effectiveness of the image-modal adversarial sample in this embodiment will be insufficient. To overcome this problem, in this embodiment, according to the distribution law of the perception heat map, the trained neural network model (image-point cloud fusion perception model) is classified as a global perception model or a local perception model, so as to make full use of the structural differences of the image-point cloud perception model, and specifically set the deployment scheme of the adversarial patch to overcome the problem of insufficient effectiveness of the single-modal adversarial sample, which is beneficial to improving the adversarial effect of the adversarial sample.
[0043] In the above steps S101 to S104, by perturbing the original image and generating the corresponding perceptual heat map, according to the similarity between the perceptual heat map and the actual object distribution, the perceptual type of the neural network model is determined. Based on the perceptual type of the neural network model, different adversarial patch deployment schemes are selected to perturb the original image with the adversarial patches. Compared with the related technology, this embodiment does not need to perturb the point cloud data, and only generates adversarial samples in the image modality, with simple deployment and low cost. Moreover, this embodiment makes full use of the differences in the image-point cloud perception model structure, overcomes the problem of insufficient effectiveness of single-modal adversarial samples, and is beneficial to improving the adversarial effect of adversarial samples.
[0044] In one embodiment, randomly initializing the mask according to the first image can be achieved by the following method: Assume that x is the first image, l is the point cloud data corresponding to the first image. In this embodiment, only the image modality is disturbed. Therefore, after superimposing the global perturbation p under the control of the mask M, the second image x' is generated, x' = x ⊙ (1 - M) + p ⊙ M. Where ⊙ is the Hadamard product operator, and p is the variable to be optimized, specifically referring to the global perturbation of the same size as the first image x.
[0045] Furthermore, in order to better control the granularity of the perceptual heat map and the convergence speed of the optimization process, this embodiment introduces hyperparameters γ and s to generate the mask M, and the specific calculation is as follows: ; where [i, j] is the pixel position, i is the abscissa, j is the ordinate, m is the variable to be optimized, and tanh is the hyperbolic tangent function.
[0046] In one embodiment, as Figure 4 shown, according to the similarity between the perceptual heat map and the actual object distribution, determining the perceptual type of the neural network model includes the following steps: Step S201, input the first image and the point cloud data into the neural network model for training, and generate an object detection frame on the first image.
[0047] Step S202, set the area surrounded by the object detection frame to 1, and set the rest to 0 to obtain a binary matrix of the same size as the first image.
[0048] Step S203, determine whether the perceptual heat map and the binary matrix meet the preset conditions: ; where x is the first image, M x is the perceptual heat map generated for the first image (i.e., the optimized mask), is the binary matrix, and E x$\overline{\alpha}$ is the mean value of the expression obtained for the perceptual heat maps generated from different first images, $\beta$ is a hyperparameter that can take the value of 3, $\odot$ is the Hadamard product operator, and $\sum$ is the summation operator.
[0049] Step S204: If the perceptual heat map and the binary matrix meet the above conditions, it is determined that the perceptual heat map is similar to the actual object distribution, and the perceptual type of the neural network model is determined to be a local perception model.
[0050] Step S205: If the perceptual heat map and the binary matrix do not meet the above conditions, it is determined that the perceptual heat map is not similar to the actual object distribution, and the perceptual type of the neural network model is determined to be a global perception model.
[0051] In this embodiment, by analyzing the distribution law of the perceptual heat map, a quantitative discrimination rule is designed to ensure accurate determination of the perceptual type of the image-point cloud perception model.
[0052] In one embodiment, as Figure 5 shown, modeling the adversarial patch based on the deployment location includes the following steps: Step S301: Determine the position parameters of the adversarial patch.
[0053] For ease of understanding, Figure 6 a schematic diagram of the process of deploying the adversarial patch on the ground is provided. Combining Figure 5 and Figure 6 shown, assuming that the adversarial patch has a rectangular structure, the rectangular adversarial patch will be deployed on the ground, and several position parameters h, w, x, y, and α are used to describe the deployment scheme of the adversarial patch, as follows: Set the parameters h and w to control the height and width of the adversarial patch respectively, set the parameters x and y to control the x and y coordinates of the center of the adversarial patch in the three-dimensional radar coordinate system (the x-y plane is parallel to the ground), and set the parameter α to control the rotation angle of the adversarial patch in the x-y plane.
[0054] Step S302: According to the position parameters, project the adversarial patch onto the pixel coordinate system of the first image according to the deployment location to obtain a pixel perturbation.
[0055] Obtain the device parameters and the physical position parameters of the adversarial patch; among them, the device parameters include the external parameters of the lidar and the internal parameters of the camera; determine the mapping relationship according to the device parameters, substitute the physical position parameters of the adversarial patch into the mapping relationship, solve to obtain the pixel coordinates; use the bilinear interpolation algorithm to calculate the pixel values corresponding to the pixel coordinates to obtain the pixel perturbation.
[0056] Step S303: Superimpose the pixel perturbation on the first image to obtain a third image.
[0057] In this embodiment, for any given adversarial patch, it is modeled through the corresponding deployment scheme, and the three-dimensional position of the vertex of the adversarial patch deployed in the lidar coordinate system is calculated. According to the camera imaging model, using the camera-lidar parameters (which can be obtained by joint calibration of the camera and lidar), the pixel position of the three-dimensional position projected onto the pixel coordinate system is calculated. Finally, the pixel values of the corresponding area are calculated using bilinear interpolation, and the simulated perturbation is superimposed on the original first image to obtain the third image.
[0058] In one embodiment, after superimposing the pixel perturbation on the first image to obtain the third image, the method further includes: Performing data random augmentation processing on the adversarial patch, and superimposing the adversarial patch after data random augmentation processing on the first image to obtain a fourth image; using multiple fourth images as samples to input into the neural network model, and calculating the average value of the neural network loss; taking minimizing the average value of the neural network loss as the optimization objective, using the gradient descent algorithm to perform gradient optimization on the pixel values of the adversarial patch, and using the optimized adversarial patch as the new adversarial patch.
[0059] In this embodiment, the data random augmentation processing process includes at least one of the following: (1) Randomly transforming the size and physical position parameters of the adversarial patch. Considering the errors in the size of the adversarial patch and the actual deployed physical position, the size and physical position of the adversarial patch can be randomly transformed by randomly transforming the position parameters (h, w, x, y, and α) of the adversarial patch.
[0060] (2) Randomly transforming the contrast, saturation, and brightness of the adversarial patch. Considering the changes in the illumination conditions in the physical world, the contrast, saturation, and brightness of the color of the adversarial patch are randomly transformed within a certain range to simulate the color of the adversarial patch captured by the camera under different illumination conditions.
[0061] (3) During the training process of the neural network model, superimposing pixel-level random perturbations on the adversarial patch. Considering the printing errors of the printer colors, during the training of the neural network model, pixel-level random perturbations are superimposed on the color of the adversarial patch.
[0062] (4) During the training process of the neural network model, superimposing pixel-level random perturbations on the adversarial patch, and adding random noise to the injected random perturbations. In practical applications, images are often affected by noise from the camera itself and environmental interference. To solve this problem, a small amount of random Gaussian noise can be added to the perturbed image (the second image). The random Gaussian noise may not exactly conform to the Gaussian distribution, and to a certain extent, it can improve the physical realizability of the adversarial samples.
[0063] In this embodiment, after determining the deployment location of the adversarial patch according to the perception type and modeling the adversarial patch based on the deployment location, by systematically analyzing a large number of physical factors affecting the induced perturbation (i.e., the error in the deployment location of the physical-world adversarial patch, the change in illumination, the noise of the camera itself, and the printer color deviation), data random augmentation is performed on the adversarial patch using the expected transformation concept, and the patch after data random augmentation is superimposed on the first image to obtain the fourth image. Using several fourth images as sample inputs, calculating the average value of the neural network loss of the neural network model, and using it as an optimization target to optimize the adversarial patch can compensate for environmental changes and printer color errors and enhance the physical robustness of the adversarial samples.
[0064] In one embodiment, multiple fourth images are used as sample inputs to the neural network model, and the average value of the neural network loss is calculated; with the goal of minimizing the average value of the neural network loss as the optimization target, the gradient descent algorithm is used to perform gradient optimization on the pixel values of the adversarial patch, which can be achieved by the following method: The adversarial patch after data random augmentation processing is superimposed into the first image to obtain the fourth image; multiple fourth images are used as sample inputs to the neural network model for processing to obtain the sum of the detection box confidences and the smoothness of the adversarial patch, and the sum of the detection box confidences and the smoothness of the adversarial patch is added to obtain the neural network loss. The average value of the neural network losses of multiple fourth images is obtained, with the goal of minimizing the average value of the neural network loss as the optimization target, the gradient descent algorithm is used to perform gradient optimization on the pixel values of the adversarial patch, and the optimized adversarial patch is used as the new adversarial patch.
[0065] In this embodiment, the perturbation superimposed image (i.e., the fourth image) after data random augmentation is used as the input to the neural network model, and the generation of the adversarial patch is transformed into an end-to-end optimization problem. To generate a visually more natural adversarial patch, the total variation loss L tv is introduced to optimize the color smoothness of the adversarial patch. The specific calculation formula is as follows: ; where x i,j represents the pixel value at the i-th row and j-th column in the adversarial patch.
[0066] For the image-point cloud fusion perception model, in this embodiment, the adversarial loss L adv and the total variation loss L tvThe sum is used as the loss function, and the Adam optimizer is utilized to optimize the pixel values of the adversarial patch through the gradient descent algorithm; it is determined whether the original image with the adversarial patch perturbation satisfies the requirements of the adversarial sample; if not, gradient optimization is performed to update the pixel values of the adversarial patch, and the next iteration is carried out; otherwise, the iteration is stopped, and the current adversarial patch is used as the new adversarial patch to enhance the adversarial effect of the adversarial patch.
[0067] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages, and these steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0068] In one embodiment, Figure 7 Another flowchart of an adversarial patch generation method for an image-point cloud fusion perception model is provided, as Figure 7 shown, and the method includes the following steps: Step S401, obtaining the structure and parameters of the image-point cloud fusion perception model; Step S402, based on the image-point cloud fusion perception model, generating corresponding perception heatmaps for different original images respectively, and determining the perception type of the image-point cloud fusion perception model; Step S403, determining the deployment position of the adversarial patch based on the perception type of the image-point cloud fusion perception model, and projecting the adversarial patch in the digital domain to the deployment position based on the camera imaging model; Step S404, performing data augmentation on the generated adversarial patch to generate multiple randomly transformed images with the adversarial patch as the input of the image-point cloud fusion perception model; Step S405, using the gradient descent algorithm to iteratively optimize and update the pixel values of the adversarial patch, and obtaining a physically feasible adversarial patch when the iteration stop condition is reached or the maximum number of iterations is exceeded.
[0069] Based on the same inventive concept, an embodiment of the present application further provides an electronic device for implementing the above-mentioned method for generating adversarial patches for an image-point cloud fusion perception model. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0070] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0071] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program: Obtain a first image, randomly initialize a mask and a global perturbation according to the first image, and superimpose the global perturbation on the first image according to the mask to obtain a second image; Obtain the point cloud data corresponding to the first image, input the second image and the point cloud data into a neural network model for training, and obtain the perception heat map of the neural network model; Determine the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution; wherein, the perception types of the neural network model include a global perception model and a local perception model; Determine the deployment position of the adversarial patch according to the perception type, and model the adversarial patch based on the deployment position to obtain a third image; wherein, the deployment positions of the global perception model include the environmental background in the first image, and the deployment positions of the local perception model include the target objects in the first image.
[0072] It should be noted that specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated in this embodiment.
[0073] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an adversarial patch generation method for an image-point cloud fusion perception model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0074] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0075] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of any of the above method embodiments.
[0076] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the steps of any of the above method embodiments.
[0077] In one embodiment, when the computer program is executed by a processor, it implements the following method steps: Obtain a first image, randomly initialize a mask and a global perturbation according to the first image, and superimpose the global perturbation on the first image according to the mask to obtain a second image; Obtain point cloud data corresponding to the first image, input the second image and the point cloud data into a neural network model for training, and obtain a perception heat map of the neural network model; Determine the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution; among them, the perception type of the neural network model includes a global perception model and a local perception model; Determine the deployment location of the adversarial patch according to the perception type, and model the adversarial patch based on the deployment location to obtain a third image; wherein, the deployment location of the global perception model includes the environmental background in the first image, and the deployment location of the local perception model includes the target object in the first image.
[0078] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Obtain the sum of detection box confidences and the mask area generated by processing via a neural network model; Taking the minimization of the sum of detection box confidences and the mask area as the optimization objective, use the gradient descent algorithm to perform gradient optimization on the mask and the global perturbation to obtain an optimized mask; Use the optimized mask as the perception heatmap of the neural network model.
[0079] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Set hyperparameters γ and s, and set variables m and global perturbation p; Generate a mask according to the following formula: ; wherein, M is the mask, [i, j] is the pixel position, and tanh is the hyperbolic tangent function.
[0080] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Input the first image and the point cloud data into the neural network model for training, and generate object detection boxes on the first image; Set the area surrounded by the object detection boxes to 1, and set the rest to 0 to obtain a binary matrix with the same size as the first image; Determine whether the perception heatmap and the binary matrix satisfy the following conditions: ; wherein, x is the first image, M x is the perception heatmap generated for the first image, is the binary matrix, E x is the mean value of the expressions obtained for the perception heatmaps generated by different first images, β is a hyperparameter, ⊙ is the Hadamard product operator, and ∑ is the summation operator; If the perception heatmap and the binary matrix satisfy the above conditions, it is determined that the perception heatmap is similar to the actual object distribution, and the perception type of the neural network model is determined to be a local perception model; If the perception heatmap and the binary matrix do not satisfy the above conditions, it is determined that the perception heatmap is not similar to the actual object distribution, and the perception type of the neural network model is determined to be a global perception model.
[0081] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Determine the position parameters of the adversarial patch; According to the position parameters, project the adversarial patch into the pixel coordinate system of the first image according to the deployment position to obtain pixel perturbations; Superimpose the pixel perturbations on the first image to obtain a third image.
[0082] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Obtain the device parameters and the physical position parameters of the adversarial patch; wherein, the device parameters include the external parameters of the lidar and the internal parameters of the camera; Determine the mapping relationship according to the device parameters, substitute the physical position parameters of the adversarial patch into the mapping relationship, and solve to obtain the pixel coordinates; Use the bilinear interpolation algorithm to calculate the pixel values corresponding to the pixel coordinates to obtain pixel perturbations.
[0083] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Perform data random augmentation processing on the adversarial patch, and superimpose the adversarial patch after data random augmentation processing on the first image to obtain a fourth image; Use multiple fourth images as samples and input them into the neural network model to calculate the average value of the neural network loss; Taking the minimization of the average value of the neural network loss as the optimization goal, use the gradient descent algorithm to perform gradient optimization on the pixel values of the adversarial patch, and use the optimized adversarial patch as the new adversarial patch.
[0084] In one embodiment, when the computer program is executed by a processor, the following method steps are implemented: Randomly transform the size and physical position parameters of the adversarial patch; Randomly transform the contrast, saturation, and brightness of the adversarial patch; During the training process of the neural network model, superimpose pixel-level random perturbations on the adversarial patch; During the training process of the neural network model, superimpose pixel-level random perturbations on the adversarial patch, and add random noise to the injected random perturbations.
[0085] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0086] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0087] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0088] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating adversarial patches for an image-point cloud fusion perception model, characterized in that: include: Acquire a first image, randomly initialize a mask and a global perturbation according to the first image, and superimpose the global perturbation on the first image according to the mask to obtain a second image; Acquire point cloud data corresponding to the first image, input the second image and the point cloud data into a neural network model for training, and obtain a perceptual heat map of the neural network model; Determining the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution; wherein the perception type of the neural network model includes a global perception model and a local perception model; Determine the deployment position of the adversarial patch according to the perception type, and model the adversarial patch based on the deployment position to obtain a third image; wherein the deployment position of the global perception model includes the environmental background in the first image, and the deployment position of the local perception model includes the target object in the first image.
2. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 1, characterized in that: Inputting the second image and the point cloud data into a neural network model for training to obtain a perceptual heat map of the neural network model includes: Obtaining a total confidence level of a detection frame and a mask area generated by processing the neural network model; Taking minimizing the sum of the confidence scores of the detection box and the area of the mask as the optimization goal, a gradient descent algorithm is used to perform gradient optimization on the mask and the global perturbation to obtain an optimized mask; The optimized mask is used as a perceptual heat map of the neural network model.
3. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 1, characterized in that: Randomly initializing a mask according to the first image, comprising: Set the hyperparameters γ and s, and set the variable m and global perturbation p; The mask is generated according to the following formula: ; Wherein, M is the mask, [i, j] is the pixel position, and tanh is the hyperbolic tangent function.
4. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 1, characterized in that: Determining the perception type of the neural network model according to the similarity between the perception heat map and the actual object distribution includes: Inputting the first image and the point cloud data into the neural network model for training, and generating an object detection frame on the first image; The area surrounded by the object detection frame is set to 1, and the rest of the area is set to 0, to obtain a binary matrix of the same size as the first image; Determine whether the perceptual heat map and the binary matrix meet the following conditions: ; Wherein, x is the first image, M x is the perceptual heat map generated for the first image, is the binary matrix, E x is the mean of the expression obtained for the perceptual heat maps generated by different first images, β is a hyperparameter, ⊙ is a Hadamard product operator, and ∑ is a summation operator; If the perception heat map and the binary matrix meet the above conditions, it is determined that the perception heat map and the actual object distribution are similar, and the perception type of the neural network model is determined to be a local perception model; If the perception heat map and the binary matrix do not satisfy the above conditions, it is determined that the perception heat map and the actual object distribution are not similar, and the perception type of the neural network model is determined to be a global perception model.
5. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 1, characterized in that: Modeling the adversarial patch based on the deployment position to obtain a third image includes: Determining position parameters of the adversarial patch; According to the position parameter, projecting the adversarial patch into a pixel coordinate system of the first image at the deployment position to obtain a pixel perturbation; The pixel disturbance is superimposed on the first image to obtain the third image.
6. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 5, characterized in that: Projecting the adversarial patch into the pixel coordinate system according to the deployment position to obtain a pixel perturbation includes: Acquire device parameters and physical position parameters of the adversarial patch; wherein the device parameters include external parameters of the laser radar and internal parameters of the camera; Determine a mapping relationship according to the device parameters, substitute the physical position parameters of the adversarial patch into the mapping relationship, and solve to obtain pixel coordinates; The pixel value corresponding to the pixel coordinate is calculated using a bilinear interpolation algorithm to obtain the pixel disturbance.
7. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 5, characterized in that: After superimposing the pixel disturbance on the first image to obtain the third image, the method further includes: Performing random data enhancement processing on the adversarial patch, and superimposing the adversarial patch after the random data enhancement processing onto the first image to obtain a fourth image; Inputting a plurality of the fourth images as samples into the neural network model, and calculating an average value of the neural network loss; Taking minimizing the average value of the neural network loss as the optimization goal, a gradient descent algorithm is used to perform gradient optimization on the pixel value of the adversarial patch, and the optimized adversarial patch is used as a new adversarial patch.
8. The method for generating adversarial patches for an image-point cloud fusion perception model according to claim 7, characterized in that: Performing random data enhancement processing on the adversarial patch includes at least one of the following: Randomly transforming the size and physical position parameters of the adversarial patch; Randomly transforming the contrast, saturation, and brightness of the adversarial patch; During the training process of the neural network model, superimposing pixel-level random perturbations on the adversarial patch; During the training process of the neural network model, pixel-level random perturbations are superimposed on the adversarial patch, and random noise is added to the injected random perturbations.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Method and device for evaluating robustness of neural network image classification model
CN114239685A
Remote sensing image-oriented patch deployable attack resisting method, device and equipment
CN116844052A
Complex road target detection method based on multi-modal fusion aerial view
CN117058646A
End-to-end method for generating imperceptible adversarial patch
CN118229954A
Picture high-quality arbitrary style migration method based on local and global style learning
CN119251069A
Cited By
Stereoscopic vision data reconstruction robustness evaluation method and system based on multiple modes
CN120451420A
Robustness evaluation method and system for multimodal stereo vision data reconstruction
CN120451420B
Testing method and device for calculating robustness of imaging model
CN121582717A
Semantic-guided diffusion model-based autonomous driving adversarial scene generation method and device
CN122527022A