Adversarial Sample Generation Method Based on Multi-Sensor Fusion for Autonomous Driving

By perceiving and perturbing the three-dimensional object model of multi-sensors in the autonomous driving system, target adversarial samples are generated, and the problem of poor adversarial testing in the prior art is solved, and more effective adversarial testing is achieved.

CN119723251BActive Publication Date: 2025-07-01TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510213997.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-01
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

When conducting adversarial testing of autonomous driving systems, the prior art tests are only conducted for image recognition, and the actual scenarios of multi-sensor fusion cannot be effectively simulated, resulting in poor testing results.

Method used

By perceiving three-dimensional object models in multiple perspectives, visual and radar-aware data are acquired and radar-aware data are perturbed to generate candidate adversarial samples. Then, based on the visual loss data, radar loss data, alignment loss data, and smooth loss data, the candidate adversarial samples are adjusted to generate the target adversarial samples.

Benefits of technology

The generated target adversarial samples can more effectively simulate obstacles after adding perturbations in multi-sensor fusion scenarios, improving the effectiveness and robustness of adversarial testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723251B_ABST
    Figure CN119723251B_ABST
Patent Text Reader

Abstract

The present invention provides an adversarial sample generation method based on multi-sensor fusion for autonomous driving, which can be applied to the field of adversarial network technology. The above method includes: for each of multiple perspectives, respectively performing perception on a three-dimensional object model to obtain perception data, where the perception data includes visual perception data and radar perception data; obtaining visual loss data based on the visual perception data and the perturbed visual perception data obtained by perturbing the visual perception data, and obtaining radar loss data based on the radar perception data and the perturbed radar perception data obtained by perturbing the radar perception data; calculating an alignment loss and a smooth loss respectively based on the candidate adversarial samples generated from the perturbed perception data of multiple perspectives and the visual perception data to obtain alignment loss data and smooth loss data; adjusting the candidate adversarial samples based on the visual loss data, the radar loss data, the alignment loss data and the smooth loss data to generate target adversarial samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of adversarial networks, and more particularly to an adversarial sample generation method based on multi-sensor fusion for autonomous driving, an adversarial testing method and device for an autonomous driving system. Background Art

[0002] With the development of deep learning technology, autonomous driving technology has gradually matured. Neural networks have achieved great success in fields such as image classification and object detection. The application of deep neural networks in visual and radar perception is very extensive. However, deep neural networks are vulnerable to adversarial attacks.

[0003] In related technologies, when testing autonomous driving technology, adversarial testing is often only carried out for image recognition in the autonomous driving system. However, current autonomous driving systems often use multiple sensors for recognition. Testing the autonomous driving system with only the adversarial samples generated for image recognition testing will result in poor adversarial testing effects. Summary of the Invention

[0004] In view of the above problems, the present invention provides an adversarial sample generation method based on multi-sensor fusion for autonomous driving.

[0005] According to a first aspect of the present invention, there is provided an adversarial sample generation method based on multi-sensor fusion for autonomous driving, including: for each of a plurality of perspectives, respectively performing perception on a three-dimensional object model to obtain perception data, where the three-dimensional object model is used to simulate obstacles encountered during autonomous driving, and the perception data includes visual perception data and radar perception data; obtaining visual loss data based on the visual perception data and perturbed visual perception data obtained by perturbing the visual perception data, and obtaining radar loss data based on the radar perception data and perturbed radar perception data obtained by perturbing the radar perception data; calculating an alignment loss and a smooth loss respectively based on a candidate adversarial sample generated from the perturbed perception data of the plurality of perspectives and the visual perception data to obtain alignment loss data and smooth loss data; adjusting the candidate adversarial sample based on the visual loss data, the radar loss data, the alignment loss data and the smooth loss data to generate a target adversarial sample, where the target adversarial sample is used to simulate an obstacle with increased perturbation.

[0006] According to an embodiment of the present invention, adjusting the candidate adversarial sample based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data to generate a target adversarial sample includes: obtaining a total loss function value based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data; using a stochastic gradient descent algorithm to solve for target perturbation data that minimizes the total loss function value; and adjusting the candidate adversarial sample based on the target perturbation data to obtain a target adversarial sample.

[0007] According to an embodiment of the present invention, the three-dimensional object model is constructed by the following method: obtaining respective images and sampling attribute data of a target object from multiple perspectives, where the sampling attribute data includes the acquisition positions of the image acquisition devices corresponding to different perspectives and the acquisition angles relative to the target object; inputting the multiple images and the sampling attribute data into a trained volume radiation field model to obtain the volume radiation field of the target object; and obtaining the three-dimensional object model based on the volume radiation field and the sampling attribute data.

[0008] According to an embodiment of the present invention, the sampling attribute data further includes the transmittance, opacity, and color of each of multiple sampling points, and the multiple sampling points are distributed on multiple rays from the optical center of the image acquisition device to the target object. Obtaining the three-dimensional object model based on the volume radiation field and the sampling attribute data includes: for each ray, obtaining a ray rendering color based on the sampling point transmittance, the sampling point opacity, and the sampling point color; and rendering the volume radiation field based on the ray rendering colors of the multiple rays to obtain the three-dimensional object model.

[0009] According to an embodiment of the present invention, the method further includes: using a mean squared error loss function to obtain a mean squared loss value based on the pixel values of the three-dimensional object model and the pixel values of the target object; and adjusting the three-dimensional object model based on the mean squared loss value.

[0010] According to an embodiment of the present invention, the visual loss data is calculated by the following formula: where represents the visual loss data, N represents the total number of perspectives, represents the color value of the perturbed visual perception data at the i-th perspective, represents the color value of the visual perception data at the i-th perspective. The radar loss data is calculated by the following formula: where represents the radar loss data, represents the cross-entropy loss function, represents the point cloud prediction mask of the perturbed radar perception data, Represents the ground truth mask of the point cloud for radar perception data.

[0011] According to an embodiment of the present invention, the above alignment loss data is calculated by the following formula: , where represents the above alignment loss data, n represents the total number of points in the point cloud formed by the candidate adversarial samples, represents the weight of the k-th point, represents the feature of the k-th point in the candidate adversarial sample, represents the feature of the k-th point in the visual perception data; the above smooth loss data is calculated by the following formula: , where represents the above smooth loss data, represents the pixel value at the position (i, j) on the above visual perception data, represents the pixel value at the position (i + 1, j) on the above visual perception data, represents the pixel value at the position (i, j + 1) on the above visual perception data.

[0012] The second aspect of the present invention provides an adversarial testing method for an autonomous driving system, including: obtaining a target adversarial sample, where the target adversarial sample is obtained by using the method of the first aspect; using the target adversarial sample to test the autonomous driving system to obtain the adversarial test result of the autonomous driving system, and the adversarial test result is used to improve the safety of the autonomous driving system.

[0013] The third aspect of the present invention provides an adversarial sample generation device based on autonomous driving multi-sensor fusion, including: a perception module, configured to respectively perceive a three-dimensional object model for each of multiple perspectives to obtain perception data, where the three-dimensional object model is used to simulate obstacles encountered during the autonomous driving process, and the perception data includes visual perception data and radar perception data; a perturbation module, configured to obtain visual loss data based on the visual perception data and the perturbed visual perception data obtained by perturbing the visual perception data, and obtain radar loss data based on the radar perception data and the perturbed radar perception data obtained by perturbing the radar perception data; a loss data calculation module, configured to calculate an alignment loss and a smooth loss respectively based on the candidate adversarial samples generated from the perturbed perception data of the multiple perspectives and the visual perception data to obtain alignment loss data and smooth loss data; a target adversarial sample generation module, configured to adjust the candidate adversarial samples based on the visual loss data, the radar loss data, the alignment loss data, and the smooth loss data to generate target adversarial samples, and the target adversarial samples are used to simulate obstacles with increased perturbations.

[0014] The fourth aspect of the present invention provides an adversarial testing device for an autonomous driving system, including: an acquisition module for acquiring a target adversarial sample, where the target adversarial sample is obtained by using the method of the first aspect; a testing module for testing the autonomous driving system by using the target adversarial sample to obtain an adversarial test result of the autonomous driving system, and the adversarial test result is used to improve the safety of the autonomous driving system.

[0015] The fifth aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] The sixth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0017] The seventh aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0018] The present invention perceives a three-dimensional object model from multiple perspectives to obtain visual perception data and radar perception data, and simultaneously perturbs the visual perception data and radar perception data to obtain candidate adversarial samples. By calculating the alignment loss and smooth loss between the candidate adversarial samples and the perception data, and adjusting the candidate adversarial samples through the visual loss data, radar loss data, alignment loss data, and smooth loss data to generate target adversarial samples, the corresponding relationship between the generated target adversarial samples and the three-dimensional object model is ensured. By using the smooth loss data to adjust the candidate adversarial samples to reduce the difference in adjacent pixel values, the naturalness and robustness of the generated target adversarial samples in the physical scene are ensured, thereby improving the effect of adversarial testing. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0020] Figure 1 Shows an application scenario diagram of an adversarial sample generation method for autonomous driving multi-sensor fusion, an adversarial testing method for an autonomous driving system, and a device according to an embodiment of the present invention.

[0021] Figure 2 Shows a flowchart of an adversarial sample generation method for autonomous driving multi-sensor fusion according to an embodiment of the present invention.

[0022] Figure 3 Shows a flowchart of generating a target adversarial sample according to an embodiment of the present invention.

[0023] Figure 4 Shows a flowchart of an adversarial testing method for an autonomous driving system according to an embodiment of the present invention.

[0024] Figure 5 Shows a structural block diagram of an adversarial sample generation device based on multi-sensor fusion for autonomous driving according to an embodiment of the present invention.

[0025] Figure 6 Shows a structural block diagram of an adversarial testing device for an autonomous driving system according to an embodiment of the present invention.

[0026] Figure 7 Shows a block diagram of an electronic device suitable for implementing an adversarial sample generation method based on multi-sensor fusion for autonomous driving according to an embodiment of the present invention. Detailed implementation manners

[0027] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0031] With the development of deep learning technology, neural networks have achieved great success in fields such as image classification and object detection. However, neural networks also have some deficiencies, and one of the most important problems is adversarial sample attacks. Adversarial sample attacks refer to making tiny modifications to the original input, which causes the output result of the neural network to be incorrect. This attack method may lead to serious security problems.

[0032] With the rapid development of deep learning technology, autonomous driving technology has gradually matured, and among them, the application of DNN (Deep Neural Network) in visual and radar perception is particularly extensive. However, research shows that DNN models are vulnerable to adversarial attacks, that is, by applying tiny perturbations to the input data, the prediction results of the model can be significantly changed. This poses a serious safety hazard to autonomous driving systems that rely on precise perception.

[0033] Existing two-dimensional adversarial attack methods mainly target image classification, but these methods have certain limitations in actual three-dimensional scenarios. In addition, current three-dimensional adversarial attack methods mostly focus on modifying the texture or geometry of three-dimensional models. However, the consistency and transferability of these methods under different perspectives are poor, and it is difficult to effectively test multi-modal perception systems.

[0034] Embodiments of the present invention provide an adversarial sample generation method based on autonomous driving multi-sensor fusion, including: for each of multiple perspectives, respectively performing perception on a three-dimensional object model to obtain perception data, where the three-dimensional object model is used to simulate obstacles encountered during the autonomous driving process, and the perception data includes visual perception data and radar perception data; obtaining visual loss data based on the visual perception data and the perturbed visual perception data obtained by perturbing the visual perception data, and obtaining radar loss data based on the radar perception data and the perturbed radar perception data obtained by perturbing the radar perception data; calculating an alignment loss and a smoothing loss respectively based on the candidate adversarial samples generated from the perturbed perception data of multiple perspectives and the visual perception data to obtain alignment loss data and smoothing loss data; adjusting the candidate adversarial samples based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data to generate target adversarial samples, where the target adversarial samples are used to simulate obstacles with increased perturbations.

[0035] Figure 1 Shows an application scenario diagram of the adversarial sample generation method for autonomous driving multi-sensor fusion, the adversarial test method for autonomous driving systems, and the device according to the embodiments of the present invention.

[0036] As Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0037] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0038] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablets, laptop portable computers, and desktop computers, etc.

[0039] The server 105 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server may analyze and process data such as received user requests, etc., and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests, etc.) to the terminal device.

[0040] It should be noted that the adversarial sample generation method based on multi-sensor fusion for autonomous driving and the adversarial testing method for autonomous driving systems provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the adversarial sample generation device based on multi-sensor fusion for autonomous driving and the adversarial testing device for autonomous driving systems provided by the embodiments of the present invention can generally be disposed in the server 105. The adversarial sample generation method based on multi-sensor fusion for autonomous driving and the adversarial testing method for autonomous driving systems provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the adversarial sample generation device based on multi-sensor fusion for autonomous driving and the adversarial testing device for autonomous driving systems provided by the embodiments of the present invention can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0041] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0042] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 3 the described scenario, and will describe in detail the adversarial sample generation method based on multi-sensor fusion for autonomous driving of the embodiments of the present invention through

[0043] Figure 2 FIG. shows a flowchart of the adversarial sample generation method based on multi-sensor fusion for autonomous driving according to an embodiment of the present invention.

[0044] As Figure 2 shown, the adversarial sample generation method based on multi-sensor fusion for autonomous driving in this embodiment includes operations S210 to S240.

[0045] In operation S210, for each of multiple perspectives, the three-dimensional object model is respectively perceived to obtain perception data.

[0046] Among them, the three-dimensional object model is used to simulate obstacles encountered during autonomous driving, and the perception data includes visual perception data and radar perception data.

[0047] According to an embodiment of the present invention, the above multi-sensors may include a visual perception sensor and a radar perception sensor.

[0048] According to an embodiment of the present invention, the above three-dimensional object can be obtained by modeling a target object based on a neural radiance field, or can be obtained by modeling and further rendering a target object based on methods such as a grid or a voxel grid. The present invention does not limit this.

[0049] According to an embodiment of the present invention, the above perspectives include the angle, illumination, and background of the three-dimensional object. By adjusting and perceiving the angle, illumination, and background of the three-dimensional object, perception data under multiple perspectives can be obtained, thereby ensuring robustness in a real physical environment.

[0050] In operation S220, visual loss data is obtained based on visual perception data and perturbed visual perception data obtained by perturbing the visual perception data, and radar loss data is obtained based on radar perception data and perturbed radar perception data obtained by perturbing the radar perception data.

[0051] According to an embodiment of the present invention, the three-dimensional object model can be perturbed by the following formula (1) to obtain perturbed visual perception data and perturbed radar perception data.

[0052]

[0053] Among them, represents the set of perturbed visual perception data and perturbed radar perception data, represents the three-dimensional object model, represents a randomly initialized adversarial perturbation.

[0054] According to an embodiment of the present invention, by perturbing the three-dimensional object model, the visual perception data and the radar perception data can be perturbed simultaneously to obtain perturbed visual perception data and perturbed radar perception data.

[0055] According to an embodiment of the present invention, visual loss data can be obtained based on visual perception data and perturbed visual perception data through a first preset loss function, and radar loss data can be obtained based on radar perception data and perturbed radar perception data through a second preset loss function.

[0056] In operation S230, an alignment loss and a smoothness loss are calculated respectively based on candidate adversarial samples generated from perturbed perception data of multiple perspectives and visual perception data to obtain alignment loss data and smoothness loss data.

[0057] In operation S240, the candidate adversarial samples are adjusted based on the visual loss data, the radar loss data, the alignment loss data, and the smoothness loss data to generate target adversarial samples.

[0058] Among them, the target adversarial samples are used to simulate an obstacle after increased perturbation.

[0059] The present invention perceives a three-dimensional object model from multiple perspectives to obtain visual perception data and radar perception data, and simultaneously perturbs the visual perception data and radar perception data to obtain candidate adversarial samples. By calculating the alignment loss and the smoothing loss for the candidate adversarial samples and the perception data, and adjusting the candidate adversarial samples through the visual loss data, radar loss data, alignment loss data, and smoothing loss data to generate target adversarial samples, the corresponding relationship between the generated target adversarial samples and the three-dimensional object model is ensured. The candidate adversarial samples are adjusted through the smoothing loss data to reduce the difference in adjacent pixel values, ensuring the naturalness and robustness of the generated target adversarial samples in the physical scenario, thereby improving the effect of adversarial testing.

[0060] According to an embodiment of the present invention, adjusting the candidate adversarial samples based on the visual loss data, radar loss data, alignment loss data, and smoothing loss data to generate target adversarial samples includes: obtaining a total loss function value based on the visual loss data, radar loss data, alignment loss data, and smoothing loss data; using the stochastic gradient descent algorithm to solve for the target perturbation data that minimizes the total loss function value; and adjusting the candidate adversarial samples based on the target perturbation data to obtain the target adversarial samples.

[0061] According to an embodiment of the present invention, the total loss function can be calculated by the following formula (2).

[0062]

[0063] Wherein, represents the total loss function, represents the visual loss data, represents the radar loss data, represents the alignment loss data, represents the smoothing loss data, is a hyperparameter.

[0064] According to an embodiment of the present invention, appropriate hyperparameters can be adjusted according to the scenario to generate target adversarial samples that can perform sufficient adversarial testing.

[0065] According to an embodiment of the present invention, optimization algorithms such as the stochastic gradient descent algorithm or the batch gradient descent algorithm can be used to iteratively optimize the total loss value and adjust the perturbation to the three-dimensional object model, thereby minimizing the total loss.

[0066] According to an embodiment of the present invention, the effectiveness of the generated target adversarial samples can be verified on a dataset, and the aforementioned effectiveness includes the attack success rate for visual perception sensors and radar perception sensors.

[0067] According to an embodiment of the present invention, by integrating visual loss data, radar loss data, alignment loss data, and smoothing loss data, the performance of the model in different perception tasks and the consistency and smoothness of the model output are comprehensively considered, and the target perturbation data is obtained through the stochastic gradient descent algorithm to adjust the candidate adversarial samples, so as to obtain the target adversarial samples that can maximize the adversarial test of the autonomous driving system.

[0068] According to an embodiment of the present invention, the three-dimensional object model is constructed by the following method: acquiring the images and sampling attribute data of the target object from multiple perspectives respectively, wherein the sampling attribute data includes the acquisition positions of the image acquisition devices corresponding to different perspectives and the acquisition angles relative to the target object; inputting the multiple images and sampling attribute data into the trained volume radiation field model to obtain the volume radiation field of the target object; and obtaining the three-dimensional object model based on the volume radiation field and the sampling attribute data.

[0069] According to an embodiment of the present invention, images of the target object from multiple perspectives can be collected in the scene to establish a three-dimensional data set , where represents the images from different perspectives. The volume radiation field of the three-dimensional object can be generated by the trained volume radiation field model, and then the three-dimensional object model can be obtained by rendering the volume radiation field based on the volume radiation field and the sampling attribute data.

[0070] According to an embodiment of the present invention, the above sampling attribute data further includes the transmittance, opacity, and color of each of the multiple sampling points, and the multiple sampling points are distributed on multiple rays from the optical center of the image acquisition device to the target object. Obtaining the three-dimensional object model based on the volume radiation field and the sampling attribute data includes: for each ray, obtaining the ray rendering color based on the sampling point transmittance, sampling point opacity, and sampling point color; and rendering the volume radiation field based on the ray rendering colors of the multiple rays to obtain the three-dimensional object model.

[0071] According to an embodiment of the present invention, sampling can be performed on the target object in the three-dimensional scene along the line of sight or the light direction, and the density, color, and transparency of the sampling points jointly determine the final pixel color of the ray.

[0072] The ray rendering color can be calculated by the following formula (3).

[0073]

[0074] Where represents the final pixel color rendered along the ray r, N represents the number of sampling points on the ray, represents the transmittance from the starting point of the ray to the i-th sampling point, is the volume density or absorption coefficient of the i-th sampling point, represents the distance between adjacent sampling points, represents the opacity of the i-th sampling point, represents the color of the i-th sampling point.

[0075] According to the embodiments of the present invention, the above three-dimensional object model can be obtained by calculating each ray to render the color first and then further rendering the volume radiation field according to the above formula (3).

[0076] According to the embodiments of the present invention, the volume density and color can be predicted through a neural radiation field model. The above neural radiation field model receives the input five-dimensional coordinates (x, y, z, θ, ψ), where x, y, z are the position coordinates in three-dimensional space, and θ, ψ represent the viewing angles (azimuth angle and elevation angle respectively), outputs the volume density of each sampling point and the emission amplitude value (color) related to the viewing angle at this position, and predicts the transmittance and opacity of the sampling point through the volume density, so as to render the volume radiation field to obtain a three-dimensional object model.

[0077] According to the embodiments of the present invention, by obtaining the images and sampling attribute data of the target object from multiple perspectives and inputting them into the trained volume radiation field model to obtain the volume radiation field of the target object, the details of the target object can be accurately captured, so as to generate a more realistic three-dimensional model; by calculating the sampling attributes to obtain the ray rendering color of each ray, and finally rendering the volume radiation field based on the ray rendering color to obtain a three-dimensional object model, comprehensively considering the scattering and absorption of light when passing through the object, by calculating the transmittance, opacity and color of the sampling points on each ray, a more accurate rendering result can be obtained, so as to better simulate the lighting and material effects in the real world.

[0078] According to the embodiments of the present invention, the above method further includes: using the mean square error loss function to obtain the mean square loss value based on the pixel values of the three-dimensional object model and the pixel values of the target object; adjusting the three-dimensional object model based on the mean square loss value.

[0079] According to the embodiments of the present invention, the volume radiation field can be rendered through a neural radiation field model, and the neural radiation field model can be adjusted by the following formula (4).

[0080]

[0081] Wherein, represents the mean square error value of the pixel color, represents the number of rays participating in the neural radiation field training process, r represents a single ray, and R represents the set of all sampling rays, represents the final pixel color rendered along the ray r, Represents the true pixel color on the ray r.

[0082] According to an embodiment of the present invention, the neural radiance field model can be adjusted to minimize the mean square error value in the above formula (4), thereby adjusting the three-dimensional object model.

[0083] According to an embodiment of the present invention, the above visual loss data is calculated by the following formula (5):

[0084]

[0085] Wherein, Represents the visual loss data, N represents the total number of viewpoints, Represents the color value of the perturbed visual perception data at the i-th viewpoint, Represents the color value of the visual perception data at the i-th viewpoint.

[0086] The radar loss data is calculated by the following formula (6):

[0087] (6)

[0088] Wherein, Represents the radar loss data, Represents the cross-entropy loss function, Represents the point cloud prediction mask of the perturbed radar perception data, Represents the point cloud true mask of the radar perception data.

[0089] According to an embodiment of the present invention, the alignment loss data is calculated by the following formula (7):

[0090] (7)

[0091] Wherein, Represents the alignment loss data, n represents the total number of points in the point cloud formed by the candidate adversarial samples, Represents the weight of the k-th point, Represents the feature of the k-th point in the candidate adversarial sample, Represents the feature of the k-th point in the visual perception data.

[0092] The smooth loss data is calculated by the following formula (8):

[0093] (8)

[0094] Wherein, Represents the smooth loss data, Represents the pixel value at the (i, j) position on the visual perception data, represents the pixel value at the position (i+1, j) on the visual perception data, represents the pixel value at the position (i, j+1) on the visual perception data.

[0095] According to the embodiments of the present invention, the above alignment loss data is used to ensure the point-to-point correspondence between the generated radar point cloud and the real point cloud. The above smoothing loss data adopts the total variation strategy to reduce the difference between adjacent pixel values, ensuring the naturalness of the generated target adversarial samples in the physical scene. By jointly optimizing the alignment loss and the smoothing loss, smoother and more natural target adversarial samples can be generated.

[0096] Figure 3 Shows a flowchart of generating target adversarial samples according to an embodiment of the present invention.

[0097] First, construct a dataset through the input multi-view images, and generate the volume radiation field of the three-dimensional object. Then, sample along the rays and render the volume radiation field to obtain the three-dimensional object model. By perturbing the visual perception data and the radar perception data, calculate the visual loss data, radar loss data, alignment loss data, and smoothing loss data to adjust the candidate adversarial samples. Finally, obtain the target adversarial samples, realizing the joint attack on the visual and radar sensors, and ensuring the robustness and naturalness under physical conditions. Finally, the relevant data can be input into the neural radiation field model to generate new views and perform the next round of iteration.

[0098] Figure 4 Shows a flowchart of an adversarial testing method for an autonomous driving system according to an embodiment of the present invention.

[0099] As Figure 4 shown, the adversarial sample generation method based on autonomous driving multi-sensor fusion in this embodiment includes operation S410 to operation S420.

[0100] In operation S410, obtain the target adversarial sample.

[0101] Among them, the target adversarial sample is obtained according to the above adversarial sample generation method based on autonomous driving multi-sensor fusion.

[0102] In operation S420, use the target adversarial sample to test the autonomous driving system to obtain the adversarial test result of the autonomous driving system.

[0103] Among them, the adversarial test result is used to improve the safety of the autonomous driving system.

[0104] Based on the above adversarial sample generation method based on autonomous driving multi-sensor fusion, the present invention also provides an adversarial sample generation device based on autonomous driving multi-sensor fusion. The following will be combined with Figure 5Describe the device in detail.

[0105] Figure 5 The structural block diagram of an adversarial sample generation device based on multi - sensor fusion for autonomous driving according to an embodiment of the present invention is shown.

[0106] As Figure 5 shown, the adversarial sample generation device 500 based on multi - sensor fusion for autonomous driving in this embodiment includes a perception module 510, a perturbation module 520, a loss data calculation module 530, and a target adversarial sample generation module 540.

[0107] The perception module 510 is used to perceive a three - dimensional object model for each of multiple perspectives to obtain perception data. The three - dimensional object model is used to simulate obstacles encountered during autonomous driving, and the perception data includes visual perception data and radar perception data. In one embodiment, the perception module 510 can be used to perform the operation S210 described above, which will not be elaborated here.

[0108] The perturbation module 520 is used to obtain visual loss data based on the visual perception data and the perturbed visual perception data obtained by perturbing the visual perception data, and obtain radar loss data based on the radar perception data and the perturbed radar perception data obtained by perturbing the radar perception data. In one embodiment, the perturbation module 520 can be used to perform the operation S220 described above, which will not be elaborated here.

[0109] The loss data calculation module 530 is used to calculate the alignment loss and the smooth loss respectively based on the candidate adversarial samples generated from the perturbed perception data of multiple perspectives and the visual perception data, to obtain the alignment loss data and the smooth loss data. In one embodiment, the loss data calculation module 530 can be used to perform the operation S230 described above, which will not be elaborated here.

[0110] The target adversarial sample generation module 540 is used to adjust the candidate adversarial samples based on the visual loss data, the radar loss data, the alignment loss data, and the smooth loss data to generate target adversarial samples. The target adversarial samples are used to simulate obstacles with increased perturbations. In one embodiment, the target adversarial sample generation module 540 can be used to perform the operation S240 described above, which will not be elaborated here.

[0111] According to an embodiment of the present invention, any combination of the perception module 510, the perturbation module 520, the loss data calculation module 530, and the target adversarial sample generation module 540 may be implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the perception module 510, the perturbation module 520, the loss data calculation module 530, and the target adversarial sample generation module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the perception module 510, the perturbation module 520, the loss data calculation module 530, and the target adversarial sample generation module 540 may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.

[0112] Based on the above adversarial testing method for an autonomous driving system, the present invention also provides an adversarial testing device for an autonomous driving system. The following will be combined with Figure 6 to describe the device in detail.

[0113] Figure 6 Fig. shows a structural block diagram of an adversarial testing device for an autonomous driving system according to an embodiment of the present invention.

[0114] As Figure 6 shown, the adversarial testing device 600 for an autonomous driving system according to this embodiment includes an acquisition module 610 and a testing module 620.

[0115] The acquisition module 610 is configured to acquire a target adversarial sample, which is obtained by using an adversarial sample generation method based on autonomous driving multi-sensor fusion. In one embodiment, the acquisition module 610 may be configured to perform the operation S410 described above, which will not be elaborated here.

[0116] The testing module 620 is configured to test the autonomous driving system by using the target adversarial sample to obtain an adversarial testing result of the autonomous driving system, and the adversarial testing result is used to improve the safety of the autonomous driving system. In one embodiment, the testing module 620 may be configured to perform the operation S420 described above, which will not be elaborated here.

[0117] According to an embodiment of the present invention, any plurality of modules among the acquisition module 610 and the test module 620 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 610 and the test module 620 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware or in an appropriate combination of any several of them. Alternatively, at least one of the acquisition module 610 and the test module 620 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0118] Figure 7 The block diagram of an electronic device suitable for implementing an adversarial sample generation method based on multi-sensor fusion for autonomous driving according to an embodiment of the present invention is shown.

[0119] As Figure 7 shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 705 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0120] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.

[0121] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 705 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. The drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage portion 705 as needed.

[0122] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0123] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.

[0124] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program includes program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to cause the computer system to implement the method for generating adversarial samples based on multi-sensor fusion of autonomous driving provided by the embodiments of the present invention.

[0125] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described system, apparatus, module, unit, etc. can be implemented by computer program modules.

[0126] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0127] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0128] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0130] Those skilled in the art will appreciate that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0131] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A method for generating adversarial samples based on multi-sensor fusion for autonomous driving, characterized in that: The method comprises: For each of the multiple perspectives, the three-dimensional object model is sensed to obtain perception data, where the three-dimensional object model is used to simulate obstacles encountered during the autonomous driving process, and the perception data includes visual perception data and radar perception data; Obtaining visual loss data based on the visual perception data and perturbed visual perception data obtained by perturbing the visual perception data, and obtaining radar loss data based on the radar perception data and perturbed radar perception data obtained by perturbing the radar perception data; Based on the candidate adversarial samples generated by the perturbed perception data of the multiple perspectives and the visual perception data, respectively calculating the alignment loss and the smoothness loss to obtain the alignment loss data and the smoothness loss data; Adjust the candidate adversarial sample based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data to generate a target adversarial sample, where the target adversarial sample is used to simulate an obstacle after adding disturbance; The visual loss data is calculated by the following formula: in, represents visual loss data, N represents the total number of viewing angles, represents the color value of the perturbed visual perception data at the i-th viewing angle, Represents the color value of visual perception data at the i-th viewing angle; The radar loss data is calculated by the following formula: in, represents radar loss data, represents the cross entropy loss function, Represents the point cloud prediction mask of the perturbed radar perception data, Point cloud ground truth mask representing radar perception data.

2. The method according to claim 1, characterized in that The adjusting the candidate adversarial sample based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data to generate a target adversarial sample comprises: Obtaining a total loss function value based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data; Using a stochastic gradient descent algorithm, solving the target perturbation data that minimizes the total loss function value; The candidate adversarial sample is adjusted based on the target perturbation data to obtain a target adversarial sample.

3. The method according to claim 1, characterized in that The three-dimensional object model is constructed by the following method: Acquire images and sampling attribute data of the target object at the multiple viewing angles, wherein the sampling attribute data includes a capture position of an image capture device corresponding to different viewing angles and a capture angle relative to the target object; Inputting the plurality of images and the sampled attribute data into a trained volume radiation field model to obtain a volume radiation field of the target object; The three-dimensional object model is obtained based on the volume radiation field and the sampled attribute data.

4. The method according to claim 3, characterized in that The sampling attribute data also includes transmittance, opacity and color of each of the plurality of sampling points, the plurality of sampling points being distributed on a plurality of rays from the optical center of the image acquisition device to the target object, and obtaining the three-dimensional object model based on the volume radiation field and the sampling attribute data includes: For each ray, obtaining a ray rendering color based on the transmittance of the sampling point, the opacity of the sampling point and the color of the sampling point; The volume radiation field is rendered based on the ray rendering colors of each of the multiple rays to obtain the three-dimensional object model.

5. The method according to claim 4, characterized in that The method further comprises: Using a mean square error loss function, a mean square loss value is obtained based on the pixel values ​​of the three-dimensional object model and the pixel values ​​of the target object; The three-dimensional object model is adjusted based on the mean square loss value.

6. The method according to claim 1, characterized in that The alignment loss data is calculated by the following formula: in, represents the alignment loss data, n represents the total number of points in the point cloud formed by the candidate adversarial sample, represents the weight of the kth point, represents the feature of the kth point in the candidate adversarial sample, Represents the features of the kth point in the visual perception data; The smoothed loss data is calculated by the following formula: in, represents the smoothed loss data, represents the pixel value at position (i, j) on the visual perception data, represents the pixel value at the position (i+1, j) on the visual perception data, Represents the pixel value at position (i, j+1) on the visual perception data.

7. A confrontation test method for an autonomous driving system, characterized in that: The method comprises: Obtain a target adversarial sample, wherein the target adversarial sample is obtained by using the method described in any one of claims 1 to 6; The target adversarial sample is used to test the autonomous driving system to obtain an adversarial test result of the autonomous driving system.

8. An adversarial sample generation device based on multi-sensor fusion for autonomous driving, characterized in that: The device comprises: a perception module, configured to perceive the three-dimensional object model for each of the multiple perspectives to obtain perception data, wherein the three-dimensional object model is used to simulate obstacles encountered during the autonomous driving process, and the perception data includes visual perception data and radar perception data; a perturbation module, configured to obtain visual loss data based on the visual perception data and perturbed visual perception data obtained by perturbing the visual perception data, and to obtain radar loss data based on the radar perception data and perturbed radar perception data obtained by perturbing the radar perception data; A loss data calculation module, configured to calculate alignment loss and smoothing loss based on candidate adversarial samples generated by the perturbation perception data of the multiple perspectives and the visual perception data, respectively, to obtain alignment loss data and smoothing loss data; A target adversarial sample generation module, configured to adjust the candidate adversarial sample based on the visual loss data, the radar loss data, the alignment loss data, and the smoothing loss data to generate a target adversarial sample, wherein the target adversarial sample is used to simulate an obstacle after adding disturbance; The visual loss data is calculated by the following formula: in, represents visual loss data, N represents the total number of viewing angles, represents the color value of the perturbed visual perception data at the i-th viewing angle, Represents the color value of visual perception data at the i-th viewing angle; The radar loss data is calculated by the following formula: in, represents radar loss data, represents the cross entropy loss function, Represents the point cloud prediction mask of the perturbed radar perception data, Point cloud ground truth mask representing radar perception data.

9. An adversarial test device for an autonomous driving system, characterized in that: The device comprises: An acquisition module, used to acquire a target adversarial sample, wherein the target adversarial sample is obtained by using the method described in any one of claims 1 to 6; The testing module is used to test the autonomous driving system using the target adversarial sample to obtain the adversarial test result of the autonomous driving system.

Citation Information

Patent Citations

  • Adversarial sample generation method and system for image segmentation

    CN117788830A

  • Target recognition model detection method and device, equipment, medium and program product

    CN118968470A