A three-dimensional automatic driving physical confrontation generation method based on robustness evaluation

CN122780488APending Publication Date: 2026-09-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610776726.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于鲁棒性评估的三维自动驾驶物理对抗生成方法,通过引入扩散模型的先验纹理约束与全链路可导逆向渲染机制,从根本上解决了现有对抗补丁隐蔽性差与空间一致性弱的问题,能够生成既具备卓越物理鲁棒性,又在视觉上与自然场景无缝融合的高质量对抗补丁

Benefits of technology

[0015] This invention presents a 3D autonomous driving physics adversarial generation method based on robustness evaluation, overcoming the limitations of traditional adversarial attacks that rely on pixel-level gradient optimization and are prone to generating high-frequency, singular noise. By introducing a diffusion model as a prior constraint and combining fractional distillation sampling with truncated time steps, the generated physics adversarial patch is forced to highly fit the real road background texture (such as road cracks, stains, etc.) in probability distribution, achieving seamless integration with the natural scene and greatly reducing the probability of being detected by human drivers and conventional anomaly detection systems. This invention designs a fully differentiable 3D multi-view inverse perspective projection operator. This operator not only ensures the spatial coherence of the adversarial patch in the 3D physical world but also opens up a direct gradient backpropagation path from 2D perception results to the 3D generation space. By combining tensor-quantized physical environment perturbations and seamless gradient propagation, this invention effectively overcomes the "gradient breakage" problem caused by rendering simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780488A_ABST
    Figure CN122780488A_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional automatic driving physical confrontation generation method based on robustness evaluation, constructs a three-dimensional to multi-view inverse projection rendering operator based on camera projection relationship and data enhancement inverse transformation, fuses patch texture corresponding to potential features into multi-view images to obtain enhanced observation data, inputs the enhanced observation data into a frozen automatic driving BEV perception model, calculates an adversarial target loss for robustness evaluation, takes a diffusion model as texture prior, calculates a prior texture distribution loss of the enhanced observation data through fractional distillation sampling, and jointly updates the adversarial target loss and the prior texture distribution loss to only update the learnable potential features, and generates a physical test patch with road texture appearance and satisfying three-dimensional multi-view consistency. The application can generate natural and stable robustness evaluation samples in a closed test or simulation environment, and provides technical support for safety verification and reinforcement of an automatic driving perception model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving safety testing and computer vision technology, and in particular to a method for generating 3D autonomous driving physical adversarial relationships based on robustness assessment. Background Technology

[0002] With the rapid development of deep learning technology, BEV perception models based on multi-camera fusion have become a core component of mainstream autonomous driving systems. Deep neural networks, as the core of perception, are extremely vulnerable and susceptible to adversarial attacks, leading to the risk of missed detections or false positives.

[0003] Early adversarial attack research was largely confined to two-dimensional digital space, inducing model errors by directly adding minute pixel-level perturbations to digital input images. However, in real-world open-road autonomous driving scenarios, attackers typically cannot directly intercept or tamper with the vehicle's underlying sensor data streams. Therefore, adversarial patches must be deployed as physical entities in the real three-dimensional physical world. While some physical adversarial generation techniques have emerged in recent years, traditional physical adversarial methods often suffer from the following significant limitations when directly applied to such complex multi-view BEV perception systems, due to the reliance on joint observations from multiple cameras and perspectives, along with 3D-to-2D projection fusion mechanisms: Poor visual concealment: Gradient-optimized physical patches often produce high-frequency, unnatural textures, which are easily recognized by human drivers or intercepted by the system's anomaly detection module; Insufficient spatial consistency and robustness: BEV perception involves spatial overlap between multiple viewpoints. If traditional projection methods are not differentiable or simplify physical modeling, the attack effectiveness of the generated patches will decrease significantly under different lighting and viewpoints, making it impossible to achieve stable cross-viewpoint attacks; Gradient breakage problem: When simulating physical environments (such as lighting and occlusion), if a fully differentiable rendering operator is not used, the gradient flow will not be effectively transmitted to the features to be optimized, leading to convergence difficulties. Summary of the Invention

[0004] This invention provides a 3D autonomous driving physics adversarial generation method based on robustness assessment. By introducing prior texture constraints of diffusion model and full-link differentiable inverse rendering mechanism, it fundamentally solves the problems of poor concealment and weak spatial consistency of existing adversarial patches. It can generate high-quality adversarial patches that have excellent physical robustness and are visually seamlessly integrated with natural scenes.

[0005] A first aspect of this invention provides a method for generating 3D autonomous driving physics adversarial relationships based on robustness assessment, comprising the following steps: Step S1: Obtain a set of multi-view images of the autonomous driving scene at the same timestamp. and the corresponding 3D scene parameters ,in, For the number of cameras, For the first The intrinsic parameter matrix of each camera, For the first The extrinsic parameter matrix of a camera from the real physical coordinate system to the camera coordinate system. Spatial data augmentation matrix used in the preprocessing stage of the perception model for autonomous driving BEVs; Step S2: Initialize the 3D position parameters of the adversarial patch in the real physical coordinate system. and learnable latent feature representation The three-dimensional position parameters At least include the center point coordinates of the adversarial patch. Physical length and width and the normal vector of the plane containing the adversarial patch ; Step S3, using the diffusion model Represent learnable latent features Decoded as anti-patch texture Based on 3D scene parameters and three-dimensional position parameters Constructing a fully differentiable 3D multi-view inverse perspective projection operator And combined with tensor-quantized differentiable physical environment perturbations , will resist patch texture Rendering and fusion into a multi-view image collection In this process, enhanced observational data were obtained. ; Step S4 will enhance the observation data. Input to the frozen autonomous driving BEV perception model to be evaluated In the process, the adversarial target loss used for robustness assessment is calculated. ; Step S5, the learnable latent feature representation Input to diffusion model In this study, fractional distillation sampling is used to calculate the prior texture distribution loss that constrains the realism of adversarial patch textures. ; Step S6, Jointly combat target losses and prior texture distribution loss Construct the total loss function Backpropagation using an automatic differentiation framework updates only the learnable latent feature representation. Iteratively generate physics-based adversarial patches with realistic texture appearance.

[0006] In one embodiment of the present invention, the three-dimensional scene parameters in step S1 It also includes image size, camera time synchronization information, and the 3D detection bounding box of the target to be evaluated. and target category When the 3D detection box of the target to be evaluated Located in the spatial data augmentation matrix In the resulting coordinate system, based on the spatial data enhancement matrix The inverse transform will transform the 3D detection box Restored to the true physical coordinate system, the restored 3D detection box is obtained. .

[0007] In one embodiment of the present invention, step S2 specifically includes: Based on the selected target camera viewpoint and the 3D detection box restored to the true physical coordinate system. Within a preset physical distance range, sample the coordinates of candidate center points of adversarial patches. and candidate orientation angle; Based on candidate center point coordinates Candidate orientation angle and preset physical length and width Candidate 3D regions are identified and augmented based on the spatial data matrix. Transform the candidate 3D region to the enhanced coordinate system; In the enhanced coordinate system, the rotation intersection-union ratio of the candidate 3D region and the existing real target 3D detection box in the bird's-eye view plane is calculated, and when the preset avoidance conditions are met, the position parameters corresponding to the candidate 3D region are determined as the 3D position parameters of the adversarial patch. The preset avoidance condition is that the candidate 3D region has no spatial overlap with existing real targets in the scene; the initial image of the road background is input into the encoder of the diffusion model to obtain initial latent features, and the initial latent features are used as learnable latent feature representations. The initial value.

[0008] In one embodiment of the present invention, step S3 specifically includes: Step S301: Obtain the three-dimensional position parameters of the adversarial patch. and spatial data augmentation matrix and the three-dimensional position parameters Unify to the real physical coordinate system; Step S302: In the real physical coordinate system, based on the three-dimensional position parameters of the adversarial patch... Physical length and width and the normal vector of the plane containing the adversarial patch Calculate the set of coordinates of the four three-dimensional corner points of the adversarial patch. ; Step S303, combine the intrinsic parameter matrix of the i-th camera. and extrinsic parameter matrix The set of three-dimensional corner coordinates Projecting the image onto the two-dimensional image plane corresponding to the i-th camera yields a set of two-dimensional corner coordinates. The projection relationship is expressed as: in, The coordinates of a single three-dimensional corner point in the real physical coordinate system. These are the projected two-dimensional pixel coordinates. Let be the depth of field value in the coordinate system of the i-th camera; Step S304, based on the set of two-dimensional corner coordinates Determine the two-dimensional mask region from the i-th viewpoint. Traverse the two-dimensional mask region The pixel coordinates within the plane are then mapped inversely to the adversarial patch plane in three-dimensional space. Step S305, representing the learnable latent features within the adversarial patch plane. The corresponding adversarial patch texture A is bilinearly color sampled, and the sampled texture color is combined with a tensor-quantized differentiable physical environment perturbation. The data is fused into the image from the corresponding viewpoint to obtain enhanced observation data. .

[0009] In one embodiment of the present invention, in step S3, the tensor-quantized differentiable physical environment perturbation Includes differentiable illumination gain transformation, Gaussian blur transformation, and color dithering transformation; the tensor-quantized differentiable physical environment perturbation Enhanced observation data from multiple perspectives is applied to the fitted pixels of the two-dimensional mask region. The formula for cascading rendering is expressed as: in, Represents the two-dimensional projected pixel coordinates. For the first The enhanced observation image corresponding to each camera in pixel coordinates Pixel value at that location, For the first Two-dimensional mask area of ​​each camera, This represents the bilinear color sampling function. To combat patch textures, This represents the inverse mapping operator from pixel coordinates in a two-dimensional image to the three-dimensional patch plane. For the first The original image corresponding to each camera in pixel coordinates The pixel value at that location.

[0010] In one embodiment of the present invention, step S4 specifically includes: Freeze the perception model of the autonomous driving BEV to be evaluated The model parameters will enhance the observed data. Input autonomous driving BEV perception model , thus obtaining the target detection output of the autonomous driving BEV perception model; Based on the three-dimensional detection frame of the target to be evaluated and target category Construct an adversarial target, and calculate the adversarial target loss based on at least one of the following: classification difference, 3D bounding box regression difference, and orientation difference between the target detection output and the adversarial target. This causes the target detection output to deviate from the actual detection result.

[0011] In one embodiment of the present invention, step S5 specifically includes: Step S501: Randomly generate target noise and diffusion time step , among which target noise It follows a standard normal distribution; it is represented by learnable latent features according to the noise scheduling rules of the diffusion model. Add the target noise To obtain the latent features of noise addition ; Step S502, add noisy potential features diffusion time step and text prompt words input to the diffusion model In the middle, predictive conditional noise ; Step S503, based on the predicted conditional noise With the target noise The difference between them is used to construct a learningable latent feature representation. Updated fractional distillation gradient , is represented as: in, To match the diffusion time step Relevant weighting coefficients; Introducing the gradient cutoff operator To block gradient flow into the diffusion model, a differentiable prior texture distribution loss is constructed through dot product summation. The formula is expressed as: in, This represents the dot product between tensor elements.

[0012] In one embodiment of the present invention, step S6 specifically includes: Based on the losses of the opposing target and prior texture distribution loss Construct the total loss function ,in, This is the balance coefficient; Freeze the perception model of autonomous driving BEV to be evaluated and diffusion model The model parameters are backpropagated through an automatic differentiation framework to obtain the total loss function. Only update the learnable latent feature representation. And the updated learnable latent features are represented Decoded as a physical combat patch.

[0013] A second aspect of this invention proposes a three-dimensional autonomous driving physics adversarial generation system based on robustness assessment, comprising: The data acquisition module is used to acquire a set of multi-view images of an autonomous driving scene at the same time stamp and the corresponding 3D scene parameters; The spatial initialization module is used to initialize the 3D position parameters and learnable latent feature representations of the adversarial patch in a real physical coordinate system. The end-to-end differentiable rendering module is used to decode learnable latent feature representations into adversarial patch textures using a diffusion model. It constructs an end-to-end differentiable 3D multi-view inverse perspective projection operator based on 3D scene parameters and 3D position parameters, and combines tensor-quantized differentiable physical environment perturbation to fuse adversarial patch texture rendering into a multi-view image set to obtain enhanced observation data. The adversarial deception perception module is used to input enhanced observation data into the frozen perception model of the autonomous driving BEV to be evaluated and calculate the adversarial target loss for robustness assessment. The diffusion fractional distillation module is used to input learnable latent feature representations into the diffusion model and use fractional distillation sampling to calculate the prior texture distribution loss that constrains the realism of patch textures. The joint convergence optimization module is used to jointly construct the total loss function from the adversarial target loss and the prior texture distribution loss. Through backpropagation using an automatic differentiation framework, it updates only the learnable latent feature representations and iteratively generates physical adversarial patches with realistic texture appearance.

[0014] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the three-dimensional autonomous driving physics adversarial generation method based on robustness evaluation described in the above embodiments.

[0015] This invention presents a 3D autonomous driving physics adversarial generation method based on robustness evaluation, overcoming the limitations of traditional adversarial attacks that rely on pixel-level gradient optimization and are prone to generating high-frequency, singular noise. By introducing a diffusion model as a prior constraint and combining fractional distillation sampling with truncated time steps, the generated physics adversarial patch is forced to highly fit the real road background texture (such as road cracks, stains, etc.) in probability distribution, achieving seamless integration with the natural scene and greatly reducing the probability of being detected by human drivers and conventional anomaly detection systems. This invention designs a fully differentiable 3D multi-view inverse perspective projection operator. This operator not only ensures the spatial coherence of the adversarial patch in the 3D physical world but also opens up a direct gradient backpropagation path from 2D perception results to the 3D generation space. By combining tensor-quantized physical environment perturbations and seamless gradient propagation, this invention effectively overcomes the "gradient breakage" problem caused by rendering simulation.

[0016] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A schematic diagram of a three-dimensional autonomous driving physics adversarial generation method based on robustness assessment provided by the present invention; Figure 2 This invention provides an architecture diagram for a three-dimensional autonomous driving physics adversarial generation method based on robustness evaluation. Detailed Implementation

[0018] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0019] This novel 3D autonomous driving physics adversarial generation method, based on robustness assessment, aims to overcome the shortcomings of traditional adversarial attacks, such as easy detection by humans, inability to stably cope with multi-view joint observations, and gradient breakage in physical modeling. By introducing a diffusion model-driven prior texture constraint mechanism and a fully differentiable inverse physical rendering framework, combined with fractional distillation sampling for latent feature optimization, it achieves efficient and covert physical attacks against BEV perception systems, significantly improving the natural visual fit and multi-view robustness of adversarial patches in the physical world. Figure 1 and Figure 2 As shown, the specific steps include the following: Step S1: Obtain a set of multi-view images of the autonomous driving scene at the same timestamp. and the corresponding 3D scene parameters ,in, For the number of cameras, For the first The intrinsic parameter matrix of each camera, For the first The extrinsic parameter matrix of a camera from the world coordinate system to the camera coordinate system. Spatial data augmentation matrix used in the preprocessing stage of the perception model for autonomous driving BEVs.

[0020] In an embodiment of the present invention, three-dimensional scene parameters It also includes image size, camera time synchronization information, and the 3D detection bounding box of the target to be evaluated. and target category When the 3D detection box of the target to be evaluated Located in the spatial data augmentation matrix In the resulting coordinate system, according to the inverse transformation of the spatial data enhancement matrix... The three-dimensional detection box Restored to the true physical coordinate system, the restored 3D detection box is obtained. .

[0021] In practice, the surround-view camera array mounted on the autonomous vehicle acquires a collection of multi-view images at the same time stamp. , For the number of cameras (e.g.) (Representing a typical configuration of six cameras: front, rear, left front, right front, left rear, and right rear), synchronously reading the 3D scene parameters at that timestamp. .

[0022] Simultaneously, extract the perception model of the autonomous driving BEV to be evaluated. Spatial data augmentation matrices are used in the data preprocessing stage of pre-trained models such as BEVFormer or CenterPoint. (Including affine transformation operations such as random rotation, scaling, and translation). Furthermore, the original 3D bounding box of the target to be evaluated (such as a specific vehicle or pedestrian) in the perceptual coordinate system after spatial data augmentation is obtained. and target category .

[0023] Step S2: Initialize the 3D position parameters of the adversarial patch in the real physical coordinate system. and learnable latent feature representation The three-dimensional position parameters At least include the center point coordinates of the adversarial patch. Physical length and width and the normal vector of the plane containing the adversarial patch .

[0024] In an embodiment of the present invention, based on the selected target camera viewpoint and the restored 3D detection box... Within a preset physical distance range, sample the candidate center point coordinates and candidate orientation angles of the adversarial patch; the preset physical distance range is preferably 7m to 10m. Based on the candidate center point coordinates... Candidate orientation angle and preset physical length and physical width Candidate 3D regions are determined, with a preferred physical length of 4m and a preferred physical width of 2m, and are then augmented based on the spatial data matrix. The candidate 3D region is transformed to an enhanced coordinate system. In this enhanced coordinate system, the rotation intersection-union ratio (ROI) of the candidate 3D region and the existing real target 3D detection bounding box in the bird's-eye view plane is calculated. When a preset avoidance condition is met, the position parameters corresponding to the candidate 3D region are determined as the 3D position parameters of the adversarial patch. The initial image of the road background is input into the encoder of the fine-tuned diffusion model to obtain initial latent features, which are then used as learnable latent feature representations. The initial value.

[0025] Since the perception model performs data augmentation transformations on the input image or 3D spatial features during inference, this embodiment first performs an inverse transformation based on the spatial data augmentation matrix to ensure that the generated physics adversarial patch can be accurately deployed in the real world. 3D detection box Restored to the true physical coordinate system, the restored 3D detection box is obtained. .

[0026] Based on the restored 3D detection box The system samples the candidate center point coordinates and candidate orientation angles of adversarial patches within a preset physical distance range. To avoid physical collisions or overlaps between the generated patches and real targets in 3D space, the system uses the candidate center point coordinates, candidate orientation angles, and preset physical length as parameters. and width Determine candidate 3D regions and utilize Transform it to the enhanced coordinate system.

[0027] In the augmented coordinate system, the rotational intersection-union ratio (ROI) between the candidate 3D region and the existing real target 3D bounding boxes in the BEV plane is calculated. When the ROI is less than a preset threshold (e.g., 0.05), the avoidance condition is deemed met, and the position parameters at this point are established as the 3D position parameters of the adversarial patch. .

[0028] Meanwhile, to overcome the limitations of traditional pixel-level optimization, the initial image of the road background in the target area is input into the diffusion model. The encoder obtains the initial latent features and uses them as a learnable latent feature representation. The initial value.

[0029] Step S3, using the diffusion model Represent learnable latent features Mapped as adversarial patch texture Based on 3D scene parameters and three-dimensional position parameters Constructing a fully differentiable 3D multi-view inverse perspective projection operator And combined with tensor-quantized differentiable physical environment perturbations , will resist patch texture Rendering and fusion into a multi-view image collection In this process, enhanced observational data were obtained. .

[0030] In embodiments of the present invention, the aim is to include potential features Corresponding patch texture Seamless and differentiable rendering to multi-view images solves the gradient breakage problem caused by traditional physics simulations. Step S3 specifically includes: Step S301: Obtain the three-dimensional position parameters of the adversarial patch. and spatial data augmentation matrix and the three-dimensional position parameters Unify to the real physical coordinate system; Step S302: In the real physical coordinate system, based on the three-dimensional position parameters of the adversarial patch... Physical length and width and the normal vector of the plane containing the adversarial patch The set of coordinates of the four three-dimensional corner points of the adversarial patch is calculated through geometric analysis. ; Step S303, combine the intrinsic parameter matrix of the i-th camera. and extrinsic parameter matrix The set of three-dimensional corner coordinates Projecting the image onto the two-dimensional image plane corresponding to the i-th camera yields a set of two-dimensional corner coordinates. The projection relationship is expressed as: in, The coordinates of a single three-dimensional corner point in the real physical coordinate system. These are the projected two-dimensional pixel coordinates. Let be the depth of field value in the coordinate system of the i-th camera; Step S304, based on the set of two-dimensional corner coordinates Determine the two-dimensional mask region from the i-th viewpoint. Traverse the two-dimensional mask region The pixel coordinates within the plane are then mapped inversely to the adversarial patch plane in three-dimensional space. Step S305, representing the learnable latent features within the adversarial patch plane. The corresponding adversarial patch texture A is bilinearly color sampled, and the sampled texture color is combined with a tensor-quantized differentiable physical environment perturbation. The data is fused into the image from the corresponding viewpoint to obtain enhanced observation data. .

[0031] Tensor-quantized differentiable physical environment perturbation Includes differentiable illumination gain transformation, Gaussian blur transformation, and color dithering transformation; the tensor-quantized differentiable physical environment perturbation Enhanced observation data from multiple perspectives is applied to the fitted pixels of the two-dimensional mask region. The formula for cascading rendering is expressed as: in, Represents the two-dimensional projected pixel coordinates. For the first The enhanced observation image corresponding to each camera in pixel coordinates Pixel value at that location, For the first Two-dimensional mask area of ​​each camera, This represents the bilinear color sampling function. To combat patch textures, This represents the inverse mapping operator from pixel coordinates in a two-dimensional image to the three-dimensional patch plane. For the first The original image corresponding to each camera in pixel coordinates The pixel value at that location.

[0032] Step S4 will enhance the observation data. Input to the frozen autonomous driving BEV perception model to be evaluated In the process, the adversarial target loss used for robustness assessment is calculated. .

[0033] In robustness evaluation, adversarial target loss is used to characterize a decrease in target detection confidence or class shift.

[0034] In an embodiment of the present invention, step S4 specifically includes: Freeze the perception model of the autonomous driving BEV to be evaluated The model parameters will enhance the observed data. Input autonomous driving BEV perception model The target detection output of the autonomous driving BEV perception model is obtained; based on the 3D detection box of the target to be evaluated... and target category Construct an adversarial target, and calculate the adversarial target loss based on at least one of the following: classification difference, 3D bounding box regression difference, and orientation difference between the target detection output and the adversarial target. This causes the target detection output to deviate from the actual detection result.

[0035] The losses to the opposing target are: in, To force a decrease in target confidence level, To induce a 3D bounding box regression loss for localization offset, To predict the loss. This is a preset loss weight. By minimizing this loss, the perception model is made to produce either missed detections or false alarms.

[0036] Step S5, the learnable latent feature representation Input to diffusion model In this study, fractional distillation sampling (SDS) is used to compute the prior texture distribution loss that constrains the realism of adversarial patch textures. .

[0037] In an embodiment of the present invention, to give the anti-patch a highly natural physical appearance similar to road cracks or stains, this embodiment introduces a fine-tuned diffusion model. As a priori, in the optimization phase, step S5 specifically includes: Step S501: Randomly generate target noise and diffusion time step , among which target noise It follows a standard normal distribution; it is represented by learnable latent features according to the noise scheduling rules of the diffusion model. Add the target noise To obtain the latent features of noise addition ; Step S502, add noisy potential features diffusion time step and text prompt words input to the diffusion model For example, "generating a highly realistic asphalt pavement with faint tire marks, soft shadows, and fine cracks," predicting conditional noise. ; Step S503, based on the predicted conditional noise With the target noise The difference between them is used to construct a learningable latent feature representation. Updated fractional distillation gradient , is represented as: in, To match the diffusion time step Relevant weighting coefficients; To avoid memory overflow caused by calculating the second derivative of the gradient, a gradient truncation operator is introduced. To block gradient flow into the diffusion model, constraints are constructed using dot product summation. Differentiable prior texture distribution loss The formula is expressed as: in, This represents the dot product between tensor elements.

[0038] Ultimately, the gradient of the natural road surface features corresponding to backpropagation can be directly propagated to the learnable latent feature representation. .

[0039] Step S6, Jointly combat target losses and prior texture distribution loss Construct the total loss function Backpropagation using an automatic differentiation framework updates only the learnable latent feature representation. Iteratively generate physics-based adversarial patches with realistic texture appearance.

[0040] In an embodiment of the present invention, step S6 specifically includes: Based on the losses of the opposing target and prior texture distribution loss Construct the total loss function ,in, This is the balance coefficient; Freeze the perception model of autonomous driving BEV to be evaluated and diffusion model The model parameters are backpropagated through an automatic differentiation framework to obtain the total loss function. Only update the learnable latent feature representation. And the updated learnable latent features are represented Decoded as a physical combat patch.

[0041] In the automatic differentiation framework, backpropagation is performed using the Adam optimizer. Because and The gradient of the computation graph has been frozen, and it will stably propagate through the inverse perspective projection operator to the single independent variable—the learnable latent feature representation. Iterative updates Until the loss converges, the optimized... Decoding yields an image of adversarial patches that can be deployed on real physical surfaces.

[0042] This invention also proposes a three-dimensional autonomous driving physics adversarial generation system based on robustness evaluation, comprising: The data acquisition module is used to acquire a set of multi-view images of an autonomous driving scene at the same time stamp and the corresponding 3D scene parameters; The spatial initialization module is used to initialize the 3D position parameters and learnable latent feature representations of the adversarial patch in a real physical coordinate system. The end-to-end differentiable rendering module is used to decode learnable latent feature representations into adversarial patch textures using a diffusion model. It constructs an end-to-end differentiable 3D multi-view inverse perspective projection operator based on 3D scene parameters and 3D position parameters, and combines tensor-quantized differentiable physical environment perturbation to fuse adversarial patch texture rendering into a multi-view image set to obtain enhanced observation data. The adversarial deception perception module is used to input enhanced observation data into the frozen perception model of the autonomous driving BEV to be evaluated and calculate the adversarial target loss for robustness assessment. The diffusion fractional distillation module is used to input learnable latent feature representations into the diffusion model and use fractional distillation sampling to calculate the prior texture distribution loss that constrains the realism of patch textures. The joint convergence optimization module is used to jointly construct the total loss function from the adversarial target loss and the prior texture distribution loss. Through backpropagation using an automatic differentiation framework, it updates only the learnable latent feature representations and iteratively generates physical adversarial patches with realistic texture appearance.

[0043] This invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the three-dimensional autonomous driving physics adversarial generation method based on robustness evaluation described in the above embodiments.

[0044] This invention proposes a method and system for generating 3D autonomous driving physical adversarial pairs based on robustness assessment. This method targets BEV perception models based on multi-camera input, generating physical test patches with road background texture appearance, 3D physical position constraints, and multi-view projection consistency during closed testing. The method involves acquiring synchronous multi-view observation data and corresponding 3D scene parameters for the autonomous driving scenario; setting patch geometric parameters in the real physical coordinate system and initializing learnable latent features; constructing a fully differentiable 3D-to-multi-view inverse projection rendering operator based on camera projection relationships and data augmentation inverse transformation; and combining differentiable illumination, blur, and color perturbations to fuse the patch texture corresponding to the latent features into the multi-view image to obtain augmented observation data; inputting the augmented observation data into the frozen autonomous driving BEV perception model to calculate the adversarial target loss for robustness assessment; using a diffusion model fine-tuned from road texture data as a texture prior, calculating the prior texture distribution loss through fractional distillation sampling; and jointly updating only the learnable latent features using the adversarial target loss and the prior texture distribution loss to generate physical test patches with road texture appearance and satisfying 3D multi-view consistency. This invention can generate natural, stable, and robust evaluation samples in closed testing or simulation environments, providing technical support for the safety verification and hardening of autonomous driving perception models.

[0045] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0046] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for generating 3D autonomous driving physics adversarial relationships based on robustness evaluation, characterized in that, Includes the following steps: Step S1: Obtain a set of multi-view images of the autonomous driving scene at the same timestamp. and the corresponding 3D scene parameters ,in, For the number of cameras, For the first The intrinsic parameter matrix of each camera, For the first The extrinsic parameter matrix of a camera from the real physical coordinate system to the camera coordinate system. Spatial data augmentation matrix used in the preprocessing stage of the perception model for autonomous driving BEVs; Step S2: Initialize the 3D position parameters of the adversarial patch in the real physical coordinate system. and learnable latent feature representation The three-dimensional position parameters At least include the center point coordinates of the adversarial patch. Physical length and width and the normal vector of the plane containing the adversarial patch ; Step S3, using the diffusion model Represent learnable latent features Decoded as anti-patch texture Based on 3D scene parameters and three-dimensional position parameters Constructing a fully differentiable 3D multi-view inverse perspective projection operator And combined with tensor-quantized differentiable physical environment perturbations , will resist patch texture Rendering and fusion into a multi-view image collection In this process, enhanced observational data were obtained. ; Step S4 will enhance the observation data. Input to the frozen autonomous driving BEV perception model to be evaluated In the process, the adversarial target loss used for robustness assessment is calculated. ; Step S5, the learnable latent feature representation Input to diffusion model In this study, fractional distillation sampling is used to calculate the prior texture distribution loss that constrains the realism of adversarial patch textures. ; Step S6, Jointly combat target losses and prior texture distribution loss Construct the total loss function Backpropagation using an automatic differentiation framework updates only the learnable latent feature representation. Iteratively generate physics-based adversarial patches with realistic texture appearance.

2. The method according to claim 1, characterized in that, 3D scene parameters in step S1 It also includes image size, camera time synchronization information, and the 3D detection bounding box of the target to be evaluated. and target category ; When the 3D detection box of the target to be evaluated Located in the spatial data augmentation matrix In the resulting coordinate system, based on the spatial data enhancement matrix The inverse transform will transform the 3D detection box Restored to the true physical coordinate system, the restored 3D detection box is obtained. .

3. The method according to claim 1, characterized in that, Step S2 specifically includes: Based on the selected target camera viewpoint and the 3D detection box restored to the true physical coordinate system. Within a preset physical distance range, sample the coordinates of candidate center points of adversarial patches. and candidate orientation angle; Based on candidate center point coordinates Candidate orientation angle and preset physical length and width Candidate 3D regions are identified and augmented based on the spatial data matrix. Transform the candidate 3D region to the enhanced coordinate system; In the enhanced coordinate system, the rotation intersection-union ratio of the candidate 3D region and the existing real target 3D detection box in the bird's-eye view plane is calculated, and when the preset avoidance conditions are met, the position parameters corresponding to the candidate 3D region are determined as the 3D position parameters of the adversarial patch. The preset avoidance condition is that the candidate 3D region has no spatial overlap with existing real targets in the scene; the initial image of the road background is input into the encoder of the diffusion model to obtain initial latent features, and the initial latent features are used as learnable latent feature representations. The initial value.

4. The method according to claim 1, characterized in that, Step S3 specifically includes: Step S301: Obtain the three-dimensional position parameters of the adversarial patch. and spatial data augmentation matrix and the three-dimensional position parameters Unify to the real physical coordinate system; Step S302: In the real physical coordinate system, based on the three-dimensional position parameters of the adversarial patch... Physical length and width and the normal vector of the plane containing the adversarial patch Calculate the set of coordinates of the four three-dimensional corner points of the adversarial patch. ; Step S303, combine the intrinsic parameter matrix of the i-th camera. and extrinsic parameter matrix The set of three-dimensional corner coordinates Projecting the image onto the two-dimensional image plane corresponding to the i-th camera yields a set of two-dimensional corner coordinates. The projection relationship is expressed as: in, The coordinates of a single three-dimensional corner point in the real physical coordinate system. These are the projected two-dimensional pixel coordinates. Let be the depth of field value in the coordinate system of the i-th camera; Step S304, based on the set of two-dimensional corner coordinates Determine the two-dimensional mask region from the i-th viewpoint. Traverse the two-dimensional mask region The pixel coordinates within the plane are then mapped inversely to the adversarial patch plane in three-dimensional space. Step S305, representing the learnable latent features within the adversarial patch plane. The corresponding adversarial patch texture A is bilinearly color sampled, and the sampled texture color is combined with a tensor-quantized differentiable physical environment perturbation. The data is fused into the image from the corresponding viewpoint to obtain enhanced observation data. .

5. The method according to claim 4, characterized in that, In step S3, the tensor-quantized differentiable physical environment perturbation It includes guided illumination gain transformation, Gaussian blur transformation, and color dithering transformation; The tensor-quantized differentiable physical environment perturbation Enhanced observation data from multiple perspectives is applied to the fitted pixels of the two-dimensional mask region. The formula for cascading rendering is expressed as: in, Represents the two-dimensional projected pixel coordinates. For the first The enhanced observation image corresponding to each camera in pixel coordinates Pixel value at that location, For the first Two-dimensional mask area of ​​each camera, This represents the bilinear color sampling function. To combat patch textures, This represents the inverse mapping operator from pixel coordinates in a two-dimensional image to the three-dimensional patch plane. For the first The original image corresponding to each camera in pixel coordinates The pixel value at that location.

6. The method according to claim 1, characterized in that, Step S4 specifically includes: Freeze the perception model of the autonomous driving BEV to be evaluated The model parameters will enhance the observed data. Input autonomous driving BEV perception model , thus obtaining the target detection output of the autonomous driving BEV perception model; Based on the three-dimensional detection frame of the target to be evaluated and target category Construct an adversarial target, and calculate the adversarial target loss based on at least one of the following: classification difference, 3D bounding box regression difference, and orientation difference between the target detection output and the adversarial target. This causes the target detection output to deviate from the actual detection result.

7. The method according to claim 1, characterized in that, Step S5 specifically includes: Step S501: Randomly generate target noise and diffusion time step , among which target noise It follows a standard normal distribution; it is represented by learnable latent features according to the noise scheduling rules of the diffusion model. Add the target noise To obtain the latent features of noise addition ; Step S502, add noisy potential features diffusion time step and text prompt words input to the diffusion model In the middle, predictive conditional noise ; Step S503, based on the predicted conditional noise With the target noise The difference between them is used to construct a learningable latent feature representation. Updated fractional distillation gradient , is represented as: in, To match the diffusion time step Relevant weighting coefficients; Introducing the gradient cutoff operator To block gradient flow into the diffusion model, a differentiable prior texture distribution loss is constructed through dot product summation. The formula is expressed as: in, This represents the dot product between tensor elements.

8. The method according to claim 1, characterized in that, Step S6 specifically includes: Based on the losses of the opposing target and prior texture distribution loss Construct the total loss function ,in, This is the balance coefficient; Freeze the perception model of autonomous driving BEV to be evaluated and diffusion model The model parameters are backpropagated through an automatic differentiation framework to obtain the total loss function. Only update the learnable latent feature representation. And the updated learnable latent features are represented Decoded as a physical combat patch.

9. A three-dimensional autonomous driving physics adversarial generation system based on robustness evaluation, characterized in that, include: The data acquisition module is used to acquire a set of multi-view images of an autonomous driving scene at the same time stamp and the corresponding 3D scene parameters; The spatial initialization module is used to initialize the 3D position parameters and learnable latent feature representations of the adversarial patch in a real physical coordinate system. The end-to-end differentiable rendering module is used to decode learnable latent feature representations into adversarial patch textures using a diffusion model. It constructs an end-to-end differentiable 3D multi-view inverse perspective projection operator based on 3D scene parameters and 3D position parameters, and combines tensor-quantized differentiable physical environment perturbation to fuse adversarial patch texture rendering into a multi-view image set to obtain enhanced observation data. The adversarial deception perception module is used to input enhanced observation data into the frozen perception model of the autonomous driving BEV to be evaluated and calculate the adversarial target loss for robustness assessment. The diffusion fractional distillation module is used to input learnable latent feature representations into the diffusion model and use fractional distillation sampling to calculate the prior texture distribution loss that constrains the realism of patch textures. The joint convergence optimization module is used to jointly construct the total loss function from the adversarial target loss and the prior texture distribution loss. Through backpropagation using an automatic differentiation framework, it updates only the learnable latent feature representations and iteratively generates physical adversarial patches with realistic texture appearance.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the steps of a three-dimensional autonomous driving physics adversarial generation method based on robustness assessment as described in any one of claims 1 to 8.