Method and apparatus for physical adversarial attack in three-dimensional detection
Patent Information
- Application Number
- CN202611038338.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-13
AI Technical Summary
[0005]本发明提出了一种用于三维检测中物理对抗攻击方法及装置,以解决现有面向三维检测器的对抗攻击缺乏跨模态统一的对抗表示,分别对图像与点云施加的扰动无法对应到同一物理物体,导致物理可实现性受限与跨模态一致性不足;同时,在异质属性上分散的优化信号与固定位姿下的过拟合,使得对抗物体在跨场景与跨位姿条件下攻击效果显著退化的技术问题
1)本发明中由于相机观测与激光雷达观测由同一组共享三维高斯基元经统一的坐标变换实例化得到,对抗物体在两种模态下天然保持几何一致,避免了分别扰动两个模态所引入的不一致,提升了攻击的物理合理性。
Smart Images

Figure CN122574830B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence security technology, specifically to a method and apparatus for physical adversarial attacks in 3D detection. Background Technology
[0002] 3D object detection is fundamental to autonomous driving perception systems, and its output directly supports downstream path planning and behavioral decisions. Modern 3D detectors are typically built upon cameras, LiDAR, or a fusion of both: cameras provide rich visual semantics, LiDAR provides precise geometric structure, and the fusion of camera and LiDAR information has become the mainstream paradigm for 3D perception in autonomous driving, achieving excellent performance. However, the adversarial robustness of 3D detectors has long been under-researched, especially in scenarios that more closely resemble real-world adversarial object attacks, where relevant security assessment methods remain lacking.
[0003] Existing research has demonstrated effective attacks against pure camera detectors, pure LiDAR detectors, and multimodal fusion detectors. However, most of these methods directly target sensor observations, i.e., perturbing image input, LiDAR input, or both simultaneously, rather than constructing a single adversarial object acting on a shared 3D geometry. As a result, even when attacking multiple modalities simultaneously, the resulting perturbations may not correspond to the same underlying physical object, thus weakening the physical plausibility and cross-modal consistency of the attack. Although existing work on object space or physical adversarial approaches has shown that a single adversarial object can simultaneously affect camera and LiDAR perception, a unified adversarial representation that maintains geometric consistency across image and point cloud observations for modern 3D detectors remains lacking. A physically meaningful cross-modal attack should be implemented as a single object whose camera appearance and LiDAR response originate from the same 3D source and remain consistent despite changes in viewpoint, placement, and scene configuration.
[0004] Beyond the limitations at the representation level, the second problem lies in the optimization process. Even with a shared adversarial object, its attack effectiveness can significantly degrade in different scenarios. On one hand, the optimization signals generated by the camera and LiDAR paths for different object attributes are often unbalanced, causing the learned object to overfit to the gradient-dominated mode. On the other hand, optimization under fixed placement makes the adversarial object highly sensitive to pose changes, resulting in decreased robustness when the object's position and orientation change. Therefore, current technologies cannot construct a unified adversarial object across modalities, nor can they optimize it to remain effective under different scenario conditions and bridge the domain gap between digital simulation and the physical world. Summary of the Invention
[0005] This invention proposes a physical adversarial attack method and apparatus for 3D detection, which addresses the lack of a unified cross-modal adversarial representation in existing adversarial attacks for 3D detectors. The perturbations applied to images and point clouds cannot be mapped to the same physical object, resulting in limited physical realizability and insufficient cross-modal consistency. At the same time, the dispersed optimization signals on heterogeneous attributes and overfitting under fixed poses cause the attack effect of adversarial objects to degrade significantly under cross-scene and cross-pose conditions.
[0006] To address the aforementioned technical problems, this invention provides a method for physical adversarial attacks in 3D detection. The method generates adversarial objects that can be inserted into a scene perceived by a camera and lidar to deceive the 3D detector; it includes the following steps: Step S1: Encode the potential of the adversarial object The input is a generator decoder, which decodes the data to obtain a shared three-dimensional Gaussian set representing the adversarial object; Step S2: Perform object-to-world pose transformation and sensor extrinsic parameter transformation on the shared 3D Gaussian element set in sequence to obtain the first 3D Gaussian element set located in the camera coordinate system and the second 3D Gaussian element set located in the lidar coordinate system. Step S3: Perform differentiable Gaussian rasterization on the first 3D Gaussian primitive set, and compare the rasterization result with the camera image of the scene. Synthesize the data to obtain adversarial camera observations; then perform visibility perception based on spherical projection on the second 3D Gaussian element set. The composite data is then combined with the background point cloud of the scene to obtain counter-laser radar observations. Step S4: Input the adversarial camera observations and the adversarial lidar observations into the 3D detector, and determine the matching confidence loss based on the prediction results of the 3D detector for the adversarial object; Step S5: Apply the matching confidence loss to the latent encoding The adversarial object is generated by iteratively updating and minimizing the matching confidence loss.
[0007] Preferably, each three-dimensional Gaussian element in the shared three-dimensional Gaussian element set includes a center position parameter. Axis alignment dimension parameters , by unit quaternion Induced rotation matrix Opacity and color vectors The covariance matrix of each of the three-dimensional Gaussian elements The expression is: .
[0008] Preferably, the generative decoder is a frozen pre-trained generative decoder; the sparse voxel structure of the generative decoder is determined by the target category reference image through a sparse structure flow model; the latent encoding The multi-channel features are stored on the occupied voxels in the sparse voxel structure; the generator decoder decodes the multi-channel features and generates a multi-channel feature for each occupied voxel. Three-dimensional Gaussian elements, among which It is a preset positive integer.
[0009] Preferably, step S2 includes: rotating the object to the world using the object-to-world rotation matrix. Object-to-world translation vector With global scale factor The center of each of the three-dimensional Gaussian elements in the shared three-dimensional Gaussian element set. According to the formula Mapped to the center in the world coordinate system ; the covariance matrix of each of the three-dimensional Gaussian elements According to the formula Mapped to the covariance matrix in the world coordinate system ; Transformation from the world to the vehicle External parameters from the car to the camera The three-dimensional Gaussian element set in the world coordinate system is mapped to the camera coordinate system to obtain the first three-dimensional Gaussian element set. Through the world to vehicle transformation External parameters from vehicle to lidar The second three-dimensional Gaussian element set is obtained by mapping the three-dimensional Gaussian element set in the world coordinate system to the lidar coordinate system; wherein... Rotation matrix from the world to the vehicle Let the translation vector be from the world to the vehicle. For the rotation matrix from the car to the camera, Let be the translation vector from the car to the camera. For the rotation matrix from vehicle to lidar, This is the translation vector from the vehicle to the lidar.
[0010] Preferably, the differentiable Gaussian rasterization of the first three-dimensional Gaussian unit set in step S3 includes: for each camera viewpoint, projecting the first three-dimensional Gaussian unit set onto the image plane corresponding to the camera viewpoint using camera intrinsic parameters to obtain an object image; and comparing the object image with the camera image of the scene. Perform based on cumulative opacity per pixel Synthesize to obtain the adversarial camera observations ,in This is an index for the camera's viewpoint.
[0011] Preferably, in step S3, visibility perception of the second three-dimensional Gaussian element set is performed based on spherical projection. The synthesis includes: Project the center of each 3D Gaussian element in the second 3D Gaussian element set onto the origin of the lidar. Using a spherical coordinate system as the reference, the elevation angle, azimuth angle, and distance of each of the three-dimensional Gaussian elements are obtained; For each beam of the lidar, candidate three-dimensional Gaussian elements are selected according to a preset angle threshold, and the α response between the beam and each candidate three-dimensional Gaussian element is calculated in the form of Mahalanobis distance. The candidate three-dimensional Gaussian elements along the beam are sorted by distance and then processed from forward to backward. Synthesizing yields soft depth; Based on the soft depth, the three-dimensional point coordinates and point intensities are recovered. The recovered three-dimensional points are merged with the background point cloud of the scene. The occluded background points are removed according to the first echo rule per beam to obtain the counter-laser observation.
[0012] Preferably, the iterative update employs a focused gradient, which modifies the matching confidence loss with respect to the latent encoding. The gradient is obtained by focusing; The focusing includes: dividing the shared 3D Gaussian primitive set into multiple attribute groups according to attributes; and focusing the gradient of each attribute group in the current iteration. Norm is denoted as ,Will The exponential moving average of the norm is denoted as ; According to the formula Calculate the weight of each attribute group ,in This is a preset concentration index. This indicates that all attribute groups are being traversed. The gradient of each attribute group is calculated according to the weights. Weighted, and mapped back to the latent encoding via a chain rule. The focusing gradient is obtained.
[0013] Preferably, the nominal placement of the adversary object includes a nominal rotation matrix. Nominal translation vector With nominal scaling factor ; In each iteration of the iterative update, the pose perturbation distribution is... Independent sampling The pose perturbation, the first Individual pose perturbation Including rotational disturbances With translational disturbance , It is a preset positive integer. ; Will Each pose perturbation is applied to the nominal placement. ,get The pose after the perturbation, the first The rotation component of the pose after the perturbation With translation components Each satisfies and ; For each of the perturbation poses, the corresponding adversarial camera observations and the corresponding adversarial lidar observations are obtained according to step S3, and the corresponding focusing gradient is obtained by focusing. right The pose-averaged focusing gradient is obtained by averaging the focusing gradients. The latent encoding is then performed based on this pose-averaged focusing gradient. Iterative updates will be performed.
[0014] Preferably, the method further includes: in the latent encoding After the update is completed, the latent encoding is processed through the grid branch of the generator decoder. Decode the data to obtain an explicit 3D mesh; The texture map of the explicit 3D mesh is generated based on the color of the shared 3D Gaussian meta-set; The explicit 3D mesh and the texture map are exported as a standard 3D asset file, which is used to deploy the adversarial object in the simulator.
[0015] The present invention also provides a device for physical adversarial attacks in 3D detection, for generating adversarial objects inserted into a scene sensed by a camera and lidar to deceive a 3D detector, the device comprising: The decoding module is configured to decode the latent code of the adversarial object. The input is a generator decoder, which decodes the data to obtain a shared three-dimensional Gaussian set representing the adversarial object; The transformation module is configured to sequentially perform object-to-world pose transformation and sensor extrinsic parameter transformation on the shared three-dimensional Gaussian element set to obtain a first three-dimensional Gaussian element set located in the camera coordinate system of the camera and a second three-dimensional Gaussian element set located in the lidar coordinate system of the lidar. The instantiation module is configured to perform differentiable Gaussian rasterization on the first three-dimensional Gaussian primitive set, and then compare the rasterization result with the camera image. Synthesize the data to obtain adversarial camera observations; then perform visibility perception based on spherical projection on the second 3D Gaussian element set. Synthesis: The synthesis result is merged with the background point cloud of the scene to obtain counter-laser radar observations; The loss module is configured to input the adversarial camera observations and the adversarial lidar observations into the 3D detector, and determine the matching confidence loss based on the prediction results of the 3D detector for the adversarial object; The optimization module is configured to optimize the latent encoding based on the matching confidence loss. The adversarial object is generated by iteratively updating the data to minimize the matching confidence loss.
[0016] The beneficial effects of the present invention include at least the following: 1) In this invention, since camera observation and lidar observation are instantiated from the same set of shared three-dimensional Gaussian elements through a unified coordinate transformation, the adversarial object naturally maintains geometric consistency in the two modes, avoiding the inconsistencies introduced by disturbing the two modes separately, and improving the physical rationality of the attack.
[0017] 2) This invention optimizes adversarial objects in the low-dimensional latent space of the pre-trained generator decoder, rather than directly optimizing all Gaussian properties in a high-dimensional space with weak constraints. The search is restricted to reasonable 3D object manifolds learned by the decoder, thereby reducing the effective search space while preserving the reasonable structure of the 3D object.
[0018] 3) This invention optimizes the attribute group that is guided to carry more attack-related signals by introducing gradient focusing on the Gaussian attribute group, thereby alleviating the gradient dispersion of the camera and lidar paths on heterogeneous attributes; due to the introduction of object pose resampling, the same adversarial object is exposed to diverse placements and viewpoints during the optimization process, thereby significantly improving the robustness of the attack under pose changes. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a physical adversarial attack method for 3D detection provided in an embodiment of the present invention. Figure 2 A schematic diagram illustrating the principle of instantiating a shared 3D Gaussian set into adversarial camera observation and adversarial lidar observation via dual-pathway, as provided in this embodiment of the invention; Figure 3 A schematic diagram illustrating a qualitative comparison of attack effects provided in an embodiment of the present invention; Figure 4This is a distribution diagram of the attack success rate of an adversary object under positional changes, provided in an embodiment of the present invention. Figure 5 This is a distribution chart of attack success rate of an adversary object under rotational changes, provided in an embodiment of the present invention. Figure 6 This is a diagram illustrating the relationship between the concentration index and the attack success rate, provided in an embodiment of the present invention. Figure 7 This is a schematic diagram showing the comparison before and after an attack in the simulator provided in an embodiment of the present invention; Figure 8 A schematic diagram showing the comparison between simulated point cloud and real scanned point cloud provided in an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating the deployment results of adversarial objects in multiple scenarios of a simulator, as provided in an embodiment of the present invention. Figure 10 A schematic diagram illustrating the visualization of multiple types of adversarial objects provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a physical countermeasure attack device for three-dimensional detection provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0021] This embodiment provides a method for physical adversarial attacks in 3D detection. The method generates adversarial objects to be inserted into a scene perceived by cameras and LiDAR to deceive the 3D detector. In autonomous driving, the 3D detection scene is represented by synchronized multi-view camera images and a LiDAR point cloud. The 3D detector takes the camera images and LiDAR point cloud as input and outputs a set of predicted 3D bounding boxes with category labels and confidence levels. This embodiment employs an adversarial object insertion attack, inserting an object of a target category into the scene and suppressing its detectability, causing the detector to assign it a low confidence level or not generate a matching prediction, while keeping the rest of the scene unchanged. This abstracts the scenario where physically introduced roadside objects, temporary obstacles, or camouflaged targets may not be sufficiently detected across multiple perception modalities.
[0022] For example, let the camera image be LiDAR point cloud is The three-dimensional detector is ;by Indicates the inserted adversary object, with Indicates will The camera and lidar observations obtained after inserting the original scene remain unchanged; This represents the confidence score of the detection prediction that matches the inserted target; if no matching prediction exists, then... The attack problem is transformed into minimizing the matching confidence: (1) To solve the above problem, this embodiment does not directly optimize sensor observations, but instead parameterizes the adversarial object as a shared set of three-dimensional Gaussian primitives and optimizes it in the low-dimensional latent space of the generative model. For example... Figure 1 As shown, the method of this embodiment includes the following steps.
[0023] Step S1: Encode the latent adversarial object The input is a generator decoder, which decodes the shared 3D Gaussian primitives representing the adversarial object.
[0024] Specifically, in autonomous driving, camera images and LiDAR point clouds are different sensing results of the same physical scene. For a cross-modal adversarial target, this means that its image appearance and LiDAR response should be bound to the same potential object, rather than being constructed independently in two modalities. Therefore, this embodiment represents the adversarial object as a set of shared 3D Gaussian primitives for instantiation across the same 3D source.
[0025] Each 3D Gaussian element in the shared 3D Gaussian element set includes a center position parameter. Axis alignment dimension parameters , by unit quaternion Induced rotation matrix Opacity and color vectors The spatial distribution of each three-dimensional Gaussian element is represented by the following equation: (2) in For any position in three-dimensional space, Indicates the first A three-dimensional Gaussian element at position The response value at the location; The covariance matrix of each three-dimensional Gaussian element Represented as: (3) Adversarial objects are represented as ,in It defines a continuous three-dimensional object in the scene.
[0026] Directly optimizing all Gaussian properties would result in a high-dimensional and weakly constrained search space. Therefore, this embodiment optimizes the adversarial object in a low-dimensional latent space defined by the pre-trained image and the 3D generative model.
[0027] In this embodiment, the generator decoder is a frozen pre-trained generator decoder; the sparse voxel structure of the generator decoder is determined by the target category reference image through a sparse structure flow model; latent encoding... The multi-channel features are stored on occupied voxels in a sparse voxel structure; the generator decoder decodes the multi-channel features and generates a generator for each occupied voxel. A three-dimensional Gaussian element, This is a preset positive integer. The decoding process can be represented as: (4) in This represents the frozen pre-trained generator-decoder. During optimization, both the generator backbone and the decoder remain fixed; the attack only updates the latent encoding. The generative model thus acts as a parameterized prior, preserving the reasonable structure of the 3D object while reducing the effective search space.
[0028] In one specific implementation, the generative decoder employs a decoding branch from a pre-trained image to a 3D generative model, with decoding performed in three stages. The first stage is occupancy prediction: an independent sparse structured flow model first predicts which voxels are occupied in the voxel grid, retaining only voxels with occupancy values greater than zero, resulting in a sparse set of occupied voxel coordinates. This sparse structure is determined once by the reference image and remains unchanged during adversarial optimization. The second stage is latent-to-feature decoding: latent encoding is stored as multi-channel features on each occupied voxel. This is first denormalized using pre-computed channel-wise statistics, then contextualized in the spatial neighborhood by multiple sparse self-attention modules, and finally projected onto the multi-dimensional output channels of each voxel by a linear layer. The third stage is attribute segmentation and activation: the multidimensional output is continuously segmented into five attribute groups based on position, color, scale, rotation, and opacity. Each group corresponds to several Gaussians on each voxel, and each group undergoes specific activation. The position offset is superimposed on the voxel center, normalized, and scaled according to the object's bounding box. The scale parameter is activated with soft addition to avoid degenerate Gaussians. The rotation quaternion is normalized. The opacity is mapped to the effective range using a function. The color is stored as a zero-order spherical harmonic coefficient. The resulting set of Gaussians collectively defines the adversarial object. Throughout the optimization process, only the voxel-by-voxel multichannel latent encoding is updated; all decoder parameters, sparse voxel structure, and attribute activations remain unchanged, thus ensuring that the adversarial perturbation always lies within the manifold learned by the generative model.
[0029] Step S2: Perform object-to-world pose transformation and sensor extrinsic parameter transformation sequentially on the shared 3D Gaussian primitive set to obtain the first 3D Gaussian primitive set located in the camera coordinate system of the camera and the second 3D Gaussian primitive set located in the lidar coordinate system of the lidar.
[0030] Specifically, the shared 3D Gaussian element set is initially defined in the standard object coordinate system, independent of any particular scene or sensor. This embodiment performs the transformation first through an object-to-world rotation matrix. Object-to-world translation vector With global scale factor The center of each 3D Gaussian element in the shared 3D Gaussian element set will be... Mapped to the center in the world coordinate system using the following formula : (5) The covariance matrix of each three-dimensional Gaussian element Mapped to the covariance matrix in the world coordinate system using the following formula : (6) Subsequently, for each synchronized sensor observation, the calibrated vehicle pose and sensor extrinsic parameters are used to further map the object in the world coordinate system to the corresponding sensor coordinate system. Specifically, this is achieved through a world-to-vehicle transformation. External parameters from the car to the camera The first 3D Gaussian element set is obtained by mapping the 3D Gaussian element set in the world coordinate system to the camera coordinate system; this is achieved through world-to-vehicle transformation. External parameters from vehicle to lidar The three-dimensional Gaussian element set in the world coordinate system is mapped to the lidar coordinate system of the lidar to obtain the second three-dimensional Gaussian element set.
[0031] Taking the mapping at the center as an example, its composite transformation can be expressed as: (7) in Let the coordinates of the center of this 3D Gaussian element be in the corresponding sensor coordinate system. Take at the camera path , Take in the lidar path ; Rotation matrix from the world to the vehicle Let the translation vector be from the world to the vehicle. For the rotation matrix from the car to the camera, Let be the translation vector from the car to the camera. For the rotation matrix from vehicle to lidar, The translation vector from the vehicle to the LiDAR is denoted as , where the vehicle is the vehicle equipped with both the camera and the LiDAR.
[0032] The above transformations are all given by the scene's sensor calibration parameters. The covariance is also transformed accordingly using composite rotation. Thus, all perceptual pathways are instantiated from a shared, world-anchored adversarial object, and their final observations remain consistent with the scene's calibration sensor geometry, thereby ensuring cross-modal geometric consistency in construction. The principle of instantiating a shared 3D Gaussian element set into two modal observations via dual pathways is as follows: Figure 2 As shown.
[0033] Step S3: Perform differentiable Gaussian rasterization on the first three-dimensional Gaussian set, and perform α synthesis with the camera image of the scene to obtain adversarial camera observation; perform visibility perception α synthesis based on spherical projection on the second three-dimensional Gaussian set, and merge the synthesis result with the background point cloud of the scene to obtain adversarial lidar observation.
[0034] For the first 3D Gaussian set of camera paths, a differentiable Gaussian rasterization is performed. Specifically, for each camera viewpoint, the first 3D Gaussian set of Gaussian elements is projected onto the image plane corresponding to that camera viewpoint using the camera's intrinsic parameters, and differentiable Gaussian rasterization is performed to obtain the object image; the object image is then compared with the camera image of the scene. Perform based on cumulative opacity per pixel Synthesis, to obtain adversarial camera observations ,in This is an index for the camera's viewpoint. The compositing process can be represented as: (8) in This represents the set of Gaussian elements after the transformation in step S2. Indicates the first Camera intrinsics from a single camera perspective. This represents the corresponding pose transformation. This indicates the background image is processed based on the cumulative opacity per pixel. Synthesis. Since the rendering is based on the same placed 3D object and is completed using calibrated camera geometry, the resulting appearance remains consistent across camera views; the synthesized image is fed into the network after passing through the detector's original preprocessing flow, and this path remains differentiable for Gaussian sets and even the latent encoding.
[0035] Unlike camera paths, lidar observations cannot be obtained by directly rendering shared objects as a continuous appearance because the detector expects sparse measurements organized according to beam structure. Therefore, this embodiment instantiates the lidar path by performing a differentiable approximation of the measurement process using sensor perception, thus preserving the acquisition cues most relevant to downstream 3D detection. Specifically, the second 3D Gaussian set is processed as follows.
[0036] First, a spherical projection is performed, projecting the center of each 3D Gaussian element in the second 3D Gaussian element set onto the origin of the lidar. Using a spherical coordinate system as the reference, the elevation, azimuth, and distance of each three-dimensional Gaussian element are obtained; its three-dimensional covariance is obtained through the Jacobian transformation from Cartesian to spherical coordinates. Projecting onto the angular domain yields the two-dimensional angular covariance used for evaluating the overlap between the beam and the Gaussian beam: (9) in For regularization terms, It is a second-order identity matrix.
[0037] Then, the pairing and response calculation of the beam and Gaussian are performed: for each beam of the lidar, candidate three-dimensional Gaussian elements are selected according to a preset angle threshold; for a set of beam-Gaussian pairs, the distance between the beam and each candidate three-dimensional Gaussian element is calculated in the form of Mahalanobis distance. response: (10) in The angular offset between the beam and the center of the Gaussian beam. Set the opacity to Gaussian. Then proceed... Synthesis: Candidate 3D Gaussian elements along the beam are sorted by distance and then processed from forward to backward. Synthesize to obtain the soft depth. Let the permutation sorted by depth be denoted as . Then, transmittance, weight, and soft depth are given by the following formulas, (11) (12) in For the first After sorting the beams, the first Transmittance before Gauss For its composite weight, The soft depth of the beam. For the sorted number The echo distance corresponding to each Gaussian, and the hit confidence of the beam are taken as... As a soft measure of effectiveness, beams with a hit confidence level below a preset threshold are discarded as missed tests.
[0038] Next, the 3D point coordinates and point intensity are recovered based on soft depth: for each hit beam, the 3D coordinates are recovered as follows: ,in For a unit beam direction, the recovered points inherit the ring index and azimuth binning of that beam, thus naturally presenting the angular sparsity and ring structure of a real lidar scan. Each generated point is assigned an intensity value consistent with the point cloud intensity range. Before synthesis, the intensity of each beam and Gaussian pair is: (13) in The intensity of the beam and the Gaussian pair, Gaussian reflectivity For its surface normal, The angle between the beam direction and the surface normal. For reference distance, The echo distance is given; reflectivity is estimated by mapping Gaussian color through perceived brightness.
[0039] Finally, scene fusion is performed: the recovered 3D points are merged with the background point cloud of the scene, and occluded background points are removed according to the first echo rule per beam. That is, for each key occupied by the injection point, which is composed of the ring index and the azimuth bin, background points that are farther away than the injection echo distance are removed. Physical modeling is used to insert the occlusion of objects on more distant surfaces, thereby obtaining counter-LiDAR observation: (14) in Indicates the configuration of the lidar sensor. This represents the background scan after removing points replaced by closer simulated echoes. In this embodiment, the lidar path is a detector-oriented differentiable approximation, and sensor-level phenomena such as material reflection, multipath effects, and atmospheric scattering are not modeled.
[0040] After obtaining the adversarial observations from the two aforementioned pathways, the entire instantiation path from the adversarial object's latent encoding to adversarial camera observations and adversarial lidar observations is designed to be differentiable, thus enabling the gradient of the detection loss in subsequent steps to be backpropagated to the latent encoding. Specifically, starting from the latent encoding, through Gaussian decoding, world coordinate system placement, spherical projection, ... Throughout the synthesis and 3D coordinate recovery process: soft depth is used as a continuously weighted average instead of a hard first echo, thus providing smooth gradients for Gaussian position, scale, and opacity; binary hit and miss decisions are separated from the computation graph, so that the gradient flow of surviving point coordinates and intensity remains unaffected; on the detector side, the hard voxelization used by standard LiDAR detectors is replaced by soft voxelization, allowing the gradient of detection loss to propagate back through the generated point coordinates to the Gaussian properties and eventually back to the latent encoding.
[0041] Step S4: Input the adversarial camera observations and adversarial lidar observations into the 3D detector, and determine the matching confidence loss based on the prediction results of the 3D detector for the adversarial object.
[0042] Specifically, adversarial camera observations and adversarial lidar observations are input into a 3D detector to obtain predicted 3D bounding boxes with category labels and confidence levels. The inserted adversarial object is then matched with predicted boxes of the same category whose center distance does not exceed a preset matching radius, and the box with the highest confidence level is selected. As a matching confidence loss Zero is used when there is no matching prediction. Minimize That is, suppressing the detectability of the inserted adversarial object, consistent with Equation (1); where the thresholds such as the matching radius are adaptively determined according to the actual bounding box size.
[0043] Step S5: Encode the latent based on the matching confidence loss Iterative updates are performed to minimize the matching confidence loss, thereby generating adversarial objects.
[0044] In this embodiment, the shared 3D Gaussian representation ensures consistent instantiation at the cross-modal level. However, the optimization process is still affected by cross-scene domain bias, which mainly comes from two aspects: firstly, feature-level bias, i.e., the camera and LiDAR paths provide unbalanced gradient intensities on different attribute groups; secondly, instance-level bias, i.e., the attack effect varies significantly under different object poses. To mitigate these two types of bias, this embodiment employs two complementary strategies: gradient focusing and pose resampling. The camera path and LiDAR path respectively provide matching confidence losses. Two gradients regarding latent encoding and Their sum defines the basic update direction: (15) To address feature-level bias, this embodiment introduces gradient focusing on the decoded Gaussian properties. Let all Gaussian properties obtained from latent encoding and decoding be denoted as... , of which Each attribute group is denoted as .
[0045] Iterative updates employ focused gradients, obtained by focusing the gradient of the matching confidence loss with respect to the latent encoding. Focusing involves: dividing the Gaussian attributes into multiple attribute groups; and for each attribute group, calculating its gradient in the current iteration. Norm is denoted as And the exponential moving average of this norm over iterations is denoted as Then, the weights of each attribute group are calculated using the following formula. : (16) in For the first The smoothness norm of the group, This is a preset concentration index. This indicates iterating through all attribute groups. When... At this point, the update degenerates into a uniform weighting of each group; The larger the value, the more it tends to emphasize attribute groups that have stronger and more stable gradient responses.
[0046] The gradients of each attribute group are then weighted according to their respective weights to obtain the weighted attribute gradients. : (17) Then, by mapping back to the latent encoding using the chain rule, we obtain the focused gradient: (18) This mechanism smooths out noise in the gradient statistics at each step and alleviates attribute-level gradient dispersion, allowing potential updates to be driven more by attack-related attribute groups without imposing any hard modality-to-attribute assignments.
[0047] To address instance-level bias, this embodiment resamples the object pose around the nominal placement during optimization. The adversarial object nominal placement includes a nominal rotation matrix. Nominal translation vector With nominal scaling factor In each iteration of the iterative update, the pose perturbation distribution is changed from the preset distribution. Independent sampling The pose perturbation, the first Individual pose perturbation Including rotational disturbances With translational disturbance , It is a preset positive integer. Applying a pose perturbation to the nominal placement yields the perturbed pose, whose rotational and translational components satisfy the following: (19) For each perturbation-triggered pose, the corresponding adversarial camera observation and the corresponding adversarial lidar observation are obtained according to step S3, and the corresponding focusing gradient is obtained according to the focusing. The average focusing gradient of each perturbation-triggered pose is obtained. In this way, the same adversarial object is exposed to diverse placements and viewpoints during the optimization process, thereby improving its robustness beyond a single pose configuration.
[0048] The pose-averaged focusing gradient can be expressed as: (20) in For the first The pose-averaged focusing gradient of the next iteration. The iteration number, For the first The potential encoding of the next iteration. and These are the camera images and LiDAR point clouds of the sampled keyframes, respectively. This represents the complete instantiation pipeline consisting of steps S1 to S3. This represents the confidence level in matching the insertion adversarial target. Finally, the latent encoding is updated using projective gradient descent with momentum: (twenty one) (twenty two) in For the first The momentum accumulation in the next iteration. Step size, As the momentum factor, the projection operator projects the updated latent code onto the clean initialization. Bounded feasible set superior, This is a preset perturbation budget. Combining the scene distribution and pose perturbation distribution, the overall optimization objective is to minimize the expected matching detection confidence across both distributions, i.e.: (twenty three) in Indicates the expectation. The scene distribution used for training, The aforementioned pose perturbation distribution, This indicates sampling keyframes from the scene distribution. This indicates that pose perturbations are sampled from the pose perturbation distribution.
[0049] In one implementation, the generative model generates 3D objects via a conditional flow model, mapping noise to a structured latent representation, which is then decoded into a Gaussian by the decoder. Adversarial optimization is performed in the noise space: clean initialization is done with random noise samples conditioned on a target class reference image; during each forward propagation, the current adversarial noise is propagated through the flow model solver to obtain a denoised structured latent representation. This flow step is performed without calculating gradients, and its output is separated into new leaf nodes, subsequently mapped into a Gaussian by the decoder. The gradient of the detection loss is backpropagated only through the decoder path. Throughout the optimization process, the flow model, decoder, and all detector weights remain frozen; the only variable updated is the latent encoding in the noise space. At the end of each training epoch, the validation set attack success rate is monitored, and the checkpoint with the highest attack success rate is selected as the final adversarial object to avoid overfitting to recent mini-batches.
[0050] After the latent encoding update is completed and the adversarial object is obtained, in one alternative embodiment, the method further includes exporting a standard 3D asset that can be deployed in the simulator. Specifically, the latent encoding is decoded by generating a mesh branch of the decoder to obtain an explicit 3D mesh; a texture map of the explicit 3D mesh is generated based on the colors of a shared 3D Gaussian primitive set; the explicit 3D mesh and texture map are exported as a standard 3D asset file, which is used to deploy the adversarial object in the simulator. In one specific implementation, the mesh export is performed in four stages: the first stage is mesh extraction, where the mesh decoder predicts the directed distance values, vertex deformations, and colors of each voxel on the voxel mesh, and then obtains the mesh using an isosurface extraction algorithm; the second stage is post-processing, where the mesh is reduced in surface area, filled with holes, and parameterized; the third stage is texture baking, where Gaussian representations are rendered from multiple perspectives and the multi-view observations are projected onto the parameterized mesh to obtain a texture map; the fourth stage is export, where coordinate system transformation is performed, materials are assigned, and the mesh is exported in a standard binary format. Since the mesh branch and the Gaussian branch share the same underlying encoding, the geometry and rough appearance are well preserved, and the resulting assets are as close as possible to the optimized Gaussian object.
[0051] To illustrate the effectiveness of this invention, test data from a specific embodiment is provided below. This embodiment is implemented on the nuScenes dataset, which includes synchronized six-camera images, 32-line LiDAR scans, and 3D object annotations. Under general settings, keyframes were sampled from 28,130 keyframes across 700 training scenes to optimize a single shared adversarial object. After fixing the optimized object, it was tested on 6,019 validation keyframes with different placements and poses. No further optimization was performed for any specific scene during the testing process.
[0052] The tests covered six 3D detectors across three perception paradigms, including two pure camera detectors, BEVDet and DETR3D; two pure LiDAR detectors, CenterPoint and VoxelNeXt; and two multimodal fusion detectors, BEVFusion and TransFusion. For detectors that rely on voxelized point cloud input, soft voxelization was only employed in the LiDAR agent on the attack side to enable gradient backpropagation; the detector architecture and its inference process remained unchanged during the tests.
[0053] The test uses two metrics: attack success rate and confidence decrease. For each frame, the inserted target is matched with the highest confidence prediction of the same category that meets the fixed center distance and IoU threshold. If there is no matching prediction, it is considered undetected and assigned a score of zero. The attack success rate is the percentage of frames in which the confidence of the matched target is lower than a given threshold, given at thresholds of 0.3 and 0.5, denoted as ASR@0.3 and ASR@0.5, respectively. The confidence decrease is the average decrease in the confidence of the matched target from clean input to adversarial input under the same matching protocol.
[0054] Regarding the attack effectiveness, this embodiment achieved a strong attack effect on all six detectors mentioned above. A comparison before and after the attack is shown below. Figure 3 As shown, Figure 3 MiuGS, as indicated in the text, refers to the method of this invention. Adv3D, BEV-Robust, and CL-FusionAttack are existing comparative methods. Adv3D is an insert camera modality adversarial method, while BEV-Robust and CL-FusionAttack are adversarial methods targeting the camera detector and the camera-LiDAR fusion detector, respectively. On the fusion detectors BEVFusion and TransFusion, the confidence of the matched target decreased from approximately 0.41 and 0.45 in the clean state to approximately 0.14 and 0.04 in the adversarial state, respectively, with ASR@0.3 reaching approximately 85.3% and 95.3%, respectively. On the pure camera detector BEVDet, the target confidence decreased from approximately 0.64 to approximately 0.07, with ASR@0.3 reaching approximately 87.5%, significantly better than the existing interpolation-based control method Adv3D (ASR@0.3 approximately 22.7%) under the same settings. On the pure lidar detectors CenterPoint and VoxelNeXt, ASR@0.5 reached approximately 64.2% and 76.5%, respectively. In the above tests, the average accuracy and overall performance of other objects in the scene remained basically unchanged, with minimal collateral impact. Combining the six detectors, the average attack success rate of this embodiment is approximately 78.8%, with the strongest performance on the fusion detector.
[0055] Regarding pose robustness, the optimized adversarial object was tested on a 5×5 translation mesh with 12 rotation angles. The ASR@0.3 distribution under positional changes is as follows: Figure 4 As shown, the ASR@0.3 distribution under rotation is as follows: Figure 5 As shown. By Figure 4 As can be seen, the ASR@0.3 score is close to perfect under various placements at close and medium ranges, only gradually decreasing at long range or with large lateral offsets. The average ASR@0.3 on the translation grid is approximately 88.3%. Figure 5 As can be seen, the ASR@0.3 of the attack is between approximately 91% and 98% at various rotation angles, showing relatively stable overall performance. This data indicates that pose resampling reduces overfitting to a single placement, ensuring the attack remains effective under pose variations.
[0056] Regarding gradient focusing, the concentration index Tests were conducted using values of 0, 0.5, 1, 1.5, and 2. The ASR@0.3 value on the three types of detectors was randomly determined. Changes such as Figure 6 As shown. By Figure 6 It is evident that this effect is related to the detector type: fusion detectors exhibit higher and more stable ASR@0.3 across all values; pure lidar detectors show... Significant improvements were achieved at the [specific location] level, with ASR@0.3 increasing from approximately 9.9% to approximately 51.3%; while pure camera detectors showed improvements at the uniformly weighted [specific location]. The highest ASR (at 0.3) was approximately 55%. Based on this, the concentration index... According to the calibration of the detector: when the gradient intensity of the lidar path and the camera path is unbalanced on the attribute group, take a larger value; when the contributions of each attribute group are relatively balanced, take zero or a smaller value.
[0057] Regarding the role of core modules, the effects of gradient focusing and pose resampling were tested on the fusion detector BEVFusion, the pure camera detector BEVDet, and the pure lidar detector VoxelNeXt. The results show that introducing pose resampling alone improves the ASR@0.3 of all three detectors from approximately 78.6%, 55.0%, and 6.2% to approximately 84.7%, 87.5%, and 9.9%, respectively, resulting in improvements across all three detectors. The effect of gradient focusing is detector-dependent; on VoxelNeXt, when both gradient focusing and pose resampling are enabled simultaneously, the ASR@0.3 is further improved to approximately 51.3%. The impact is smaller on BEVFusion and decreases on BEVDet. Therefore, the complete solution performs best on BEVFusion and VoxelNeXt, while the best performance on BEVDet is achieved without gradient focusing, consistent with the aforementioned concentration index calibration method.
[0058] For cross-detector transfer, a single-source protocol was adopted, where adversarial objects were optimized only on each source detector, and then tested on all six detectors, with ASR@0.3 used to measure the transfer performance. The results show that transfer within the same perceptual paradigm is strong, especially between two fused detectors, with ASR@0.3 for mutual transfer between BEVFusion and TransFusion being approximately 75.9% and 72.8%, respectively. Cross-paradigm transfer is more limited; adversarial objects trained on the camera have the highest ASR@0.3 on the LiDAR detector (approximately 2.1%), while those trained on the LiDAR detector have the highest (approximately 2.3%). Among the source detectors, adversarial objects trained through fusion have the widest transfer range. This indicates that because adversarial objects are represented by a unified 3D representation consistent across modalities, the same object can be directly evaluated on different detectors, and objects optimized by both perceptual flows have a relatively wider transfer range.
[0059] Regarding its potential in the physical world, further testing was conducted in an independent simulator. A sensor configuration consistent with the dataset style was employed, along with a fusion detector fine-tuned on simulator-style data. The optimized adversarial object was converted into a mesh asset via the aforementioned mesh export process and then imported into the simulator. A comparison of before and after the attack was performed in the simulator. Figure 7 As shown in the figure. Test data shows that the attack effect in the simulator is weakened compared to the dataset, with ASR@0.3 at approximately 36.5%, ASR@0.5 at approximately 44.3%, and a confidence decrease of approximately 0.32. This weakening is consistent with the information loss from Gaussian to mesh transformation and the domain differences in scene appearance, material rendering, and sensor modeling in the simulator. Even so, the transformed adversarial object still causes a significant decrease in detection confidence, indicating that the learned adversarial object retains some effectiveness after mesh transformation and domain migration. In addition, the fidelity test of the simulated LiDAR path is as follows: Figure 8 As shown, the simulated point cloud can be reliably detected by the downstream model. The detector's confidence levels on simulated cars and simulated traffic cones are approximately 0.79 and 0.61, respectively, which are close to the actual scans of approximately 0.82 and 0.65, indicating that the attack effect comes from the optimization itself rather than from simulation-induced bias. The deployment results of the adversarial object in various urban and highway scenarios, under various lighting and weather conditions in the simulator are shown below. Figure 9 As shown, the detection confidence can be suppressed in various scenarios. The method is not limited to a single object category; visualization results for multiple categories such as cars, bicycles, cardboard boxes, trash cans, guardrails, obstacles, and warning signs are also available. Figure 10 As shown, clean objects can be reliably detected, and the detection confidence is suppressed after the adversarial optimization.
[0060] The calibration of the matching threshold and optimization hyperparameters in the above test is explained below. In this embodiment, all matching thresholds are determined by the bird's-eye view diagonal of the ground truth bounding box. The calculations ensure that the evaluation criteria scale naturally with the object size. Specifically, the matching radius, evaluability threshold, and success threshold are obtained by multiplying the diagonal by their respective coefficients and truncating to their respective upper and lower bounds. In one implementation, the matching radius coefficient is 0.6 with a lower bound of 1.5m and an upper bound of 4.0m; the evaluability threshold coefficient is 0.06 with a range of 0.10 to 0.35; and the success threshold coefficient is 0.02 with a range of 0.03 to 0.12. To achieve fair comparisons among detectors with different sensitivities, a target location is only included in the evaluation if the clean detection confidence on all compared detectors simultaneously exceeds the evaluability threshold. This prevents a detector from failing to detect a clean target from being counted as having a high attack success rate on that detector.
[0061] This embodiment also provides a device for physical adversarial attacks in 3D detection. For example... Figure 11 The device is used to generate adversarial objects that are inserted into the scene perceived by the camera and lidar to deceive the 3D detector. The device includes a decoding module, a transformation module, an instantiation module, a loss module, and an optimization module.
[0062] The decoding module is configured to encode the potential of adversarial objects. The input is a generator decoder, which decodes the shared 3D Gaussian primitives representing the adversarial object.
[0063] The transformation module is configured to sequentially perform object-to-world pose transformation and sensor extrinsic parameter transformation on the shared 3D Gaussian element set to obtain a first 3D Gaussian element set located in the camera coordinate system of the camera and a second 3D Gaussian element set located in the lidar coordinate system of the lidar.
[0064] The instantiation module is configured to perform differentiable Gaussian rasterization on the first 3D Gaussian primitive set, and then compare the rasterization result with the camera image of the scene. Synthesizing to obtain adversarial camera observations; and performing visibility perception on the second 3D Gaussian ensemble based on spherical projection. The synthesis process merges the synthesized result with the background point cloud of the scene to obtain counter-laser radar observations.
[0065] The loss module is configured to input adversarial camera observations and adversarial lidar observations into a 3D detector, and determine the matching confidence loss based on the 3D detector's prediction results of the adversarial object.
[0066] The optimization module is configured to perform latent encoding based on matching confidence loss. Iterative updates are performed to minimize the matching confidence loss, thereby generating adversarial objects. The specific implementation methods of each module of the device are the same as the corresponding steps in the aforementioned method embodiments, and will not be repeated here.
[0067] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; only preferred embodiments of the present invention are illustrated. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. As long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0068] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the appended claims.
Claims
1. A method for physical adversarial attacks in 3D detection, the method being used to generate adversarial objects that are inserted into a scene perceived by a camera and a lidar to deceive a 3D detector; characterized in that, Includes the following steps: Step S1: Encode the potential of the adversarial object The input is a generator decoder, which decodes the data to obtain a shared three-dimensional Gaussian set representing the adversarial object; Step S2: Perform object-to-world pose transformation and sensor extrinsic parameter transformation on the shared 3D Gaussian element set in sequence to obtain the first 3D Gaussian element set located in the camera coordinate system and the second 3D Gaussian element set located in the lidar coordinate system. Step S3: Perform differentiable Gaussian rasterization on the first three-dimensional Gaussian unit set, and compare the rasterization result with the camera image of the scene. Synthesize the data to obtain adversarial camera observations; then perform visibility perception based on spherical projection on the second 3D Gaussian element set. The composite data is then combined with the background point cloud of the scene to obtain counter-laser radar observations. Step S4: Input the adversarial camera observations and the adversarial lidar observations into the 3D detector, and determine the matching confidence loss based on the prediction results of the 3D detector for the adversarial object; Step S5: Apply the matching confidence loss to the latent encoding The adversarial object is generated by iteratively updating and minimizing the matching confidence loss. The generator decoder is a frozen pre-trained generator decoder; the sparse voxel structure of the generator decoder is determined by the target category reference image through a sparse structure flow model; the latent encoding The multi-channel features are stored on the occupied voxels in the sparse voxel structure; the generator decoder decodes the multi-channel features and generates a multi-channel feature for each occupied voxel. Three-dimensional Gaussian elements, among which It is a preset positive integer; In step S3, visibility perception based on spherical projection is performed on the second three-dimensional Gaussian element set. The synthesis includes: Project the center of each 3D Gaussian element in the second 3D Gaussian element set onto the origin of the lidar. Using a spherical coordinate system as the reference, the elevation angle, azimuth angle, and distance of each of the three-dimensional Gaussian elements are obtained; For each beam of the lidar, candidate three-dimensional Gaussian elements are selected according to a preset angle threshold, and the α response between the beam and each candidate three-dimensional Gaussian element is calculated in the form of Mahalanobis distance. The candidate three-dimensional Gaussian elements along the beam are sorted by distance and then processed from forward to backward. Synthesizing yields soft depth; Based on the soft depth, the three-dimensional point coordinates and point intensities are recovered. The recovered three-dimensional points are merged with the background point cloud of the scene. The occluded background points are removed according to the first echo rule per beam to obtain the counter-laser observation. The iterative update employs a focused gradient, which is achieved by applying the matching confidence loss to the latent encoding. The gradient is obtained by focusing; The focusing includes: dividing the shared 3D Gaussian primitive set into multiple attribute groups according to attributes; and focusing the gradient of each attribute group in the current iteration. Norm is denoted as ,Will The exponential moving average of the norm is denoted as ; According to the formula Calculate the weight of each attribute group ,in This is a preset concentration index. This indicates that all attribute groups are being traversed. The gradient of each attribute group is calculated according to the weights. Weighted, and mapped back to the latent encoding via a chain rule. The focusing gradient is obtained.
2. The method according to claim 1, characterized in that, Each three-dimensional Gaussian element in the shared set of three-dimensional Gaussian elements includes a center position parameter. Axis alignment dimension parameters , by unit quaternion Induced rotation matrix Opacity and color vectors The covariance matrix of each of the three-dimensional Gaussian elements The expression is: .
3. The method according to claim 2, characterized in that, Step S2 includes: rotating the object to the world using the object rotation matrix. Object-to-world translation vector With global scale factor The center of each of the three-dimensional Gaussian elements in the shared three-dimensional Gaussian element set. According to the formula Mapped to the center in the world coordinate system ; the covariance matrix of each of the three-dimensional Gaussian elements According to the formula Mapped to the covariance matrix in the world coordinate system ; Transformation from the world to the vehicle External parameters from the car to the camera The three-dimensional Gaussian element set in the world coordinate system is mapped to the camera coordinate system to obtain the first three-dimensional Gaussian element set. Through the world to vehicle transformation External parameters from vehicle to lidar The second three-dimensional Gaussian element set is obtained by mapping the three-dimensional Gaussian element set in the world coordinate system to the lidar coordinate system; wherein... Rotation matrix from the world to the vehicle Let the translation vector be from the world to the vehicle. For the rotation matrix from the car to the camera, Let be the translation vector from the car to the camera. For the rotation matrix from vehicle to lidar, This is the translation vector from the vehicle to the lidar.
4. The method according to claim 3, characterized in that, Step S3, the differentiable Gaussian rasterization of the first 3D Gaussian unit set, includes: for each camera viewpoint, projecting the first 3D Gaussian unit set onto the image plane corresponding to the camera viewpoint using camera intrinsic parameters to obtain an object image; and comparing the object image with the camera image of the scene. Perform based on cumulative opacity per pixel Synthesize to obtain the adversarial camera observations ,in This is an index for the camera's viewpoint.
5. The method according to claim 1, characterized in that, The nominal placement of the adversary object includes a nominal rotation matrix. Nominal translation vector With nominal scaling factor ; In each iteration of the iterative update, the pose perturbation distribution is... Independent sampling The pose perturbation, the first Individual pose perturbation Including rotational disturbances With translational disturbance , It is a preset positive integer. ; Will Each pose perturbation is applied to the nominal placement. ,get The pose after the perturbation, the first The rotation component of the pose after the perturbation With translation components Each satisfies and ; For each of the perturbation poses, the corresponding adversarial camera observations and the corresponding adversarial lidar observations are obtained according to step S3, and the corresponding focusing gradient is obtained by focusing. right The pose-averaged focusing gradient is obtained by averaging the focusing gradients. The latent encoding is then performed based on this pose-averaged focusing gradient. Iterative updates will be performed.
6. The method according to claim 1, characterized in that, The method further includes: in the latent encoding After the update is completed, the latent encoding is processed through the grid branch of the generator decoder. Decode the data to obtain an explicit 3D mesh; The texture map of the explicit 3D mesh is generated based on the color of the shared 3D Gaussian meta-set; The explicit 3D mesh and the texture map are exported as a standard 3D asset file, which is used to deploy the adversarial object in the simulator.
7. A physical adversarial attack device for 3D detection, used to generate adversarial objects inserted into a scene sensed by a camera and lidar to deceive the 3D detector, characterized in that, The device includes: The decoding module is configured to decode the latent code of the adversarial object. The input is a generator decoder, which decodes the data to obtain a shared three-dimensional Gaussian set representing the adversarial object; The transformation module is configured to sequentially perform object-to-world pose transformation and sensor extrinsic parameter transformation on the shared three-dimensional Gaussian element set to obtain a first three-dimensional Gaussian element set located in the camera coordinate system of the camera and a second three-dimensional Gaussian element set located in the lidar coordinate system of the lidar. The instantiation module is configured to perform differentiable Gaussian rasterization on the first three-dimensional Gaussian primitive set, and then compare the rasterization result with the camera image. Synthesize the data to obtain adversarial camera observations; then perform visibility perception based on spherical projection on the second 3D Gaussian element set. Synthesis: The synthesis result is merged with the background point cloud of the scene to obtain counter-laser radar observations; The loss module is configured to input the adversarial camera observations and the adversarial lidar observations into the 3D detector, and determine the matching confidence loss based on the prediction results of the 3D detector for the adversarial object; The optimization module is configured to optimize the latent encoding based on the matching confidence loss. The adversarial object is generated by iteratively updating the data to minimize the matching confidence loss.
Citation Information
Patent Citations
Inrush fluid type Gaussian rendering structure method for dynamic scene modeling
CN121095411A
Privacy preserving visual localization
US20260105688A1