Multi-view 3D countermeasure attack method based on Gaussian splashing
Through the multi-view 3D adversarial attack method based on Gaussian splash, the problem of difficulty in spoofing deep classification networks in the prior art is solved, and efficient 3D adversarial sample generation is achieved at multi-view angles, which improves the success rate and computing efficiency of the attack.
Patent Information
- Application Number
- CN202510160681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing three-dimensional adversarial attack methods are difficult to effectively deceive the deep classification network from multiple perspectives, and the calculation cost is high, making it difficult to directly control the location and scale of the disturbance.
Using a multi-view 3D adversarial attack method based on Gaussian splatter, the original 3D scene is constructed, the Gaussian contribution is calculated, and a tiny 3D perturbation is generated, and the rendering loss, adversarial loss, position loss and color loss is optimized to ensure that the perturbation maintains the adversarial effect and visual nature in multiple viewing angles.
The generated 3D adversarial samples can effectively deceive the classifier from multiple perspectives, improve the success rate and generalization ability of the attack, and reduce the computational cost, achieving efficient three-dimensional scene reconstruction and multi-view continuous rendering of adversarial samples.
Smart Images

Figure CN119990252A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and artificial intelligence security, and specifically relates to a multi-view 3D adversarial attack method based on Gaussian splashing. The method generates and optimizes 3D adversarial perturbations in 3D scenes to deceive deep classification networks under multiple viewpoints. The method is suitable for target detection and other deep learning-based visual tasks, and can be used for security assessment, privacy protection, and adversarial testing of artificial intelligence systems. Background Art
[0002] In the field of adversarial attacks, a variety of adversarial attack methods have been proposed in recent years, which are mainly divided into two categories: two-dimensional adversarial attacks and three-dimensional adversarial attacks. The following are several major adversarial attack methods and their characteristics:
[0003] 1. Two-dimensional adversarial attack method
[0004] Early adversarial attacks were mainly aimed at two-dimensional images, and typical methods included the gradient-based fast gradient sign method (FGSM), projected gradient descent (PGD), and perturbation optimization methods for the physical world (such as AdvPatch). However, these methods usually only generate adversarial samples for specific perspectives and are difficult to adapt to multi-perspective changes in real scenes, resulting in poor attack generalization.
[0005] 2. Three-dimensional adversarial attack method
[0006] In order to overcome the limitations of two-dimensional adversarial attacks in the physical world, researchers have begun to explore adversarial attack methods based on three-dimensional data. Existing three-dimensional adversarial attacks mainly include methods based on point cloud, mesh, and neural radiation field (NeRF).
[0007] (1) Point cloud-based adversarial attacks
[0008] Adversarial attacks on point clouds mainly include three types of methods: point cloud perturbation, point cloud deletion, and point cloud addition. For example, Somepalli et al. proposed to perform gradient optimization on the position of points on point cloud classifiers such as PointNet++, causing the classifier to misidentify the target object. This causes the input point cloud to shift in the high-dimensional feature space, leading to model misclassification. Xi et al. studied minimizing the model's dependence on specific points, causing the model to misjudge after deleting key points. Although point cloud attack methods can effectively deceive point cloud-based DNNs, since point clouds are discrete data, it is difficult to naturally expand to the entire three-dimensional surface, resulting in unstable attack effects under different viewing angles. In addition, point cloud perturbations may cause irregular surfaces of objects during visualization, affecting applications in the real world.
[0009] (2) Mesh-based adversarial attacks
[0010] Mesh-based adversarial attacks usually optimize the geometry or texture of the mesh so that the rendered image can mislead the classifier. By adjusting the vertex positions of the mesh, the three-dimensional shape of the object is slightly deformed. For example, researchers use gradients to optimize the positions of mesh vertices so that the rendered image can deceive the target classifier from multiple perspectives. In addition, the AdvTexture method optimizes the color of the object surface so that it can achieve misleading classification from multiple perspectives. This method is usually combined with perceptual loss to ensure that the modified texture is visually consistent with the original object. Although the mesh attack method can act directly on the three-dimensional model, the following problems still exist: Poor optimization convergence: The geometric deformation of the mesh involves a highly non-convex optimization problem, which is easy to fall into the local optimum, making the attack less stable. Obvious visual forgery traces: Large movements of mesh vertices may lead to unreasonable geometric deformations and reduce concealment.
[0011] (3) NeRF-based adversarial attacks
[0012] In recent years, NeRF technology has been used in adversarial attack scenarios to exploit its continuous rendering characteristics to generate multi-view consistent adversarial samples. For example, NeRFail attempts to optimize its volume density and color field during NeRF training to generate adversarial rendered images under multiple viewpoints. However, since NeRF uses implicit scene modeling, its optimization process is computationally expensive and it is difficult to directly control the location and scale of the perturbations. In addition, NeRF requires a lot of computing resources for forward rendering, making it difficult to deploy in practical applications.
[0013] In summary, current adversarial attack schemes can achieve attacks in the 3D world, but due to the limitations of NeRF's implicit modeling characteristics, it is still impossible to generate explicit 3D adversarial samples. Recently, 3D Gaussian Splatting (3DGS) technology has brought significant progress, which greatly improves the speed of 3D reconstruction and rendering, making explicit construction of 3D objects possible. How to use 3DGS to generate explicit 3D adversarial samples to deceive DNNs in a wide range of viewing angles is a key technical problem that needs to be solved urgently. Summary of the invention
[0014] In view of the above problems, the present invention proposes a multi-view 3D counterattack method based on Gaussian splashing, and the method comprises the following steps:
[0015] 1. A multi-view 3D counterattack method based on Gaussian splashing, characterized by comprising the following steps:
[0016] S1, using multi-view images, the original 3D scene G(Θ) is constructed using the Gaussian splashing technique (where Θ represents a set of Gaussian parameters, including mean μ, covariance Σ, transparency α, and color c);
[0017] S2, contribution calculation, calculate the gradient contribution of each Gaussian in the original 3D scene to the target classifier output, and select the Gaussian with higher contribution;
[0018] S3, construct 3D perturbation, split the selected high-contribution Gaussian, and generate small 3D perturbations, where the parameters of the split 3D perturbation are different from the original Gaussian only in scale;
[0019] S4, optimizes 3D perturbations, using rendering loss, adversarial loss, position loss, and color loss to optimize 3D perturbations, ensuring that the perturbations can maintain both adversarial effects and visual naturalness under multiple viewing angles;
[0020] S5, obtain 3D adversarial samples and verify their ability to deceive the target classifier under multiple views.
[0021] 2. The multi-view 3D counterattack method based on Gaussian splashing according to claim 1, wherein the step S2 of calculating the contribution specifically comprises the following steps:
[0022] S21, for each image I at each viewing angle v, the adversarial loss uses the following cross entropy loss as the objective function, defined as:
[0023] L adv (I) = CE(f(I),y tar ) (16)
[0024] Where CE represents the cross entropy loss function, f(I) represents the predicted output of the target classifier for image I, and y tar is the preset target category;
[0025] S22, render the corresponding image I(v,G(Θ)) from multiple perspectives v∈V, sum the adversarial losses of all perspectives, and get the total loss:
[0026]
[0027] S23, for the total loss J(Θ), calculate the gradient of each Gaussian parameter Θ = {μ, Σ, α, c} respectively, and get the gradient It can be expressed as:
[0028]
[0029] S24, in order to reflect the contribution of each Gaussian in the overall adversarial attack, the normalization method is used to calculate the gradient of each parameter The contribution of is normalized using the following function:
[0030]
[0031] Where p∈{μ,Σ,α,c}, |*| represents the absolute value function, min(*) and max(*) represent the minimum and maximum values calculated, respectively;
[0032] S25, add the normalized gradients of all parameters in each Gaussian to get the contribution S of each Gaussian:
[0033]
[0034] 3. The multi-view 3D counterattack method based on Gaussian splashing according to claim 1, wherein the step S3 of constructing the 3D perturbation specifically comprises the following steps:
[0035] S31, select the k Gaussians with the highest contribution S as the splitting objects, so as to construct the 3D perturbation δ = {δ μ ,δ Σ ,δ α ,δ c};
[0036] S32, in the splitting operation, the mean, color and opacity are kept consistent with the original Gaussian. This part can be expressed as:
[0037]
[0038] S33, the variance of the disturbance is set to the minimum value among all Gaussians in order to obtain a tiny Gaussian, which can be expressed as:
[0039]
[0040] Among them, min(*) represents the operation of calculating the minimum value.
[0041] 4. The multi-view 3D counterattack method based on Gaussian splashing according to claim 1, wherein the step S4 of optimizing the 3D perturbation specifically comprises the following steps:
[0042] S41, after splitting out the perturbation, the 3D object parameters are expressed as
[0043] S42, for each view v, the rendered image I(v,G(Θ′)) is classified by the target classifier f, and the cross entropy loss function is used to measure the difference between the classification result of the rendered image and the preset target category y_tar. The adversarial loss of all view angles v∈V is accumulated, and the total adversarial loss is defined as:
[0044]
[0045] S43, in order to prevent the perturbation from deviating from the target surface, the position loss is introduced to constrain the position of the Gaussian. Let the Gaussian mean of the 3D perturbation be δ μ , the corresponding Gaussian mean in the original object is μ, then the position loss can be measured using the chamfer distance function D c To measure the difference between the two, it is defined as:
[0046]
[0047] S44, define color loss to ensure that the perturbation is highly consistent with the original object in color. Map the 3D perturbation from the world coordinate system to the camera coordinate system. The transformation process is as follows:
[0048] u=W·δ μ (25)
[0049] Where W is the observation matrix. Normalize u to get the position in the camera coordinate system. The ray r passing through the three-dimensional perturbation can be obtained according to the following formula:
[0050]
[0051] r=o+γ·d (27)
[0052] Where o is the camera origin and γ is the distance parameter. The color corresponding to the 3D perturbation is rendered as C(r), and the color loss of the 3D perturbation can be expressed as:
[0053]
[0054] Among them, R δ represents the set of rays of all three-dimensional perturbations in all viewpoints, and gt represents the true color of the corresponding ray;
[0055] S45, after weighting the above three losses with hyperparameters, the total optimization loss function is formed:
[0056]
[0057] Among them, λ adv , pos , rend and λ col are the hyperparameters for adjusting adversarial, position, rendering, and color losses, respectively;
[0058] S46, through L total Perform back propagation, calculate the gradient of the perturbation parameter δ, and use the gradient descent method to update the perturbation parameter δ, that is:
[0059]
[0060] Among them, δ Represents the learning rate.
[0061] Beneficial effects: Compared with the prior art, the present invention provides a multi-view 3D adversarial attack method based on Gaussian splashing, which can mislead the deep classification network at a wider viewing angle and produce the following beneficial effects:
[0062] 1. Explicit 3D adversarial samples: Gaussian splatting technology is used to generate explicit 3D adversarial samples, which overcomes the limitations of NeRF implicit modeling and enables adversarial samples to effectively deceive deep neural networks over a wide range of viewing angles.
[0063] 2. High visual quality: By optimizing the position and color of the perturbations, the generated adversarial samples are visually indistinguishable from the original 3D objects, ensuring high visual quality and concealment. The variance of the perturbations is set to the minimum value among all Gaussians, which significantly improves the visual imperceptibility of the perturbations.
[0064] 3. Multi-view attack capability: The generated 3D adversarial samples can effectively deceive the classifier from multiple viewpoints, significantly improving the success rate and generalization ability of the attack. By optimizing from multiple viewpoints, the stability and multi-view consistency of the adversarial samples from different viewpoints are ensured.
[0065] 4. Efficiency: By utilizing the efficient rendering capability of Gaussian splatting technology, the present invention can quickly generate adversarial samples in actual scenes, significantly reducing the computational cost compared to NeRF, and achieving efficient three-dimensional scene reconstruction and multi-perspective continuous rendering of adversarial samples.
[0066] 5. Wide applicability: The present invention is applicable to target detection and other deep learning-based visual tasks, and can be used for security assessment, privacy protection, and adversarial testing of artificial intelligence systems. By generating and optimizing 3D adversarial perturbations in 3D scenes, the security and robustness of the system can be tested in a variety of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 The figure is an overall flow chart of the multi-view 3D counter-attack method based on Gaussian splashing of the present invention.
[0068] Figure 2 , Figure 3 and Figure 4 The figure shows the details of the multi-view 3D counter-attack method based on Gaussian splashing of the present invention, as well as a comparison chart of the visualization results with other methods.
[0069] Figure 5 and Figure 6This is a comparison chart of the results of the multi-view 3D adversarial attack method based on Gaussian splashing of the present invention and other methods on multiple data sets. DETAILED DESCRIPTION
[0070] The multi-view 3D adversarial attack method based on Gaussian splashing of the present invention generates 3D adversarial samples that are adversarial in multiple viewpoints while maintaining good natural visual effects for attacking deep classification networks by constructing the original 3D scene, calculating the contribution of Gaussian to the attack, constructing 3D perturbations and optimizing 3D perturbations.
[0071] The invention will be further described below in conjunction with specific embodiments.
[0072] Embodiment 1:
[0073] The original 3D scene is constructed by Gaussian splashing technology using image data collected from multiple perspectives. The Gaussian parameter set of the original 3D scene includes mean μ, covariance Σ, transparency α and color c.
[0074] like Figure 1 As shown in the calculation saliency map, the gradient contribution of each Gaussian in the original 3D scene to the target classifier output is calculated. Specifically, for each image from each perspective, the adversarial loss uses the cross entropy loss as the objective function. The corresponding images are rendered from multiple perspectives, and the adversarial losses of all perspectives are summed to obtain the total loss. For the total loss, the gradient of each Gaussian parameter is calculated separately. In order to reflect the contribution of each Gaussian in the overall adversarial attack, the contribution of each parameter gradient is calculated using the normalization method. The normalized gradients of all parameters in each Gaussian are added to obtain the contribution of each Gaussian.
[0075] like Figure 1 As shown in Gaussian splitting, the selected high-contribution Gaussian is split to generate a small 3D perturbation. In the splitting operation, the mean, color, and opacity are kept consistent with the original Gaussian; the variance of the perturbation is set to the minimum value of all Gaussians in order to obtain a small Gaussian.
[0076] like Figure 1As shown in the loss section, rendering loss, adversarial loss, position loss and color loss are used to optimize the 3D perturbation to ensure that the perturbation can maintain both adversarial effect and visual naturalness under multiple viewpoints. Specifically, for each viewpoint, the rendered image is classified by the target classifier, and the cross entropy loss function is used to measure the difference between the classification result of the rendered image and the preset target category. The adversarial losses of all viewpoints are accumulated to obtain the total adversarial loss. In order to prevent the perturbation from deviating from the target surface, the position loss is introduced to constrain the position of the Gaussian. The color loss is defined to ensure that the perturbation is highly consistent with the original object in color. The above three losses are weighted by hyperparameters to form the total optimization loss function. The gradient of the perturbation parameters is calculated by backpropagating the 3D perturbation parameters, and the gradient descent method is used to update the perturbation parameters. The optimized Gaussian is combined for final rendering, and the adversarial sample is output, and its ability to deceive the target classifier is verified under multiple viewpoints.
[0077] The present invention provides a multi-view 3D adversarial attack method based on Gaussian splashing, which uses Gaussian splashing technology to achieve efficient reconstruction of 3D adversarial samples. It generates small and efficient 3D perturbations by calculating and optimizing the gradient contribution of Gaussian, and achieves attack effects and ensures visual quality through optimization of multiple losses. The present invention not only ensures that the generated adversarial samples have good generalization ability against classifiers under multiple viewpoints, but also are highly natural visually with the original 3D objects.
[0078] Figure 2 , Figure 3 and Figure 4 The figure shows the details of the multi-view 3D adversarial attack method based on Gaussian splashing of the present invention, as well as the comparison of the visualization results with other methods. It can be seen that the multi-view 3D adversarial attack method based on Gaussian splashing is significantly better than other methods in terms of visual naturalness, and the generated adversarial samples are almost imperceptible, thus achieving efficient deception effects under multiple perspectives.
[0079] Figure 5 and Figure 6 The figure is a comparison chart of the results of the multi-view 3D adversarial attack method based on Gaussian splashing of the present invention and other methods on multiple data sets. It can be seen that the 3D adversarial object generation method based on 3D Gaussian splashing has better deception ability than other methods and in multiple viewpoints.
[0080] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0081] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A multi-view 3D adversarial attack method based on Gaussian splashing, characterized in that: The following steps are involved: S1, using multi-view images, the original 3D scene G(Θ) is constructed using the Gaussian splashing technique (where Θ represents a set of Gaussian parameters, including mean μ, covariance Σ, transparency α, and color c); S2, contribution calculation, calculate the gradient contribution of each Gaussian in the original 3D scene to the target classifier output, and select the Gaussian with higher contribution; S3, construct 3D perturbation, split the selected high-contribution Gaussian, and generate small 3D perturbations, where the parameters of the split 3D perturbation are different from the original Gaussian only in scale; S4, optimizes 3D perturbations, using rendering loss, adversarial loss, position loss, and color loss to optimize 3D perturbations, ensuring that the perturbations can maintain both adversarial effects and visual naturalness under multiple viewing angles; S5, obtain 3D adversarial samples and verify their ability to deceive the target classifier under multiple views.
2. The multi-view 3D counterattack method based on Gaussian splashing as claimed in claim 1, characterized in that: The step S2 of calculating the contribution degree specifically includes the following steps: S21, for each image I at each viewing angle v, the adversarial loss uses the following cross entropy loss as the objective function, defined as: Where CE represents the cross entropy loss function, f(I) represents the predicted output of the target classifier for image I, and y tar is the preset target category; S22, render the corresponding image I(v,G(Θ)) from multiple perspectives v∈V, sum the adversarial losses of all perspectives, and get the total loss: S23, for the total loss J(Θ), calculate the gradient of each Gaussian parameter Θ = {μ, Σ, α, c} respectively, and get the gradient It can be expressed as: S24, in order to reflect the contribution of each Gaussian in the overall adversarial attack, the normalization method is used to calculate the gradient of each parameter The contribution of is normalized using the following function: Where p∈{μ,Σ,α,c}, |*| represents the absolute value function, min(*) and max(*) represent the minimum and maximum values calculated, respectively; S25, add the normalized gradients of all parameters in each Gaussian to get the contribution of each Gaussian 3. The multi-view 3D counterattack method based on Gaussian splashing as claimed in claim 1, characterized in that: The step S3 of constructing the 3D perturbation specifically includes the following steps: S31, Select Contribution The top k Gaussians are used as split objects to construct the 3D perturbation δ = {δ μ ,δ Σ ,δ α ,δ c }; S32, in the splitting operation, the mean, color and opacity are kept consistent with the original Gaussian. This part can be expressed as: S33, the variance of the disturbance is set to the minimum value among all Gaussians in order to obtain a tiny Gaussian, which can be expressed as: d Σj =min(Σ) (7) Among them, min(*) represents the operation of calculating the minimum value.
4. The multi-view 3D counterattack method based on Gaussian splashing as claimed in claim 1, characterized in that: The step S4 of optimizing the 3D disturbance specifically comprises the following steps: S41, after splitting out the perturbation, the 3D object parameters are expressed as S42, for each view v, the rendered image I(v,G(Θ′)) is classified by the target classifier f, and the cross entropy loss function is used to measure the difference between the classification result of the rendered image and the preset target category y_tar. The adversarial loss of all view angles v∈V is accumulated, and the total adversarial loss is defined as: S43, in order to prevent the perturbation from deviating from the target surface, the position loss is introduced to constrain the position of the Gaussian. Let the Gaussian mean of the 3D perturbation be δ μ , the corresponding Gaussian mean in the original object is μ, then the position loss can be measured using the chamfer distance function To measure the difference between the two, it is defined as: S44, define color loss to ensure that the perturbation is highly consistent with the original object in color. Map the 3D perturbation from the world coordinate system to the camera coordinate system. The transformation process is as follows: u=W·δ μ (10) Where W is the observation matrix. Normalize u to get the position in the camera coordinate system. The ray r passing through the three-dimensional perturbation can be obtained according to the following formula: r=o+γ·d (12) Where o is the camera origin and γ is the distance parameter. The color corresponding to the 3D perturbation is rendered as C(r), and the color loss of the 3D perturbation can be expressed as: in, represents the set of rays of all three-dimensional perturbations in all viewpoints, and gt represents the true color of the corresponding ray; S45, after weighting the above three losses with hyperparameters, a total optimization loss function is formed: Among them, λ adv , pos , rend and λ col are the hyperparameters for adjusting adversarial, position, rendering, and color losses, respectively; S46, through Perform back propagation, calculate the gradient of the perturbation parameter δ, and use the gradient descent method to update the perturbation parameter δ, that is: Among them, ∈ δ Represents the learning rate.