3D adversarial object generation method based on 3D Gaussian splashing
Through the 3D adversarial object generation method based on 3D Gaussian splash, the problem of poor attacking face recognition system in multi-view angles is solved, and the effect of efficiently deceiving face recognition system in the physical world is achieved, which significantly improves the attack success rate and visual reality.
Patent Information
- Application Number
- CN202510160581.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has poor attacks on face recognition systems from multiple perspectives and is difficult to achieve stable adversarial attacks in the physical world.
Using the 3D adversarial object generation method based on 3D Gaussian splatter, 3D adversarial objects with multi-view consistency and high adversarial performance are generated by constructing the original 3D scene, precise initialization of the target object and Gaussian combination operations, and combining rendering loss, position loss and adversarial loss of meta-learning.
It significantly improves the attack success rate and visual reality of adversarial objects under multi-view conditions, enhances the ability to deceive the face recognition system, simplifies the computational complexity, and realizes real-time processing and fast response.
Smart Images

Figure CN119992628A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and artificial intelligence security, and specifically relates to a 3D adversarial object generation method based on 3D Gaussian splashing, which is used to attack face recognition systems under wider viewing angles. Specifically, the present invention focuses on generating three-dimensional adversarial objects that are consistent in multiple viewing angles, and realizes efficient deception of face recognition systems under dynamic viewing angles in the physical world, which is suitable for security attack and defense testing, privacy protection, and system robustness verification. Background Art
[0002] In the field of adversarial attacks on face recognition, many methods have been proposed in recent years to implement attacks, each with its own advantages and disadvantages. The following are several major adversarial attack methods for face recognition and their characteristics:
[0003] 1. 2D adversarial attack method for face recognition
[0004] Traditional two-dimensional adversarial attack methods, such as using gradient calculation to generate perturbations (such as FGSM, PGD) or generating adversarial samples based on style transfer (such as Adv-Makeup), have achieved certain results under laboratory conditions, but in actual physical scenes, perturbations are prone to failure due to factors such as perspective, lighting, and occlusion. Some technologies interfere with the recognition system by adding two-dimensional perturbations (such as adversarial glasses and hat patches). For example, the adversarial glasses (AdvGlasses) proposed by Sharif et al. are effective from a frontal perspective, but the attack fails due to image distortion when the perspective is offset. Such methods rely on planar perturbations and cannot adapt to changes in perspective in three-dimensional space.
[0005] 2. 3D adversarial attack method for face recognition
[0006] In order to overcome the limitations of two-dimensional adversarial attacks in actual physical scenarios, some researchers have begun to explore adversarial attack methods based on three-dimensional information. Early three-dimensional methods mainly rely on three-dimensional deformable models (3DMM). For example, AT3D reconstructs the geometric structure of the face through 3DMM and generates adversarial patches. The patches generated by AT3D have obvious signs of forgery. In addition, these methods usually only build three-dimensional models based on fixed-view images, which makes it difficult to ensure the stability of attacks under different viewpoints.
[0007] In summary, the current adversarial attacks against face recognition have achieved the attack capabilities of the physical world, but the effect is still poor under multiple perspectives. In recent years, 3D Gaussian Splatting (3DGS) technology has achieved efficient 3D scene reconstruction and multi-perspective continuous rendering by representing the 3D scene as a series of spheres obeying Gaussian distribution, providing a new idea for constructing adversarial objects. However, how to achieve accurate initialization of the target object, effective combination of the original scene and the target object, and design optimization strategies that take into account both visual realism and adversarial performance under this framework are still key technical issues that need to be solved. Summary of the invention
[0008] In view of the above problems, the present invention proposes a 3D adversarial object generation method based on 3D Gaussian splashing, and the method comprises the following steps:
[0009] 1. A 3D adversarial object generation method based on 3D Gaussian splashing, characterized by comprising the following steps:
[0010] S1, constructs the original 3D scene based on the multi-view image data using the 3D Gaussian splashing technique.
[0011] S2, initialize the adversarial object, generate a multi-view consistent image containing the target object through a text-based image editor, extract the target object mask by image segmentation, and determine the key Gaussian of the target object based on the projection ray and Gaussian contribution of the mask area, thereby initializing the adversarial object;
[0012] S3, performing Gaussian combination on the original 3D scene and the initialized target object through a Gaussian combination operation to obtain an overall scene containing the object;
[0013] S4, optimizes adversarial objects, uses rendering loss to optimize the rendering of the target object, and constrains the parameters of the target object through Gaussian position loss and meta-learning-based adversarial loss, so that it maintains realistic appearance and excellent adversarial performance under different viewing angles;
[0014] S5, adversarial object generation, combined with the optimized Gaussian for final rendering, outputs the adversarial object, and verifies its ability to deceive the facial recognition system through multi-view testing.
[0015] 2. A 3D adversarial object generation method based on 3D Gaussian splashing according to claim 1, characterized in that the object initialization phase of step S2 specifically comprises the following steps:
[0016] S21, using a pre-trained editor (e.g., DreamCatalyst) to generate multi-view consistent edited images containing the desired objects (e.g., glasses or masks) according to a given text prompt (e.g., “put on a pair of glasses for him” or “put on a mask for him”) to determine the appropriate position and shape of the objects;
[0017] S22, segmenting the edited image using a pre-trained segmentation model (e.g., SAM2) to obtain an object mask;
[0018] S23, according to the mapping relationship between the mask and the object, calculate the projection ray corresponding to the mask area, and the ray direction in the camera coordinate system can be expressed as:
[0019]
[0020] Among them, K -1 is the inverse matrix of the camera intrinsic matrix, which maps the normalized image coordinates to the camera space, and x and y are the normalized coordinates of the mask. The set of all rays can be expressed in the following form:
[0021] R M (t) = o + td (14)
[0022] Where o is the origin of the ray (the location of the camera) and t is a scalar parameter along the ray.
[0023] S24, based on the contribution of each Gaussian on the ray, the Gaussian contribution can be defined as:
[0024]
[0025] Among them, c i represents the color of the i-th Gaussian, α i represents the opacity of the i-th Gaussian. Find the Gaussian g that contributes most to this ray i , and locate the set of all Gaussians that contribute the most to all rays, expressed as follows:
[0026]
[0027] S25, by extracting these high-contribution Gaussians, achieves accurate initialization of the target object;
[0028] 3. The 3D adversarial object generation method based on 3D Gaussian splashing according to claim 1, characterized in that the object optimization stage in step S4 specifically comprises the following steps:
[0029] S41, using the scene rendering loss to constrain the difference between the rendered image and the edited image of the entire scene, specifically expressed as:
[0030] L SR =λ s1 L1(I,E)+λ sp L LPIPS (I,E) (17)
[0031] Among them, L1 represents L1 loss, L LPIPS represents the LPIPS loss, I is the rendered image of the entire scene, E is the edited image, and λ s1 and λ sp is a hyperparameter;
[0032] S42, using the object rendering loss to constrain the difference between the rendered image of the object scene and the segmented object image, specifically expressed as:
[0033] L OR =λ o1 L1(I O ,E⊙M)+λ op L LPIPS (I O ,E⊙M) (18)
[0034] Among them, I O is the rendered image of the object, M is the mask of the object, ⊙ represents the dot product operation, λ o1 and λ op is a hyperparameter;
[0035] S43, using Gaussian position loss to constrain the position of Gaussians in the object to prevent Gaussians from detaching from the surface of the object, specifically expressed as:
[0036] L POS =D c (μ,μ s ) (19)
[0037] Among them, D c represents the chamfer distance metric function, μ is the Gaussian mean of the adversarial object, i.e., the position of the Gaussian, μ s is the position of the point set after statistical outlier removal (SOR) processing;
[0038] S44, using meta-learning-based adversarial loss to improve the generalization ability of attacks on multiple networks. For L face recognition models {F 1 ,F 2 ,···,F L}, randomly divided into (L-1) meta-training model sets and one meta-testing model. In meta-training, the adversarial loss of each model is as follows:
[0039]
[0040] Where i∈{1,L-1}, I(θ) represents the rendered image when the parameters of the adversarial object are θ, and cos represents the cosine distance function. For each meta-trained model, the original parameters θ can be optimized to obtain their corresponding temporary updated parameters θ i ′:
[0041]
[0042] Among them, β is the learning rate, ▽ θ Represents the gradient. For each meta-trained model updated parameter θ i ′, calculate the meta-test model F L Loss:
[0043]
[0044] The final adversarial loss can be expressed as:
[0045]
[0046] S45, combining the above rendering, position and adversarial losses, the total loss is defined as follows:
[0047] L total =λ SR L SR +λ OR L OR +λ POS L POS +λ ADV L ADV (twenty four)
[0048] Among them, λ SR ,λ OR ,λ POS and λ ADV is a hyperparameter.
[0049] Beneficial effects: Compared with the prior art, the present invention provides a 3D adversarial object generation method based on 3D Gaussian splashing, which can deceive the face recognition system at a wider viewing angle and produce the following beneficial effects:
[0050] 1. Strong versatility: Existing two-dimensional adversarial attack methods often fail to achieve perturbation failure in actual physical scenes due to factors such as perspective, lighting, and occlusion. The present invention uses 3D Gaussian splash technology to construct the original three-dimensional scene, and combines the precise initialization of the target object and Gaussian combination operations to enable the generated adversarial object to adapt to a variety of perspective changes, thereby maintaining a stable attack effect at different shooting angles, greatly improving the universality of the adversarial attack.
[0051] 2. High attack success rate: By utilizing precise target object initialization strategies and multiple optimizations of joint rendering losses (such as LPIPS and L1 losses), as well as meta-learning-based adversarial losses, the present invention significantly improves the visual realism and adversarial performance of adversarial objects. Verified by multi-view tests, its deception rate on facial recognition systems is significantly higher than that of traditional two-dimensional methods and early three-dimensional adversarial methods.
[0052] 3. High computational efficiency: The present invention adopts efficient 3D Gaussian splashing technology to realize the rapid construction of original three-dimensional scenes and multi-perspective continuous rendering, which greatly simplifies the traditional 3D modeling process and significantly reduces the computational complexity compared to 3DMM, thereby achieving real-time processing and rapid response in practical applications.
[0053] 4. Unified attack framework: This paper proposes a unified adversarial object generation framework, which can complete all the tasks from multi-view image editing, precise initialization of target objects to Gaussian combination and optimization of the overall scene in the same framework. There is no need to design and train adversarial samples for different viewpoints separately, which simplifies the overall process and reduces the difficulty of training and deployment of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is an overall flow chart of the 3D adversarial object generation method based on 3D Gaussian splashing of the present invention.
[0055] Figure 2 and Figure 3 This is a visualization result diagram and detail comparison diagram of the 3D adversarial object generation method based on 3D Gaussian splashing of the present invention and other methods.
[0056] Figure 4 This is a comparison chart of the results of the 3D adversarial object generation method based on 3D Gaussian splashing of the present invention and other methods on the man, girl, fangzhou and person datasets. DETAILED DESCRIPTION
[0057] The 3D adversarial object generation method based on 3D Gaussian splashing of the present invention generates a 3D adversarial object with high adversarial performance and natural visual effects under multi-view conditions by constructing an original 3D scene, accurately initializing the target object, and performing multiple optimizations on the overall scene, so as to attack the face recognition system.
[0058] The invention will be further described below in conjunction with specific embodiments.
[0059] Embodiment 1:
[0060] In the initial stage, facial image data from different perspectives are collected first. These image data can fully reflect the detailed features of the characters in the scene. Using 3D Gaussian splash technology, the collected multi-perspective image data is reconstructed to construct the original 3D scene. This step efficiently reconstructs the geometric information of the face in the scene, providing an accurate geometric and texture basis for the subsequent generation of adversarial objects.
[0061] like Figure 1 As shown in the object initialization stage, according to the preset text prompts (such as "put on a pair of glasses for him" or "put on a mask for him"), a set of multi-view consistent edited images containing the target object (such as glasses or masks) is generated using the pre-trained editor DreamCatalyst. The generation of the edited images aims to clearly indicate the position and shape of the target object in the overall image, providing an intuitive reference for the positioning of the target object. Subsequently, the edited images are segmented using the pre-trained image segmentation model SAM2 to obtain the mask information of the target object. The most relevant Gaussian corresponding to the mask is found through the mapping relationship between the mask area and the Gaussian representation in the three-dimensional scene. This process utilizes the geometric relationship between projection and back-projection, thereby achieving accurate mapping of two-dimensional segmentation information to three-dimensional space. Finally, by calculating and screening the contribution of the mapping results, the most representative key Gaussian for the target object is extracted, thereby completing the accurate initialization of the target object. This process ensures that the target object is accurately positioned and naturally shaped in the three-dimensional scene, laying the foundation for subsequent optimization.
[0062] After the target object is initialized, the purpose of this stage is to effectively fuse the original 3D scene with the initialized target object to facilitate subsequent optimization. First, the Gaussian combination operation is performed on the original constructed 3D scene and the initialized target object so that the target object is naturally integrated into the overall scene. This operation ensures that the geometry and texture information of the original scene are not modified, and also gives the target object continuity and consistency in the 3D space, so that subsequent object optimization can be carried out in the overall scene.
[0063] like Figure 1As shown in the object optimization stage, in order to ensure that the generated adversarial objects can achieve attack effects and high visual quality under different viewing angles, multiple optimization mechanisms are introduced. First, the appearance consistency between the overall scene and the edited image is constrained by the scene rendering loss, so that the overall rendered image is highly matched with the expected editing effect in terms of color and overall vision; secondly, the object rendering loss is introduced to constrain the local details of the target object with the segmented object image to ensure the consistency of the target object in details; thirdly, in order to prevent the Gaussian in the target object from being offset or separated from the object surface due to the optimization process, the Gaussian position loss is used to constrain the position of the Gaussian to ensure that it remains stable in the overall object structure; finally, the adversarial loss based on meta-learning is used to further improve the adversarial and generalization capabilities of the target object by training and testing on multiple facial recognition models. After the above multiple optimizations, the appearance of the target object in the overall scene has both natural and realistic visual effects and extremely high adversarial attack performance.
[0064] The present invention provides a 3D adversarial object generation method based on 3D Gaussian splashing, which uses 3D Gaussian splashing technology and pre-trained models to accurately initialize the target object, and then effectively merges the original scene with the target object through Gaussian combination and multiple optimization mechanisms to generate an adversarial object with multi-view attack capabilities. The whole process not only ensures the geometric and visual authenticity of the adversarial object, but also significantly improves the robustness and generalization ability of the adversarial attack in physical scenes.
[0065] Figure 2 and Figure 3 The visualization result diagram and detail comparison diagram of the 3D adversarial object generation method based on 3D Gaussian splashing of the present invention and other methods. It can be seen that the 3D adversarial object generation method based on 3D Gaussian splashing shows superior naturalness in appearance compared with other methods and is seamlessly integrated into the visual environment.
[0066] Figure 4 The figure is a comparison of the results of the 3D adversarial object generation method based on 3D Gaussian splashing of the present invention and other methods on the man, girl, fangzhou and person datasets. It can be seen that the 3D adversarial object generation method based on 3D Gaussian splashing is more effective than other methods in terms of the success rate of multi-view attacks on various datasets.
[0067] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0068] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A 3D adversarial object generation method based on 3D Gaussian splashing, characterized in that: The following steps are involved: S1, constructs the original 3D scene based on the multi-view image data using the 3D Gaussian splashing technique. S2, initialize the adversarial object, generate a multi-view consistent image containing the target object through a text-based image editor, extract the target object mask by image segmentation, and determine the key Gaussian of the target object based on the projection ray and Gaussian contribution of the mask area, thereby initializing the adversarial object; S3, performing Gaussian combination on the original 3D scene and the initialized target object through a Gaussian combination operation to obtain an overall scene containing the object; S4, optimizes adversarial objects, uses rendering loss to optimize the rendering of the target object, and constrains the target object parameters through Gaussian position loss and meta-learning-based adversarial loss, so that it maintains realistic appearance and excellent adversarial performance under different viewing angles; S5, adversarial object generation, combined with the optimized Gaussian for final rendering, outputs the adversarial object, and verifies its ability to deceive the facial recognition system through multi-view testing.
2. The 3D adversarial object generation method based on 3D Gaussian splashing according to claim 1, characterized in that: The object initialization phase of step S2 specifically includes the following steps: S21, using a pre-trained editor (e.g., DreamCatalyst) to generate multi-view consistent edited images containing the desired objects (e.g., glasses or masks) according to a given text prompt (e.g., "put on a pair of glasses for him" or "put on a mask for him"), which is used to determine the appropriate position and shape of the objects; S22, segmenting the edited image using a pre-trained segmentation model (e.g., SAM2) to obtain an object mask; S23, according to the mapping relationship between the mask and the object, calculate the projection ray corresponding to the mask area, and the ray direction in the camera coordinate system can be expressed as: Among them, K -1 is the inverse matrix of the camera intrinsic matrix, which maps the normalized image coordinates to the camera space, and x and y are the normalized coordinates of the mask. The set of all rays can be expressed in the following form: R M (t)=o+td (2) Where o is the origin of the ray (the location of the camera) and t is a scalar parameter along the ray. S24, based on the contribution of each Gaussian on the ray, the Gaussian contribution can be defined as: Among them, c i represents the color of the i-th Gaussian, α i represents the opacity of the i-th Gaussian. Find the Gaussian g that contributes most to this ray i , and locate the set of all Gaussians that contribute the most to all rays, expressed as follows: S25, by extracting these high-contribution Gaussians, accurate initialization of the target object is achieved.
3. The 3D adversarial object generation method based on 3D Gaussian splashing according to claim 1, characterized in that: The object optimization stage of step S4 specifically includes the following steps: S41, using the scene rendering loss to constrain the difference between the rendered image and the edited image of the entire scene, specifically expressed as: L SR λ s1 L1(I,E)+λ sp L LPIPS (I,E) (5) Among them, L1 represents L1 loss, L LPIPS represents the LPIPS loss, I is the rendered image of the entire scene, E is the edited image, and λ s1 and λ sp is a hyperparameter; S42, using the object rendering loss to constrain the difference between the rendered image of the object scene and the segmented object image, specifically expressed as: L OR =l o1 L1(I O ,E⊙M)+λ op L LPIPS (YO O ,E⊙M) (6) Among them, I O is the rendered image of the object, M is the mask of the object, ⊙ represents the dot product operation, λ o1 and λ op is a hyperparameter; S43, using Gaussian position loss to constrain the position of Gaussians in the object to prevent Gaussians from detaching from the surface of the object, specifically expressed as: L POS =D c (m,m s ) (7) Among them, D c represents the chamfer distance metric function, μ is the Gaussian mean of the adversarial object, i.e., the position of the Gaussian, μ s is the position of the point set after statistical outlier removal (SOR) processing; S44, using meta-learning-based adversarial loss to improve the generalization ability of attacks on multiple networks. For L face recognition models {F 1 ,F 2 ,···,F L }, randomly divided into (L-1) meta-training model sets and one meta-testing model. In meta-training, the adversarial loss of each model is as follows: Where i∈{1,L-1}, I(θ) represents the rendered image when the parameters of the adversarial object are θ, and cos represents the cosine distance function. For each meta-trained model, the original parameters θ can be optimized to obtain their corresponding temporary updated parameters θ i ′: Among them, β is the learning rate, ▽ θ Represents the gradient. For each meta-trained model updated parameter θ i ′, calculate the meta-test model F L Loss: The final adversarial loss can be expressed as: S45, combining the above rendering, position and adversarial losses, the total loss is defined as follows: L total =λ SR L SR +λ OR L OR +λ POS L POS +λ ADV L ADV (12) Among them, λ SR ,λ OR ,λ POS and λ ADV is a hyperparameter.
Citation Information
Cited By
Face identity protection method and system for three-dimensional Gaussian point cloud scene editing
CN122391577A