A non-rigid 3D editing method and system based on cross-modal attention guidance

By employing a differential weighting strategy and image-guided technology, the problem of multi-view consistency in non-rigid 3D editing was solved, enabling efficient and accurate non-rigid 3D editing and improving the flexibility and quality of editing.

CN120655802BActive Publication Date: 2025-11-14NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511160558.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-14
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing 3D editing techniques suffer from multi-view consistency issues in non-rigid editing tasks, and using the original model as the base model limits the ability to modify geometric features.

Method used

A camera view selection strategy guided by differential distribution weights is adopted to generate an anti-rendering basis model that conforms to the geometric features of the target result. By calculating the similarity between the editing results of each view and the guiding view, weighted optimization is performed using optimization weights. Editing is then performed by combining image-guided 3D generation technology and an Inversion-free diffusion model.

Benefits of technology

It improves the non-rigid editing capabilities and multi-view consistency of 3D editing, reduces computational load, and improves editing accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655802B_ABST
    Figure CN120655802B_ABST
Patent Text Reader

Abstract

This invention discloses a non-rigid 3D editing method and system based on cross-modal attention guidance. First, using the original Gaussian sputtering model and several different viewpoints, original frames are rendered to generate original image data. Then, a two-dimensional diffusion model is used for the first round of editing, selecting the image that best matches the expected result as the guide frame. Next, a single-image 3D generation method is used to generate an inverse rendering base model. Then, the generated 3D model is used to perform reverse and forward rendering on several cross-modal attention maps generated during the secondary editing process of the original frame to optimize the editing consistency across multiple views. Finally, the final edited frame generated from the secondary editing is used to optimize the original Gaussian sputtering model, resulting in the edited Gaussian sputtering model. This invention improves the effectiveness of the cross-modal attention guidance mechanism in non-rigid editing tasks by using a cross-modal attention map inverse rendering base model that better matches the editing goals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D editing, and mainly to a non-rigid 3D editing method and system based on cross-modal attention guidance. Background Technology

[0002] Against the backdrop of rapid development in digital content creation and computer vision, 3D editing technology is playing an increasingly crucial role as an important means of constructing virtual worlds and reconstructing real-world scenes. Traditional 2D image processing has limitations in expressing spatial structures and complex geometric relationships, making it difficult to meet the high demands for immersion and interactivity. In contrast, 3D editing technology allows for flexible and precise manipulation of the geometry, texture, lighting, and animation attributes of 3D models, providing technical support for multiple fields such as virtual reality, augmented reality, game development, film special effects, and industrial design.

[0003] With the introduction of Neural Radiation Fields (NRF), subsequent methods began editing implicit scene representations, extending the editing scope to relatively large scenes. However, limited by guided prior editing capabilities, modifications remained focused on aspects such as color, texture, or object rotation / scaling. In recent years, diffusion models have been used to edit 3D radiation fields by editing images rendered from multiple perspectives and optimizing the underlying 3D model. With the development of diffusion models, NeRF-based editing methods such as Instruct-NeRF2NeRF have begun to incorporate diffusion models to achieve high-quality editing effects. The powerful generative capabilities of diffusion models enable complex 3D editing operations, such as style transfer or altering scene features like character actions or attributes. However, due to the non-rigid nature of multi-view editing... Figure 1 The consistency issue means that current methods have many limitations in non-rigid editing tasks. Existing multi-view unification methods, such as VCEditor, use cross-modal attention map inverse rendering followed by forward rendering, but they use the original model as the base model, which limits non-rigid editing capabilities. Summary of the Invention

[0004] Objective: To address the problems existing in the aforementioned background technology, this invention provides a non-rigid 3D editing method and system based on cross-modal attention guidance. Considering that using the original model as the basis for cross-modal attention map inverse rendering would greatly limit geometric feature modification during editing, this method introduces image-guided 3D generation technology to generate an inverse rendering base model that conforms to the geometric features of the target result. To maximize the quality and generation speed of the inverse rendering base model, a differentially distributed weighted camera view selection strategy is adopted to ensure the diversity and effectiveness of the selected series of camera views. When propagating the 2D editing results to the 3D model, this invention employs a weighted optimization strategy. By calculating the similarity between the editing results of each view and the guiding view, an optimization weight is obtained. This weight is used for weighted optimization, making the editing result closer to the reference view selected by the user.

[0005] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows:

[0006] A non-rigid 3D editing method based on cross-modal attention guidance includes the following steps:

[0007] Step S1: Using the original Gaussian sputtering model and several different viewpoints, render the original frame to generate the original image data;

[0008] Step S2: Use a two-dimensional diffusion model to perform the first round of editing on the original frame. The user or the system will select the frame that best meets the expectations as the guide frame. Use a single-image 3D generation method with automatic selection and optimization of the viewpoint to generate the anti-rendering base model.

[0009] Step S3: Use the inverse rendering base model to reverse render several two-dimensional text cross-modal attention maps generated during the secondary editing process of the original frame to obtain Gaussian point cloud-text prompt word cross-modal attention, and use Gaussian point cloud-text prompt word cross-modal attention to render two-dimensional text cross-modal attention maps;

[0010] Step S4: Calculate the similarity between each viewpoint image and the reference viewpoint to obtain the optimized weights. Then, use the final edited frame generated from the secondary editing to perform weighted optimization on the original Gaussian sputtering model to obtain the edited Gaussian sputtering model.

[0011] Furthermore, the process of selecting the guide frame in step S2 can be done by the user or automatically, and is determined by a weighted score based on CLIP image-text similarity and frontal view score. The weighted score formula is as follows:

[0012] ;

[0013] in, This is the initial image edited from perspective v. "Prompt" is the editing prompt word. It is the rotation matrix between the viewpoint v and the target's frontal viewpoint obtained using PoseNet, where E is the identity matrix. This indicates that Frobenius normalization is being performed.

[0014] Furthermore, the detailed steps of the single-image 3D generation technique described in step S2 are as follows:

[0015] Step S2.1: Using the coordinate center as the center, uniformly sample the camera viewpoints around different diameters, and use zero123, taking the initial edited image of the reference view as input, to obtain images of these camera viewpoints.

[0016] Step S2.2: Calculate the feature difference index between viewpoints and generate the difference distribution weights. First, extract the image features from each viewpoint:

[0017] ;

[0018] Then calculate the average feature difference:

[0019] ;

[0020] Finally, calculate the weights of the difference distribution:

[0021] .

[0022] Step S2.3: Select several viewpoints with the highest possible differential distribution weights as training viewpoints.

[0023] Step S2.4: Using these perspectives and the images obtained from the corresponding zero123 inference, optimize the Gaussian sputtering model to obtain the cross-modal attention map inverse rendering base model.

[0024] Furthermore, the de-rendering operation performed in step S3 is shown in the following formula:

[0025] ;

[0026] in, V represents the cross-modal attention score for the i-th Gaussian ellipsoid, where V is the set of views. The original cross-modal attention map representing viewpoint v, where p represents... One pixel in , This represents the cross-modal attention score at pixel p. Opacity representing the Gaussian ellipsoid This represents the cumulative opacity of all Gaussian ellipsoids from the Gaussian ellipsoid to pixel p. The pixel p of the viewpoint v is a weight based on depth variation, and it is calculated as follows:

[0027] ;

[0028] in Let λ be the depth value of pixel p at viewpoint v, and λ be a hyperparameter for adjusting the intensity, typically set to 20. Introducing depth variation weights helps to place greater trust in geometrically more stable regions (pixels with small depth variations), reducing the impact of occlusion boundaries or noisy pixels, thereby improving cross-viewpoint consistency and editing quality.

[0029] Furthermore, the optimization weights are calculated in step S4 as follows:

[0030] ;

[0031] The loss function is:

[0032] ;

[0033] in, The edited image is for reference view. This represents the edited image from the i-th viewpoint. This represents the original image from the i-th viewpoint.

[0034] The present invention also provides a system for implementing the method, comprising:

[0035] Rendering module: Performs rendering of raw frames from multiple perspectives;

[0036] Basis generation module: Integrates zero123 and difference weight view sampler to construct inverse rendering basis model;

[0037] Cross-modal engine: Implements inverse rendering of attention maps with depth weights;

[0038] Optimization module: Calculates weights based on CLIP similarity and optimizes the Gaussian sputtering model.

[0039] Preferably, the difference weighted view sampler dynamically selects training views by using CLIP feature difference distribution weights.

[0040] The 3D generation technology used in this invention is employed to generate a 3D guided substrate. 3D generation uses text or images as a guide to generate a 3D model. This invention utilizes zero123 (Ruoshi Liu and Rundi Wu and BasileVan Hoorick and Pavel Tokmakov and Sergey Zakharov and Carl Vondrick. Zero-1-to-3: Zero-shot One Image to 3D Object. arXiv, 2303.11328.) to generate the substrate model. The core technology is based on the image generation capability of Stable Diffusion. By introducing a relative viewpoint control mechanism, it enables the generation of new views of a target object from a single image at arbitrary viewpoints, further supporting 3D reconstruction. To achieve this goal, zero123 uses a large-scale synthetic dataset (Objaverse rendered images) for fine-tuning. Training samples include pairs of images and their corresponding relative camera extrinsic parameters (rotation and translation) to learn how to map input images to images at arbitrary new viewpoints. The model employs a modified conditional Latent Diffusion architecture, fusing two conditional channels: first, the CLIP embedding of the input image is concatenated with the viewpoint transformation parameters as high-level semantic conditional information; second, the input image is directly concatenated with the image to be generated on the channel, guiding detail preservation. To achieve control over the viewpoint of the diffusion model, the authors added viewpoint information to the conditional input of U-Net. This structure allows the model to obtain controllable viewpoint generation capabilities without compromising its original image generation capabilities, achieving strong generalization on non-training categories and real images, and can be used for high-quality single-image 3D reconstruction.

[0041] To accommodate the needs of non-rigid editing, this invention introduces Inversion-free (Sihan Xu and Yidong Huang and Jiayi Pan and Ziqiao Ma and Joyce Chai. Inversion-Free Image Editing with Natural Language. arXiv, 2312.04965) as a priori for guiding 2D editing during the view editing stage. This paper proposes InEdit, an image editing framework that does not require inversion, utilizing a novel Diffusion Consistency Model (DDCM) to achieve efficient and high-quality text-guided image operations. Unlike traditional diffusion models that rely on iterative inversion and are prone to reconstruction errors, InEdit employs a non-Markovian forward process, starting with random noise to ensure consistency between the original and edited images without requiring inversion branches. This method, known as virtual inversion, significantly reduces computational overhead and accumulated errors. InEdit integrates a Unified Attention Control (UAC) mechanism, combining cross-attention and mutual self-attention control to optimize semantic and structural information, demonstrating superior editing quality and consistency on PIE-Bench. By incorporating a Latent Consistency Model (LCM), InEdit achieves faster sampling, completing editing in less than 3 seconds on a single A-40 GPU, surpassing baseline models such as Prompt-to-Prompt, SDEdit, and CycleDiff in terms of both quality and efficiency. The framework also excels in complex image-to-image translation tasks, maintains compatibility with large language models, and addresses ethical issues such as copyright infringement and potential abuse.

[0042] Beneficial effects:

[0043] (1) This invention reduces the number of cameras used for optimization by using a camera scoring and selection strategy, thereby reducing the computational load in the single-view reconstruction stage and improving the accuracy and effect of single-view reconstruction. Previous methods used uniform sampling around a single diameter, which may result in too many cameras being distributed in irrelevant areas, causing a waste of computing power and an increase in time costs. This invention reduces the waste of viewing angles by evaluating the differences between the camera's viewpoint and all other viewpoints and selecting the camera group with the greatest differences.

[0044] (2) This invention introduces a more suitable anti-rendering model base, making the 3D editing method applicable to non-rigid 3D editing. Simultaneously, it introduces depth variation weights in the anti-rendering process to improve smoothness. Existing multi-view unification mechanisms use the original model as the attention map anti-rendering base model, achieving good results in rigid editing tasks such as texture and color editing where the original model's geometric features are hardly modified. However, they are severely limited in non-rigid editing tasks, making it almost impossible to modify the geometric features of the target object. The multi-view unification mechanism of this invention makes 3D editing applicable to non-rigid tasks while also possessing high uniformity. Attached Figure Description

[0045] Figure 1 This invention provides a flowchart of the overall process of non-rigid 3D Gaussian editing, which shows the complete process from rendering the original frame to optimizing the final Gaussian sputtering model. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0047] Example 1

[0048] This invention provides a non-rigid 3D editing method based on cross-modal attention guidance, comprising the following steps:

[0049] Step S1: Using the original Gaussian sputtering model and several different viewpoints, render the original frame to generate the original image data;

[0050] Before editing can begin, a series of original frames need to be rendered using the original Gaussian model and a series of optimized perspectives for 2D editing. This invention employs the original 3D Gaussian rendering method, which can render a single image frame in milliseconds.

[0051] Step S2: Input text prompts and provide editing requirements. The Inversion-free diffusion model performs the first round of editing on the original frame based on the prompts. The core advantage of Inversion-free technology is that it eliminates the need for a time-consuming inversion process (i.e., reversing the latent encoding from the target image). Instead, it starts directly from random noise and generates the editing result through a non-Markov forward process, significantly reducing computational complexity and reconstruction error. The editing process is guided by user-provided text prompts (such as "adjust the object's pose to standing"), generating multi-view edited images. Then, the frame that best matches the expectations is selected as the guide frame, and a single-image 3D generation method is used to generate an inverse rendering base model.

[0052] The selection of the guide frame can be done by the user or automatically, and is determined by a weighted score based on CLIP image-text similarity and frontal viewpoint score. The weighted score formula is as follows:

[0053] ;

[0054] in, This is the initial image edited from perspective v. "Prompt" is the editing prompt word. It is the rotation matrix between the viewpoint v and the target's frontal viewpoint obtained using PoseNet, where E is the identity matrix. This indicates that Frobenius normalization is being performed.

[0055] Then, using the edited image of the guide frame as input, a 3D model of a single view is generated:

[0056] Using the coordinate center as the sphere's center, the camera's field of view is uniformly sampled on a sphere with a radius of 1.0 to 2.0 (typically 100-200 sampling points, pitch range [-30°, 30°], azimuth range [0°, 360°]). Using a guide frame as input, zero123 generates an image corresponding to the optimized field of view. zero123 infers the new field of view image using a pre-trained diffusion model, with inference parameters including the guidance scale (typically 7.5) and the number of sampling steps (typically 50 steps). To improve generation quality, a control network (such as ControlNet) can be introduced to ensure geometric consistency.

[0057] To avoid viewpoint redundancy, a differential distribution weighting strategy is used to dynamically select viewpoints.

[0058] First, extract image features from each viewpoint:

[0059] ;

[0060] Calculate the average feature difference:

[0061] ;

[0062] Calculate the weights of the differential distribution:

[0063] ;

[0064] Then, select several viewpoints with the highest possible differential distribution weights as training viewpoints.

[0065] Finally, using these perspectives and the corresponding zero123 inference images, the Gaussian sputtering model is optimized to obtain the cross-modal attention map inverse rendering basis model.

[0066] Step S3: Use Inversion-Free to perform secondary editing on the original frame, and use the inverse rendering base model to perform reverse and forward rendering on several cross-modal attention maps generated during the secondary editing of the original frame to optimize the consistency of multi-view 2D editing.

[0067] The formula for de-rendering is:

[0068] ;

[0069] in, V represents the cross-modal attention score for the i-th Gaussian ellipsoid, where V is the set of views. The original cross-modal attention map representing viewpoint v, where p represents... One pixel in , This represents the cross-modal attention score at pixel p. Opacity representing the Gaussian ellipsoid This represents the cumulative opacity of all Gaussian ellipsoids from the Gaussian ellipsoid to pixel p. The pixel p of the viewpoint v is a weight based on depth variation, and it is calculated as follows:

[0070] ;

[0071] in Let λ be the depth value of pixel p at viewpoint v, and λ be a hyperparameter for adjusting the intensity, typically set to 20. Introducing depth variation weights helps to place greater trust in geometrically more stable regions (pixels with small depth variations), reducing the impact of occlusion boundaries or noisy pixels, thereby improving cross-viewpoint consistency and editing quality.

[0072] Step S4: Calculate the similarity between each viewpoint image and the reference viewpoint, and generate optimized weights. Use the final edited frame generated from the secondary editing to perform weighted optimization on the original Gaussian sputtering model, resulting in the edited Gaussian sputtering model.

[0073] The optimization weights are calculated as follows:

[0074] ;

[0075] The loss function is:

[0076] ;

[0077] in, The edited image is for reference view. This represents the edited image from the i-th viewpoint. This represents the original image from the i-th viewpoint.

[0078] Example 2

[0079] This embodiment provides a system for implementing the method, including:

[0080] 1. Rendering module, whose function is to execute step S1 in embodiment 1;

[0081] The implementation process includes: loading the original Gaussian sputtering model; reading predefined optimized view parameters (pitch angle range [-30°, 30°], azimuth angle range [0°, 360°]); and calling the renderer to generate multi-view original frames.

[0082] Its output is: the original frame sequence.

[0083] 2. The basis generation module performs step S2 in Example 1. It integrates the zero123 3D generation engine and the difference weight view sampler. The workflow includes: receiving the guide frame; sampling 200 candidate views on a sphere with a radius of 1.0-2.0; calculating the CLIP feature difference distribution weight; selecting the view with the highest weight as the training view; and optimizing the generation of the anti-rendering basis model.

[0084] 3. Cross-modal engine, whose function is to execute step S3 of embodiment 1, realize the inverse rendering of attention map with depth weights, and output the optimized multi-view cross-modal attention map.

[0085] 4. Optimization module, whose function is to execute step S4 of embodiment 1. Its workflow includes calculating view similarity weights; constructing a weighted loss function; iteratively updating the Gaussian sputtering model using the Adam optimizer; and finally outputting the edited Gaussian sputtering model.

Claims

1. A non-rigid 3D editing method based on cross-modal attention guidance, characterized in that, Includes the following steps: Step S1: Render the original frame using the original Gaussian sputtering model and several different viewpoints; Step S2: Perform initial editing on the original frame using a two-dimensional diffusion model to generate a guide frame. Using this guide frame as input, generate an anti-rendering base model using a single-image three-dimensional generation method. The three-dimensional generation method used is a combination of zero123 and three-dimensional Gaussian sputtering. The specific process is as follows: (1) Fix the corresponding viewpoint of the input image at a certain viewpoint as a reference viewpoint; (2) Uniformly sample the camera's field of view on a sphere of several diameters centered at the center, and use zero123 to generate images of these field of view I. ' ; (3) Calculate the view difference index: First, use CLIP to extract images from various perspectives. i Features: f i =CLIP(I' i ) Then, the average feature difference is calculated for all image pairs: Where N represents the total number of candidate viewpoints, i, j∈[1,N]; (4) Generate differential distribution weights: (5) Select the perspectives with the highest weights in the differential distribution as the optimization perspectives; (6) Use image data from an optimized perspective to train and generate a 3D Gaussian model; Step S3: Perform secondary editing on the original frame to generate a two-dimensional text cross-modal attention map. Use the anti-rendering base model to perform anti-rendering on the attention map, outputting Gaussian point cloud-text cue word cross-modal attention, and generate an optimized two-dimensional cross-modal attention map through forward rendering. The anti-rendering operation is implemented through the following formula: Among them, G attn (i) represents the cross-modal attention score of the i-th Gaussian ellipsoid, where V is the set of viewpoints. The original cross-modal attention map representing viewpoint v, where p represents... One pixel in , O represents the cross-modal attention score at pixel p. i (p) represents the opacity of the Gaussian ellipsoid, T i (p) represents the cumulative opacity of all Gaussian ellipsoids from the Gaussian ellipsoid to pixel p: α v (p) is the weight of pixel p in viewpoint v based on depth variation, and it is calculated as follows: Where D v (p) represents the depth value of pixel p under the viewpoint v, and λ is a hyperparameter for adjusting the intensity. Step S4: Based on the cross-modal attention, calculate the similarity between each viewpoint image and the reference viewpoint to obtain the optimized weights; then use the final edited frame generated by the secondary editing to perform weighted optimization on the original Gaussian sputtering model to obtain the edited Gaussian sputtering model.

2. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that, The original Gaussian sputtering model in step S1 is a three-dimensional model reconstructed from the image dataset using a three-dimensional Gaussian sputtering method. It is represented as a set of three-dimensional Gaussian ellipsoids with spherical harmonic correction for color. The specific attributes of each three-dimensional Gaussian ellipsoid are: center μ∈R. 3 Color c∈R 3 Opacity α∈R, covariance matrix C=QSS T Q T , where Q∈R 3 Let S be a rotation matrix, and S ∈ R. 3 Given a scaling matrix, the attention attribute atn∈R n This is used to store attention scores, where n is the number of cue words and is consistent with the number of channels in the attention map generated during the diffusion model's operation.

3. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that, The selection of the guiding frame in step S2 is achieved through weighted scoring: Among them, I v This is the initial image edited from perspective v. "Prompt" is the editing prompt word. R v It is the rotation matrix between the viewpoint v and the target's frontal viewpoint obtained using PoseNet, where E is the identity matrix, ||| F This indicates that Frobenius normalization is being performed.

4. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that, The optimization weights in step S4 are calculated using CLIP similarity: The loss function is: Among them, I ref For the edited image of the reference view, I i The edited image for viewpoint i. This represents the original image from the i-th viewpoint.

5. A system for implementing the method of any one of claims 1-4, characterized in that, include: Rendering module: Performs rendering of raw frames from multiple perspectives; Basis generation module: Integrates zero123 and difference weight view sampler to construct inverse rendering basis model; Cross-modal engine: Implements inverse rendering of attention maps with depth weights; Optimization module: Calculates weights based on CLIP similarity and optimizes the Gaussian sputtering model.

6. The system according to claim 5, characterized in that, The difference-weighted view sampler dynamically selects training views by weighting the difference distribution of CLIP features.

Citation Information

Patent Citations

  • Three-dimensional model generation method based on new visual angle texture correction

    CN120182533A

  • Interactive three-dimensional Gaussian editing method and system based on three-dimensional geometry consistent attention priori

    CN120411438A