Non-rigid three-dimensional editing method and system based on cross-modal attention guidance

By guiding the difference distribution weights and depth change weights, the multi-view consistency problem of non-rigid editing in existing 3D editing is solved, and efficient and accurate non-rigid 3D editing is achieved, which is suitable for complex 3D editing tasks.

CN120655802AActive Publication Date: 2025-09-16NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511160558.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-09-16
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing 3D editing technologies suffer from multi-view consistency issues in non-rigid editing tasks, and using the original model as the base model limits the ability to modify geometric features during editing.

Method used

A camera perspective selection strategy guided by difference distribution weights is adopted. By calculating the differences and similarities of each perspective, a reverse rendering basis model that meets the target results is generated. Deep change weights are introduced to optimize the editing results, and cross-modal attention maps are used for reverse rendering and optimization.

Benefits of technology

It improves the effect and consistency of non-rigid 3D editing, reduces the amount of calculation, is suitable for complex 3D editing tasks, and improves the accuracy and smoothness of editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655802A_ABST
    Figure CN120655802A_ABST
Patent Text Reader

Abstract

The invention discloses a non-rigid three-dimensional editing method and system based on cross-modal attention guidance. The method comprises the following steps: firstly, rendering an original frame by using an original Gaussian sputtering model and a plurality of different visual angles, and generating original image data; secondly, performing first-round editing by using a two-dimensional diffusion model, and selecting a picture which is most in accordance with expectation as a guide frame; a single-image three-dimensional generation method is used to generate an anti-rendering base model; using the generated three-dimensional model to perform reverse rendering and forward rendering on a plurality of cross-modal attention maps generated in the secondary editing process of the original frame to optimize the editing consistency of multiple views; and finally, optimizing the original Gaussian sputtering model by using a final editing frame generated by secondary editing to obtain an edited Gaussian sputtering model. According to the method, the effect of the cross-modal attention guidance mechanism under the non-rigid editing task is improved by using the cross-modal attention map reverse rendering base model which better conforms to the editing target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional editing, and mainly relates to a non-rigid three-dimensional editing method and system based on cross-modal attention guidance. Background Art

[0002] Amid the rapid development of digital content creation and computer vision, 3D editing technology is playing an increasingly critical role as a key means of constructing virtual worlds and reconstructing real-world scenes. Traditional 2D image processing has limitations in expressing spatial structures and complex geometric relationships, making it difficult to meet the high demands for immersion and interactivity. 3D editing technology, on the other hand, enables flexible and sophisticated manipulation of 3D model attributes such as geometry, textures, lighting, and animation, providing technical support for a wide range of fields, including virtual reality, augmented reality, game development, film special effects, and industrial design.

[0003] With the introduction of neural radiance fields, subsequent methods began to edit implicit scene representations, which expanded the editing objects to relatively large scenes. However, limited by the editing capabilities of guided priors, modifications are still concentrated in aspects such as color, texture, or rotation / scaling of objects. In recent years, diffusion models have been used to edit 3D radiance fields by editing images rendered from multiple perspectives and optimizing the underlying 3D model. With the development of diffusion models, NeRF-based editing methods such as Instruct-NeRF2NeRF began to combine diffusion models to achieve high-quality editing effects. The powerful generation capability of the diffusion model makes complex 3D editing operations possible, such as style transfer, or changing scene features such as character actions or attributes. However, due to the multi-view of non-rigid editing Figure 1 Due to consistency issues, current methods have many limitations in non-rigid editing tasks. Existing multi-view unified methods, such as VCEditor, use a cross-modal attention map reverse rendering technique, but use the original model as the base model, which limits the non-rigid editing capabilities. Summary of the Invention

[0004] Purpose of the invention: In response to the problems existing in the above-mentioned background technology, the present invention provides a non-rigid three-dimensional editing method and system based on cross-modal attention guidance. Taking into account that using the original model as the basis for cross-modal attention map de-rendering will greatly limit the modification of geometric features during editing, the method introduces image-guided three-dimensional generation technology to generate a de-rendering base model that conforms to the geometric features of the target result. In order to maximize the quality and generation speed of the de-rendering base model, we adopted a camera perspective selection strategy guided by difference distribution weights to ensure the difference and effectiveness of the selected series of camera perspectives. When propagating the two-dimensional editing results to the three-dimensional model, the present invention adopts a weighted optimization strategy. By calculating the similarity between the editing results of each view and the guided view, the optimization weight is obtained, and this weight is used for weighted optimization, so that the editing result is closer to the reference view selected by the user.

[0005] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is: A non-rigid 3D editing method based on cross-modal attention guidance includes the following steps: Step S1: using the original Gaussian sputtering model and several different viewing angles, rendering the original frame to generate original image data; Step S2: Perform a first round of editing on the original frame using a two-dimensional diffusion model. A frame that best meets expectations is selected by the user or automatically as a guide frame. A single-image three-dimensional generation method that automatically selects an optimized viewing angle is used to generate a reverse-rendered base model. Step S3: using the inverse rendering basis model to reversely render several two-dimensional text cross-modal attention maps generated during the secondary editing process of the original frame to obtain Gaussian point cloud-text prompt word cross-modal attention, and using the Gaussian point cloud-text prompt word cross-modal attention to render the two-dimensional text cross-modal attention map; Step S4: Calculate the similarity between each view image and the reference view image to obtain an optimized weight. Then, use the final edited frame generated by the secondary editing to perform weighted optimization on the original Gaussian sputtering model to obtain an edited Gaussian sputtering model.

[0006] Furthermore, the process of selecting the guide frame in step S2 can be selected by the user or automatically, and is determined by a weighted score of the CLIP image-text similarity and the normal viewing angle score. The weighted score formula is as follows: ; in, is the first edited image of view v, Prompt is the editing prompt word, is the rotation matrix of the view v obtained using PoseNet and the target positive view, E is the unit matrix, Indicates Frobenius normalization.

[0007] Furthermore, the detailed steps of the single image 3D generation technology described in step S2 are as follows: Step S2.1: uniformly sample the camera view angles with different diameters around the coordinate center, and use zero123 to obtain images of these camera view angles with the first edited image of the reference view as input.

[0008] Step S2.2: Calculate the feature difference index between viewpoints and generate the difference distribution weight. First, extract the image features of each viewpoint: ; Then calculate the mean feature difference: ; Finally, calculate the difference distribution weight: .

[0009] Step S2.3: Select several viewing angles with the highest possible difference distribution weights as training angles.

[0010] Step S2.4: Use these perspectives and the corresponding zero123 inference images to optimize the Gaussian sputtering model and obtain the cross-modal attention map inverse rendering basis model.

[0011] Furthermore, the de-rendering operation performed in step S3 is shown in the following formula: ; in, represents the cross-modal attention score of the i-th Gaussian ellipsoid, V is the view set, represents the original cross-modal attention map of view v, and p represents A pixel in represents the cross-modal attention score at pixel p, represents the opacity of the Gaussian ellipsoid, Represents the cumulative opacity of all Gaussian ellipsoids between the Gaussian ellipsoid and pixel p. is the weight of pixel p at view angle v based on depth change, which is calculated as: ; in is the depth value of pixel p under view v, and λ is a hyperparameter for adjusting the strength, which is generally set to 20. Introducing the depth change weight helps to trust those geometrically more stable areas (pixels with small depth changes), reduce the influence of occluded boundaries or noisy pixels, and thus improve cross-view consistency and editing quality.

[0012] Furthermore, the optimization weight is calculated in step S4 as follows: ; The loss function is: ; in, is the edited image of the reference view, represents the edited image of the i-th perspective, Represents the original image of the i-th view.

[0013] The present invention also provides a system for implementing the method, comprising: Rendering module: performs multi-view original frame rendering; Base generation module: Integrates zero123 and difference weighted view sampler to build a reverse rendering base model; Cross-modal engine: Implementing attention map de-rendering with depth weights; Optimization module: Calculates weights based on CLIP similarity and optimizes the Gaussian sputtering model.

[0014] Preferably, the difference weight view sampler dynamically screens the training view by CLIP feature difference distribution weights.

[0015] The 3D generation technology used in this paper is used to generate a 3D guidance base. 3D generation uses text or images as a guide to generate a 3D model. This paper uses zero123 (Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot One Image to 3D Object. arXiv, 2303.11328.) to generate the base model. The core technology is based on the image generation capabilities of Stable Diffusion. By introducing a relative viewpoint control mechanism, it can generate new views of the target object from a single image at any viewpoint, further supporting 3D reconstruction. To achieve this, zero123 is fine-tuned using a large-scale synthetic dataset (Objaverse rendered images). The training samples include paired images and their corresponding relative camera extrinsics (rotation and translation). Zero123 learns to map the input image to an image from any new viewpoint. The model employs a modified conditional latent diffusion architecture, integrating two conditional channels: one concatenating the input image's CLIP embedding with the viewpoint transformation parameters as high-level semantic conditional information; the other channel-wise concatenating the input image with the image to be generated to guide detail preservation. To achieve control over the diffusion model's viewpoint, the authors incorporate viewpoint information into the conditional input of the U-Net. This structure allows the model to achieve controllable viewpoint generation without compromising its original image generation capabilities, achieving strong generalization to non-trained categories and real images, and enabling high-quality single-image 3D reconstruction.

[0016] To accommodate non-rigid editing, this paper introduces an inversion-free prior (Sihan Xu, Yidong Huang, Jiayi Pan, Ziqiao Ma, and Joyce Chai. Inversion-Free Image Editing with Natural Language. arXiv, 2312.04965) as a two-dimensional editing guidance prior during the view editing phase. This paper proposes InEdit, an inversion-free image editing framework that leverages a novel diffusion consistency model (DDCM) for efficient and high-quality text-guided image manipulation. Unlike traditional diffusion models that rely on iterative inversions and are prone to reconstruction errors, InEdit employs a non-Markovian forward process starting from random noise to ensure consistency between the original and edited images without the need for an inversion branch. This approach, known as virtual inversion, significantly reduces computational overhead and accumulated error. InEdit integrates a unified attention control (UAC) mechanism that combines cross-attention and mutual self-attention to optimize semantic and structural information, demonstrating excellent editing quality and consistency on the PIE-Bench. By incorporating a latent consistency model (LCM), InEdit achieves faster sampling, completing edits in less than 3 seconds on a single A-40 GPU, surpassing baseline models such as Prompt-to-Prompt, SDEdit, and CycleDiff in both quality and efficiency. The framework also excels in complex image-to-image translation tasks, maintaining compatibility with large language models while addressing ethical concerns such as copyright infringement and potential misuse.

[0017] Beneficial effects: (1) Through camera scoring and selection strategies, this invention reduces the number of cameras used for optimization, reduces the computational complexity of the single-view reconstruction stage, and improves the accuracy and effectiveness of single-view reconstruction. Previous methods used single-diameter uniform sampling, which could result in excessive distribution of cameras in irrelevant areas, resulting in wasted computing power and increased time costs. This invention evaluates the difference between a camera's perspective and all other perspectives, selects the camera group with the greatest mutual difference, and reduces perspective waste.

[0018] (2) The present invention introduces a more expected de-rendering model base, which makes the 3D editing method applicable to non-rigid 3D editing. At the same time, it introduces depth change weights in de-rendering to improve smoothness. The existing multi-view unification mechanism uses the original model as the attention map de-rendering base model, which has good results in rigid editing tasks such as texture and color that hardly modify the geometric features of the original model. However, it has very large limitations in non-rigid editing tasks and is almost unable to modify the geometric features of the target object. The multi-view unification mechanism of the present invention makes 3D editing applicable to non-rigid tasks and also has higher uniformity. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is the overall flow chart of the non-rigid three-dimensional Gaussian editing provided by the present invention, showing the complete process from original frame rendering to the final Gaussian sputtering model optimization. DETAILED DESCRIPTION

[0020] The present invention will be further described below with reference to the accompanying drawings. It should be understood that the embodiments described herein are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0021] Example 1

[0022] The present invention provides a non-rigid three-dimensional editing method based on cross-modal attention guidance, comprising the following steps: Step S1: using the original Gaussian sputtering model and several different viewing angles, rendering the original frame to generate original image data; Before starting to edit, it is necessary to use the original Gaussian model and a series of optimized viewing angles to render a series of original frames for 2D editing. The present invention adopts the original 3D Gaussian rendering method, which can render a frame of image in milliseconds.

[0023] Step S2: Input text prompts and give editing requirements, and use the Inversion-free diffusion model to perform the first round of editing on the original frame based on the prompts. The core advantage of the Inversion-free technology is that it does not require a time-consuming inversion process (i.e., inferring the latent code from the target image), but directly starts from random noise and generates the editing results through a non-Markov forward process, significantly reducing computational complexity and reconstruction error. The editing process is guided by the text prompts provided by the user (such as "adjust the object's posture to standing") to generate multi-view edited images. Then, a frame that best meets the expectations is selected as the guide frame, and a single-image 3D generation method is used to generate an inverse rendering base model; The process of selecting guide frames can be user-selected or automatically selected, and is determined by a weighted score of the CLIP image-text similarity and the orthographic view score. The weighted score formula is as follows: ; in, is the first edited image of view v, Prompt is the editing prompt word, is the rotation matrix of the view v obtained using PoseNet and the target positive view, E is the unit matrix, Indicates Frobenius normalization.

[0024] Then, the edited image of the guide frame is used as input to generate a single view of 3D: The camera viewpoint is uniformly sampled on a sphere with a radius of 1.0 to 2.0, with the coordinate center as the sphere's center. The number of sampling points is typically 100-200, with a pitch angle range of [-30°, 30°] and an azimuth range of [0°, 360°]. Using the guidance frame as input, Zero123 generates an image corresponding to the optimized viewpoint. Zero123 infers the new viewpoint image using a pretrained diffusion model. Inference parameters include the guidance scale (typically 7.5) and the number of sampling steps (typically 50). To improve generation quality, a control network (such as ControlNet) can be introduced to ensure geometric consistency.

[0025] To avoid perspective redundancy, a differential distribution weight strategy is used to dynamically screen perspectives.

[0026] First, extract the image features of each view: ; Calculate the mean feature difference: ; Calculate the difference distribution weight: ; Afterwards, several viewing angles with the highest possible difference distribution weights are selected as training angles.

[0027] Finally, these perspectives and the corresponding zero123 inference images are used to optimize the Gaussian sputtering model and obtain the cross-modal attention map de-rendering basis model.

[0028] Step S3: Perform a secondary edit on the original frame using Inversion-Free, and use the reverse rendering basis model to reversely render and forward render the cross-modal attention maps generated during the secondary editing of the original frame to optimize the consistency of the multi-view 2D editing. The formula for de-rendering is: ; in, represents the cross-modal attention score of the i-th Gaussian ellipsoid, V is the view set, represents the original cross-modal attention map of view v, and p represents A pixel in represents the cross-modal attention score at pixel p, represents the opacity of the Gaussian ellipsoid, Represents the cumulative opacity of all Gaussian ellipsoids between the Gaussian ellipsoid and pixel p. is the weight of pixel p at view angle v based on depth change, which is calculated as: ; in is the depth value of pixel p under view v, and λ is a hyperparameter for adjusting the strength, which is generally set to 20. Introducing the depth change weight helps to trust those geometrically more stable areas (pixels with small depth changes), reduce the influence of occluded boundaries or noisy pixels, and thus improve cross-view consistency and editing quality.

[0029] Step S4: Calculate the similarity between each view image and the reference view image to generate an optimization weight. Use the final edited frame generated by the secondary editing to perform weighted optimization on the original Gaussian sputtering model to obtain an edited Gaussian sputtering model.

[0030] The method for calculating the optimization weight is: ; The loss function is: ; in, is the edited image of the reference view, represents the edited image of the i-th perspective, Represents the original image of the i-th view.

[0031] Example 2

[0032] This embodiment provides a system for implementing the method, including: 1. A rendering module, whose function is to execute step S1 in embodiment 1; The implementation process includes: loading the original Gaussian sputtering model; reading the predefined optimized viewing angle parameters (pitch angle range [-30°, 30°], azimuth angle range [0°, 360°]); calling the renderer to generate multi-view original frames; Its output is: a sequence of raw frames.

[0033] 2. A base generation module, which executes step S2 in Example 1 and integrates the zero123 3D generation engine and the difference weight view sampler. The workflow includes: receiving a guide frame; sampling 200 candidate views on a sphere with a radius of 1.0-2.0; calculating the CLIP feature difference distribution weight; selecting the view with the highest weight as the training view; and optimizing and generating a de-rendered base model.

[0034] 3. A cross-modal engine, which performs step S3 of Example 1 to implement depth-weighted attention map de-rendering and outputs an optimized multi-view cross-modal attention map.

[0035] 4. An optimization module, which is used to execute step S4 of Example 1. Its workflow includes calculating view similarity weights; constructing a weighted loss function; iteratively updating the Gaussian sputtering model using the Adam optimizer; and finally outputting the edited Gaussian sputtering model.

Claims

1. A non-rigid 3D editing method based on cross-modal attention guidance, characterized in that: The following steps are involved: Step S1: Rendering the original frame using the original Gaussian sputtering model and several different viewing angles; Step S2: performing a first round of editing on the original frame using a two-dimensional diffusion model to generate a guide frame, and using the guide frame as input, generating a de-rendered base model using a single-image three-dimensional generation method; Step S3: perform secondary editing on the original frame to generate a two-dimensional text cross-modal attention map, use the de-rendering basis model to de-render the attention map, output the Gaussian point cloud-text prompt word cross-modal attention, and generate an optimized two-dimensional cross-modal attention map through forward rendering; Step S4: Based on the cross-modal attention, calculate the similarity between each perspective image and the reference perspective to obtain an optimized weight; The final edited frame generated by the secondary editing is then used to perform weighted optimization on the original Gaussian sputtering model to obtain the edited Gaussian sputtering model.

2. A non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that: The original Gaussian sputtering model in step S1 is a three-dimensional model reconstructed by using the three-dimensional Gaussian sputtering method on the image data set. Its representation is a collection of three-dimensional Gaussian ellipsoids with spherical harmonic function correction colors. The properties of each three-dimensional Gaussian are as follows: center ,color , opacity , the covariance matrix ,in, is the rotation matrix, is the scaling matrix, attention attribute Used to store attention scores, where n is the number of prompt words and is consistent with the number of channels in the attention map generated during the operation of the diffusion model.

3. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that: In step S2, the three-dimensional generation method used is the zero123 combined with three-dimensional Gaussian sputtering generation method. The specific process is as follows: (1) Fix the corresponding perspective of the input image at a certain perspective as the reference perspective; (2) Uniformly sample the camera angle of view on a sphere with several diameters centered at the center of the sphere, and use zero123 to generate images of these angles. ; (3) Calculate the view difference index: First, use CLIP to extract images from each perspective Features: ; Then the average feature difference is calculated for all image pairs: ; Where N represents the total number of candidate perspectives, ∈[1, N]; (4) Generate differential distribution weights: ; (5) Select several perspectives with the highest difference distribution weights as the optimized perspectives; (6) Use image data with optimized viewing angles to train and generate a three-dimensional Gaussian model.

4. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that: The selection of the guide frame in step S2 is achieved by weighted scoring: ; in, is the first edited image of view v, Prompt is the editing prompt word, is the rotation matrix of the view v obtained using PoseNet and the target positive view, E is the unit matrix, Indicates Frobenius normalization.

5. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that: The de-rendering operation in step S3 is implemented by the following formula: ; in, represents the cross-modal attention score of the i-th Gaussian ellipsoid, V is the view set, represents the original cross-modal attention map of view v, and p represents A pixel in represents the cross-modal attention score at pixel p, represents the opacity of the Gaussian ellipsoid, Represents the cumulative opacity of all Gaussian ellipsoids between the Gaussian ellipsoid and pixel p: ; is the weight of pixel p at view angle v based on depth change, which is calculated as: ; in is the depth value of pixel p under viewing angle v, and λ is a hyperparameter for adjusting the intensity.

6. The non-rigid 3D editing method based on cross-modal attention guidance according to claim 1, characterized in that: The optimization weights in step S4 are calculated by CLIP similarity: ; The loss function is: ; in, is the edited image of the reference view, is the edited image of view i, Represents the original image of the i-th view.

7. A system for implementing the method according to any one of claims 1 to 6, characterized in that: include: Rendering module: performs multi-view original frame rendering; Base generation module: Integrates zero123 and difference weighted view sampler to build a reverse rendering base model; Cross-modal engine: Implementing attention map de-rendering with depth weights; Optimization module: Calculates weights based on CLIP similarity and optimizes the Gaussian sputtering model.

8. The system according to claim 7, characterized in that The difference weighted view sampler dynamically filters the training view by using the CLIP feature difference distribution weights.

Citation Information

Patent Citations

  • Multi-view image reconstruction method and device based on three-dimensional Gaussian sputtering representation

    CN120070799A

  • Three-dimensional model generation method based on new visual angle texture correction

    CN120182533A

  • Interactive three-dimensional Gaussian editing method and system based on three-dimensional geometry consistent attention priori

    CN120411438A

  • Three-dimensional model construction method and system, and related device

    WO2024230843A1

  • System and method for generating three-dimensional model from virtual reality / augmented reality three-dimensional sketch, processing system for three-dimensional model, editing method, and diffusion model training method

    WO2025140611A1