Multi-view consistent 3D object editing method and system based on diffusion model
Through a multi-view consistent 3D object editing method based on a diffusion model, CLIP semantic embedding and attention mechanism are used to dynamically adjust style feature injection. Combined with gradient-guided convolution and spatial attention mechanism, the style drift and structural distortion problems in multi-view 3D editing are solved, and high-fidelity and high-stability 3D object editing is achieved.
Patent Information
- Application Number
- CN202510807754.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have problems in multi-view 3D editing, such as style drift, structural distortion, poor editing consistency, and inaccurate semantic control. Especially when there are high degrees of freedom and large-angle rotations, the 3D consistency and spatial continuity of the generated results are difficult to guarantee.
A multi-view consistent 3D object editing method based on a diffusion model is adopted. Through initial frame style editing, differential style injection and structure-aware fusion, CLIP semantic embedding and attention mechanism are used to dynamically adjust the style feature injection strength. Combined with gradient-guided convolution and spatial attention mechanism, the consistency of style transfer and the stability of geometric structure are ensured.
It effectively mitigates style drift and structural collapse in multi-view editing, ensures the overall geometric consistency and stability of 3D objects, and improves the realism and usability of the generated results.
Smart Images

Figure CN120707790A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of AI generation technology, and in particular relates to a multi-view consistency 3D object editing method and system based on a diffusion model. Background Art
[0002] With the rapid development of virtual reality, augmented reality, and metaverse content creation, the need to maintain the style consistency and structural integrity of 3D objects across multiple viewpoints has become increasingly urgent. While existing approaches based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made significant progress in reconstructing 3D scenes, they still face numerous challenges when performing style transfer or semantic editing.
[0003] First, existing methods based on image or video diffusion models mostly rely on short-term spatiotemporal consistency assumptions, making them difficult to handle 3D object editing tasks with large perspective changes. For static objects or video sequences with large rotations, traditional editing methods typically only maintain good editing results in the first few frames, while subsequent frames exhibit significant style drift, loss of edited features, or structural distortion.
[0004] Secondly, because current diffusion models primarily focus on local texture features during generation and lack effective modeling of the object's global geometric structure, style editing can easily lead to local detail misalignment, object contour distortion, and decreased geometric stability. This is especially true when the style changes dramatically or the editing intensity is high, making it difficult to ensure 3D consistency and spatial continuity in the generated results.
[0005] Furthermore, existing methods generally lack effective structural constraints when handling high-degree-of-freedom, multi-view 3D editing. This leads to uncontrolled diffusion of edited features across different viewpoints, resulting in severe semantic drift, pose distortion, or content ghosting, severely impacting the realism and usability of the resulting output. Therefore, how to effectively suppress drift, maintain cross-view consistency, and preserve the geometric stability of 3D objects while preserving the freedom of stylistic expression has become a critical technical challenge that needs to be addressed in the field of multi-view 3D editing. Summary of the Invention
[0006] In response to the problems of style drift, structural distortion, poor editing consistency and inaccurate semantic control in multi-view 3D object editing in the existing technology, the present invention proposes a multi-view consistent 3D object editing method and system based on a diffusion model.
[0007] To achieve the above object, the present invention provides a multi-view consistent 3D object editing method based on a diffusion model, the method comprising the following steps:
[0008] (1) Initial frame style editing and multi-view propagation: Use the image editing model to edit the input initial frame original image to generate a stylized edited image, and then use the single-view 3D generation model to consistently propagate the editing features to image sequences from multiple perspectives, achieving preliminary multi-view consistent style editing propagation;
[0009] (2) Differential style injection: CLIP semantic embedding is used to calculate the differential features between the original image of the initial frame and the stylized edited image, and the injection strength of the differential features is dynamically adjusted using the attention mechanism to achieve accurate adaptive injection of style features in the multi-view sequence in step (1);
[0010] (3) Structure-aware fusion: Based on step (2), during the multi-view propagation process of step (1), gradient-guided convolution is used to calculate local gradients. The fusion weights of structural information and stylized features are dynamically determined by combining the spatial attention mechanism, and are selectively integrated into the video latent variables to ensure the integrity of the 3D object structure and generate the edited 3D object.
[0011] Furthermore, step (1) includes:
[0012] (1.1) Based on the given style or text hint, the image editing model is used to perform single-view editing on the original image of the initial frame to generate a stylized edited image;
[0013] (1.2) Taking the stylized edited image as input, a single-view 3D generative model is used to propagate the editing features consistently to image sequences from multiple perspectives through an iterative denoising process of a diffusion model, thus achieving preliminary multi-view consistent style editing propagation.
[0014] Furthermore, step (2) includes:
[0015] (2.1) Extract the CLIP semantic embedding of the original image of the initial frame and the stylized edited image, which are recorded as:
[0016] E edit =f CLIP (I edit )
[0017]
[0018] Among them, E edit It indicates that the stylized edited image gets its corresponding feature vector after passing through the CLIP encoder, I edit and I0 represent the stylized edited image and the original image of the initial frame, respectively, f CLIP (·) represents the CLIP encoder function; V s ,Right now Represents the set of all original picture frame sequences, Es ,Right now Indicates that all original picture frame sequences are passed through the CLIP encoder to obtain their corresponding feature vector sets;
[0019] (2.2) Calculate the difference between the CLIP semantic embedding of the stylized edited image and the CLIP semantic embedding of the original image of the initial frame:
[0020] ΔE=E edit -E0
[0021] Among them, E0 represents the feature vector of the initial frame, E edit represents the feature vector of the stylized edited image, and ΔE represents the differential feature between the initial frame and the stylized edited image;
[0022] (2.3) Calculate the semantic similarity between the initial frame features and the original picture frame sequence feature set, dynamically determine the injection strength of the differential style through the attention mechanism, and calculate the semantic similarity matrix between the initial frame style and the original picture frame sequence style set to obtain the attention weight:
[0023]
[0024] Where L is the CLIP embedding dimension, represents its scaling factor, E0 represents the feature vector of the initial frame, represents the transpose of the feature vector of the original image frame sequence, Softmax(·) is the activation function that normalizes the vector into a probability distribution, and α represents the attention weight during style injection;
[0025] (2.4) Inject the differential style features into the diffusion model in the multi-view propagation in step (1.2) as supervision, so that the multi-view propagation model can robustly inject style from the stylized edited image while maintaining structural fidelity, guiding the diffusion model to generate a multi-view consistent and geometrically stable 3D object image sequence with the style of the edited image:
[0026]
[0027] Where i represents the number of each frame, The transpose of the attention weight of the i-th frame, E i represents the sequence features of the original image frame i, and ΔE represents the differential style features; It represents the new style feature with the style of the stylized edited image after injecting the differential style into the sequence feature of the i-th frame of the original image according to the attention weight.
[0028] Furthermore, step (3) includes:
[0029] (3.1) In the forward diffusion stage of the diffusion model, the latent variables are initialized for each frame:
[0030]
[0031] in represents the initial noise latent variable of the i-th frame, Indicates that the noise latent variable conforms to the Gaussian distribution;
[0032] (3.2) The initialized latent variables are iteratively updated through the denoising network, and the latent variables obtained by iteration at each time step are recorded. The update formula is as follows:
[0033] Z t =∈ θ (Z t-1 ,t,E0),
[0034] Among them, ∈ θ is the noise predictor of the multi-view propagation diffusion model, t represents the current time step, E0 represents the initial frame feature vector, which is used to provide structural guidance;
[0035] (3.3) During the iterative denoising process of the diffusion model in step (1.2), the latent variables recorded in step (3.2) are adaptively fused with the current diffusion latent variables.
[0036] Furthermore, step (3.3) includes:
[0037] (3.3.1) A gradient-based method is used in the convolutional layer to calculate the spatial adaptive weight map λ(x,y), so that it prioritizes the original geometric information in areas with obvious structural features, and the Sobel operator is applied to calculate the gradient G of the original structural potential features in the horizontal and vertical directions respectively. x (x,y) and G y (x,y), the specific calculation formula of the adaptation weight map is as follows:
[0038]
[0039] The formula for local fusion based on the adaptive weight map is as follows:
[0040]
[0041] Among them, max (x,y) G(x,y) represents the maximum value of the gradient amplitude at all positions (i,j) in the entire feature map, which is used for global normalization. ∈ represents a very small positive constant; Z t (x,y) and denote the latent variables recorded in step (3.2) and the latent variables in the iterative denoising process of the diffusion model in step (1.2), respectively;
[0042] (3.3.2) In the spatial attention layer that captures the global structure interaction, only the query Q t and key K t The components apply a fixed mixing factor β∈[0,1] to selectively inject structural information while maintaining the value V t The specific weighted fusion formula is as follows:
[0043]
[0044] in, Respectively represent the query, key, and value components of the new hybrid latent variable after fusion; Respectively represent the query, key, and value components of the latent variables recorded in step (3.2).
[0045] To achieve the above objectives, the present invention also provides a multi-view consistency 3D scene editing system based on a diffusion model, which includes the following modules:
[0046] Initial frame style editing module: This module receives a single-view image as input, combines it with specified style cues, uses the image editing model to generate a stylized initial frame, and then uses the single-view 3D generative model to consistently propagate this style feature to the multi-view sequence, forming a preliminary stylized 3D scene.
[0047] Multi-view differential style injection and multi-view propagation modules: Based on the CLIP encoder, they extract semantic differential features between the original frame and the stylized edited image, calculate attention correlations across views, and dynamically enhance the style of each frame based on the differential features and attention scores to achieve precisely controlled multi-view style injection.
[0048] Structure-aware adaptive fusion module: During the diffusion generation process, it records the structural information of the latent variables of each frame, adopts a local gradient-based convolutional fusion strategy and a spatial attention mechanism, and dynamically adjusts the fusion ratio of structural information and style features to ensure that the consistency and geometric structure stability of 3D objects are maintained while performing style editing.
[0049] To achieve the above-mentioned purpose, the present invention also provides an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned multi-perspective consistency 3D object editing method based on the diffusion model.
[0050] To achieve the above object, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned multi-view consistency 3D object editing method based on the diffusion model.
[0051] The beneficial effects of the present invention are as follows:
[0052] The present invention can effectively alleviate the problems of style drift, detail distortion and structural collapse in multi-perspective editing by introducing an overall process based on initial frame style propagation, differential semantic injection and structural perception fusion. By introducing the attention mechanism guided by CLIP differential semantics, precise control of style changes under various perspectives is achieved. At the same time, the consistency and stability of the overall geometric structure of 3D objects are ensured through the convolution fusion of local gradient perception and spatial attention regulation. The overall method does not require retraining of the diffusion model, and editing can be completed only in the inference stage. It is easy to deploy and has strong adaptability. It can be widely used in virtual reality, augmented reality, digital content generation and other fields, and has extremely high application value and promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings are incorporated into the specification and serve as an integral part of the specification, illustrating the implementation of the various modules involved in the present invention in accordance with the principles of the present invention. The focus of the drawings is not to limit the invention, but to explain the principles of the invention. In the drawings,
[0054] Figure 1 is an execution flow chart of each module of the 3D object editing method and system of the present invention;
[0055] Figure 2 This is a flowchart of calculating the difference in the multi-view difference style injection module of the present invention;
[0056] Figure 3 This is a flowchart of differential style injection in the multi-view differential style injection module of the present invention;
[0057] Figure 4 is a detailed flow chart of the structure-aware adaptive fusion module of the present invention;
[0058] Figure 5 is the numerical details of the Sobel operator used in this invention;
[0059] Figure 6 is the image of the edited 3D object obtained in the present invention;
[0060] Figure 7 It is a schematic diagram of an electronic device of the present invention. DETAILED DESCRIPTION
[0061] The following detailed description, which refers to the accompanying drawings, will set forth specific details of one embodiment in order to provide a comprehensive understanding of all aspects of the claimed invention. It will be apparent to those skilled in the art that certain modules and systems of the present invention may be implemented using alternative embodiments that differ from the specific details described below. The following description is intended to be illustrative rather than limiting. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention are intended to be included within the scope of protection of the invention.
[0062] See Figure 1 , Figure 2 , Figure 3 and Figure 4 The present invention provides a multi-view consistency 3D object editing method based on a diffusion model, which specifically includes the following steps:
[0063] Step 1: Initial Frame Style Editing and Multi-View Propagation. A single initial frame raw image is input. Based on a specified style or textual cue, the image editing model is used to stylize the image, generating a stylized edited image. This stylized edited image is then used as input to a single-view 3D generative model. Through an iterative denoising process using a diffusion model, the edited features are consistently propagated to a multi-view image sequence, achieving preliminary multi-view consistent style editing propagation and generating a multi-view consistent and geometrically stable 3D object image sequence. Finally, 3D reconstruction is performed based on the 3D object image sequence to obtain the edited 3D object.
[0064] In step 1, the process of generating a single-view 3D model includes:
[0065] (1.1) Initial frame style editing: Based on the given style or text hint, the image editing model is used to perform single-view editing on the input initial frame original image I0 to generate a stylized edited image I edit .
[0066] (1.2) Multi-perspective communication: I edit The model takes the image as input and undergoes iterative denoising through a diffusion model. Two modules, step 2, multi-view differential style injection, and step 3, structure-aware adaptive fusion, are introduced into the diffusion model. These two modules generate a multi-view consistent edited 3D object image sequence, ensuring that the geometric structure remains consistent across all viewpoints.
[0067] (1.3) Generating an edited 3D object: performing three-dimensional reconstruction based on the edited 3D object image sequence to obtain the edited 3D object.
[0068] Step 2: Differential Style Injection. To further improve the consistency and local controllability of style transfer, we extract the CLIP semantic embeddings of the original image and the stylized edited image, calculate the differential features between the two, and dynamically adjust the strength of the style (differential feature) injection based on the attention mechanism, achieving accurate and adaptive injection of style features in multi-view sequences.
[0069] In step 2, the differential style injection process specifically includes:
[0070] (2.1) Differential style calculation, such as Figure 2 As shown in Figure 2, the stylized edited image and the original picture frame sequence (i.e., the initial frame original image) are passed through the CLIP encoder to obtain their corresponding feature vectors:
[0071] E edit =f CLIP (I edit )
[0072]
[0073] Among them I edit represents the stylized edited image, f CLIP (·) represents the CLIP encoder function, E edit V represents the corresponding feature vector obtained after the stylized edited image passes through the CLIP encoder. s ,Right now Represents the set of all original picture frame sequences; E s ,Right now This represents the set of feature vectors corresponding to all original image frame sequences after they are passed through the CLIP encoder. By encoding all original image sequences, we can obtain the global semantic context from multiple perspectives, providing richer information support for subsequent attention matching.
[0074] (2.2) Then calculate the differential features between the stylized edited image and the first frame in the original picture frame sequence, that is, the initial frame:
[0075] ΔE=E edit -E0
[0076] Among them, E0 represents the feature vector of the initial frame, E edit The feature vector of the stylized image is represented by ΔE, which represents the difference between the original frame and the stylized image. This method highlights the style changes, suppresses the interference of the original content, and enhances the effectiveness and pertinence of the style signal.
[0077] (2.3) Style injection stage, such as Figure 3As shown, the semantic similarity between the initial frame features and the original picture frame sequence feature set is calculated, the injection strength of the differential style is dynamically determined through the attention mechanism, and the semantic similarity matrix between the initial frame style and the original picture frame sequence style set is calculated to obtain the attention weight:
[0078]
[0079] Where L is the CLIP embedding dimension, represents its scaling factor, E0 represents the feature vector of the initial frame, represents the transpose of the feature vector of the original image frame sequence. Softmax(·) is an activation function that normalizes the vector to a probability distribution. α represents the attention weight during style injection. The amount of style injection is adaptively adjusted based on the semantic similarity of each frame. Frames with high similarity receive stronger style enhancement, while frames with low similarity retain their structural integrity.
[0080] (2.4) Next, the differential features obtained in the differential style calculation are replicated n times on the 0th dimension. The injection strength of the differential style features is then controlled using the attention weights. The differential style features are then injected into the diffusion model in the multi-view propagation in step (1.2) as supervision, enabling the multi-view propagation model to robustly inject style from the stylized edited image while maintaining structural fidelity, guiding the diffusion model to generate a 3D object image sequence with a multi-view consistent and geometrically stable style of the edited image:
[0081]
[0082] Where i represents the number of each frame, i=0,…,n, The transpose of the attention weight of the i-th frame, E i represents the sequence features of the i-th frame of the original image, and ΔE represents the differential style features. This represents a new style feature that represents the stylized style of the edited image after injecting the differential style based on the attention weights into the features of the i-th frame sequence of the original image. By combining local differences with global attention and introducing a clear "structure-style" balance mechanism, this approach significantly improves the multi-view consistency and structural fidelity of the edited image.
[0083] Step 3: Structure-aware fusion. Based on step 2, during the iterative denoising process of the diffusion model in multi-view propagation (1.2), a structure-aware adaptive fusion strategy based on local gradient guidance and spatial attention mechanism is introduced to dynamically fuse the original structural latent variable with the style latent variable. That is, gradient-guided convolution is used to calculate local gradients, and the spatial attention mechanism is used to dynamically determine the fusion weights of structural information and stylized features, which are then selectively fused into the video latent variable. This ensures that the overall geometric structure and cross-view consistency of the 3D object are preserved and enhanced while the style changes.
[0084] In step 3, if Figure 4 As shown in Figure 2, the structure-aware fusion process specifically includes:
[0085] (3.1) The system uses the initial frame original image I0 to perform diffusion denoising through the multi-view diffusion model. In the forward process, the latent variables of each frame are initialized:
[0086]
[0087] in, is the initial noise latent variable of the i-th frame, It means that the noise latent variable conforms to the standard Gaussian distribution, and the subscript 0 indicates initialization, that is, the latent variable at time step 0.
[0088] (3.2) The initialized latent variable is iteratively updated through the diffusion denoising network. The system records the noise latent variable obtained after passing through the denoising network at each time step t:
[0089] Z t =∈ θ (Z t-1 ,t,E0),
[0090] Among them, ∈ θ is the noise predictor of the multi-view propagation diffusion model, t is the current time step, E0 is the initial frame feature vector, which is used to provide structural guidance. t=1,2,…,n step , which means that a total of step denoising iterations have been performed. Finally, the noise latent variable set of all time steps is obtained.
[0091] (3.3) Using the stylized edited image, at each time step t of the forward inference process of the diffusion model of multi-view propagation in step (1.2), adaptively convert The feature information in is injected into the denoising reasoning process. Take any time step t as an example, Figure 4As shown in Figure 1, the multi-view propagation diffusion model (i.e., the multi-view propagation network) has two modules, the convolution layer and the spatial attention layer, in both the upsampling and downsampling processes. In the convolution layer and the spatial attention layer modules of the downsampling part, the latent variables recorded in step (3.2) are adaptively fused with the latent variables in the multi-view propagation diffusion model in step (1.2), including the following sub-steps:
[0092] (3.3.1) Gradient Adaptive Weight Injection: In the upsampled convolutional layer, a spatially adaptive weight map is calculated based on the local gradient, so that the original geometric information is prioritized in areas with obvious structural features (high gradients). The Sobel operator is applied to calculate the gradients of the original structural potential features along the horizontal and vertical directions respectively. The specific calculation formula of the adaptive weight map is as follows:
[0093]
[0094] Among them, G x (x,y) and G y (x, y) represents the application of the Sobel operator to calculate the gradient of the original structure potential features along the horizontal and vertical directions, G(x, y) represents the gradient amplitude synthesized by the horizontal and vertical gradients at the position (x, y), which is used to measure the strength of the structural information at that location. The convolution kernel details of the Sobel operator are as follows: Figure 5 As shown in the figure, through this adaptive weight based on gradient magnitude, the network can better preserve geometric features in areas with rich structures, while appropriately enhancing semantic expression capabilities in flat areas. λ(x,y) is the adaptive weight value for each position (x,y). max (x,y) G(x,y) represents the maximum gradient magnitude at all positions (i,j) in the entire feature map and is used for global normalization. ∈ represents a very small positive constant used to prevent the denominator from being zero and to ensure numerical stability.
[0095] Local fusion is performed based on the adaptive weight map to take into account both the original geometric information and the semantic expression of the diffusion model:
[0096]
[0097] Among them, Z t (x,y) represents the latent variable value recorded in step (3.2) at the tth diffusion time step, represents the original latent variable value obtained by iterative denoising of the multi-view propagation diffusion model in step (1.2) at the tth diffusion time step. λ(x,y) is the adaptive weight value for each position (x,y) in the adaptive weight stock. Represents the new mixed latent variable obtained after weighted fusion. The new latent variable will replace the original latent variable in step (1.2) and continue the denoising process of the next time step.
[0098] (3.3.2) Query and Key controllable injection, in the upsampling spatial attention module (i.e., in the spatial attention layer that captures global structural interactions), for any time step t, only the query (Q t ) and key (K t ) components are fused using a fixed blending factor, specifically:
[0099]
[0100] Among them, β is the preset mixing factor, β∈[0,1]; the value (V t ) remains unchanged to balance structural preservation and style editing. t ,K t They represent the query and key components of the original latent variable values obtained by iterative denoising of the multi-view propagation diffusion model in step (1.2), Respectively represent the query, key and value components of the latent variables recorded in step (3.2), Represent the query, key, and value components of the new hybrid latent variable after fusion. The new latent variable will replace the original latent variable in step (1.2) and continue the denoising process for the next time step.
[0101] Through the collaborative work of the above three modules, this embodiment can achieve cross-view consistent style editing of any object while maintaining 3D structural consistency, and generate the edited 3D object (such as Figure 6 It has the advantages of high fidelity and high stability.
[0102] In addition, if Figure 1 As shown, the present invention also provides a multi-view consistency 3D scene editing system based on a diffusion model, which includes the following modules:
[0103] Initial frame style editing module: It is used to receive the input single-view image, combine it with the specified style hints, use the image editing model to generate a stylized initial frame, and then use the single-view 3D generation model to consistently propagate the style features to the multi-view sequence to form a preliminary stylized 3D scene.
[0104] Multi-view differential style injection and multi-view propagation modules: Based on the CLIP encoder, the semantic differential features of the initial frame original image and the stylized edited image are extracted, the attention correlation across views is calculated, and the style of each frame is dynamically enhanced based on the differential features and attention scores to achieve precisely controlled multi-view style injection.
[0105] Structure-aware adaptive fusion module: During the diffusion generation process, it records the structural information of the latent variables of each frame, adopts a local gradient-based convolutional fusion strategy and a spatial attention mechanism, and dynamically adjusts the fusion ratio of structural information and style features to ensure that the consistency and geometric structure stability of 3D objects are maintained while performing style editing.
[0106] Corresponding to the embodiment of the multi-view consistency 3D object editing method based on the diffusion model, the embodiment of the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-view consistency 3D object editing method based on the diffusion model as described above. Figure 7 As shown in the figure, a hardware structure diagram of any device with data processing capability for the multi-view consistency 3D object editing method based on the diffusion model provided in the embodiment of the present application is provided. Figure 7 In addition to the processor, memory, DMA controller, disk, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0107] Corresponding to the embodiment of the multi-perspective consistent 3D object editing method based on the diffusion model mentioned above, an embodiment of the present invention also provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the multi-perspective consistent 3D object editing method based on the diffusion model in the above embodiment is implemented.
[0108] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0109] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solutions disclosed in the present invention without the need for creative work should be included in the scope of protection of the present invention.
Claims
1. A multi-view consistent 3D object editing method based on a diffusion model, characterized by: The method comprises the following steps: (1) Initial frame style editing and multi-view propagation: Use the image editing model to edit the input initial frame original image to generate a stylized edited image, and then use the single-view 3D generation model to consistently propagate the editing features to image sequences from multiple perspectives, achieving preliminary multi-view consistent style editing propagation; (2) Differential style injection: CLIP semantic embedding is used to calculate the differential features between the original image of the initial frame and the stylized edited image, and the injection strength of the differential features is dynamically adjusted using the attention mechanism to achieve accurate adaptive injection of style features in the multi-view sequence in step (1); (3) Structure-aware fusion: Based on step (2), during the multi-view propagation process of step (1), gradient-guided convolution is used to calculate local gradients. The fusion weights of structural information and stylized features are dynamically determined by combining the spatial attention mechanism, and are selectively integrated into the video latent variables to ensure the integrity of the 3D object structure and generate the edited 3D object.
2. The multi-view consistent 3D object editing method based on the diffusion model according to claim 1, characterized in that: Step (1) includes: (1.1) Based on the given style or text hint, the image editing model is used to perform single-view editing on the original image of the initial frame to generate a stylized edited image; (1.2) Taking the stylized edited image as input, a single-view 3D generative model is used to propagate the editing features consistently to image sequences from multiple perspectives through an iterative denoising process of a diffusion model, thus achieving preliminary multi-view consistent style editing propagation.
3. The multi-view consistent 3D object editing method based on the diffusion model according to claim 2, characterized in that: Step (2) includes: (2.1) Extract the CLIP semantic embedding of the original image of the initial frame and the stylized edited image, which are recorded as: HAVE BEEN edit =f CLIP (I edit ) Among them, E edit It indicates that the stylized edited image gets its corresponding feature vector after passing through the CLIP encoder, I edit and I0 represent the stylized edited image and the original image of the initial frame, respectively, f CLIP (·) represents the CLIP encoder function; V s ,Right now Represents the set of all original picture frame sequences, E s ,Right now Indicates that all original picture frame sequences are passed through the CLIP encoder to obtain their corresponding feature vector sets; (2.2) Calculate the difference between the CLIP semantic embedding of the stylized edited image and the CLIP semantic embedding of the original image of the initial frame: ΔE=E edit -E0 Among them, E0 represents the feature vector of the initial frame, E edit represents the feature vector of the stylized edited image, and ΔE represents the differential feature between the initial frame and the stylized edited image; (2.3) Calculate the semantic similarity between the initial frame features and the original picture frame sequence feature set, dynamically determine the injection strength of the differential style through the attention mechanism, and calculate the semantic similarity matrix between the initial frame style and the original picture frame sequence style set to obtain the attention weight: Where L is the CLIP embedding dimension, represents its scaling factor, E0 represents the feature vector of the initial frame, represents the transpose of the feature vector of the original image frame sequence, Softmax(·) is the activation function that normalizes the vector into a probability distribution, and α represents the attention weight during style injection; (2.4) Inject the differential style features into the diffusion model in the multi-view propagation in step (1.2) as supervision, so that the multi-view propagation model can robustly inject style from the stylized edited image while maintaining structural fidelity, guiding the diffusion model to generate a multi-view consistent and geometrically stable 3D object image sequence with the style of the edited image: Where i represents the number of each frame, The transpose of the attention weight of the i-th frame, E i represents the sequence features of the original image frame i, and ΔE represents the differential style features; It represents the new style feature with the style of the stylized edited image after injecting the differential style into the sequence feature of the i-th frame of the original image according to the attention weight.
4. The multi-view consistent 3D object editing method based on the diffusion model according to claim 2, characterized in that: Step (3) includes: (3.1) In the forward diffusion stage of the diffusion model, the latent variables are initialized for each frame: in represents the initial noise latent variable of the i-th frame, Indicates that the noise latent variable conforms to the Gaussian distribution; (3.2) The initialized latent variables are iteratively updated through the denoising network, and the latent variables obtained by iteration at each time step are recorded. The update formula is as follows: WITH t =∈ θ (WITH t-1 ,t,E0), Among them, ∈ θ is the noise predictor of the multi-view propagation diffusion model, t represents the current time step, E0 represents the initial frame feature vector, which is used to provide structural guidance; (3.3) During the iterative denoising process of the diffusion model in step (1.2), the latent variables recorded in step (3.2) are adaptively fused with the current diffusion latent variables.
5. The multi-view consistent 3D object editing method based on the diffusion model according to claim 4, characterized in that: Step (3.3) includes: (3.3.1) A gradient-based method is used in the convolutional layer to calculate the spatial adaptive weight map λ(x,y), so that it prioritizes the original geometric information in areas with obvious structural features, and the Sobel operator is applied to calculate the gradient G of the original structural potential features in the horizontal and vertical directions respectively. x (x,y) and G y (x,y), the specific calculation formula of the adaptation weight map is as follows: The formula for local fusion based on the adaptive weight map is as follows: Among them, max (x,y) G(x,y) represents the maximum value of the gradient amplitude at all positions (i,j) in the entire feature map, which is used for global normalization. ∈ represents a very small positive constant; Z t (x,y) and denote the latent variables recorded in step (3.2) and the latent variables in the iterative denoising process of the diffusion model in step (1.2), respectively; (3.3.2) In the spatial attention layer that captures the global structure interaction, only the query Q t and key K t The components apply a fixed mixing factor β∈[0,1] to selectively inject structural information while maintaining the value V t The specific weighted fusion formula is as follows: in, Respectively represent the query, key, and value components of the new hybrid latent variable after fusion; Respectively represent the query, key, and value components of the latent variables recorded in step (3.2).
6. A multi-view consistency 3D scene editing system based on a diffusion model, characterized by: The system includes the following modules: Initial frame style editing module: This module receives a single-view image as input, combines it with specified style cues, uses the image editing model to generate a stylized initial frame, and then uses the single-view 3D generative model to consistently propagate this style feature to the multi-view sequence, forming a preliminary stylized 3D scene. Multi-view differential style injection and multi-view propagation modules: Based on the CLIP encoder, they extract semantic differential features between the original frame and the stylized edited image, calculate attention correlations across views, and dynamically enhance the style of each frame based on the differential features and attention scores to achieve precisely controlled multi-view style injection. Structure-aware adaptive fusion module: During the diffusion generation process, it records the structural information of the latent variables of each frame, adopts a local gradient-based convolutional fusion strategy and a spatial attention mechanism, and dynamically adjusts the fusion ratio of structural information and style features to ensure that the consistency and geometric structure stability of 3D objects are maintained while performing style editing.
7. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the multi-view consistency 3D object editing method based on the diffusion model as described in any one of claims 1-5 above.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multi-view consistency 3D object editing method based on the diffusion model is implemented.
Citation Information
Cited By
Intelligent lightweight collaborative design method and system for digital media content
CN121788079A