Multi-view texture remodeling method and system based on structure perception
Through a multi-view texture reconstruction method based on structure perception, using a pre-trained diffusion model, a noise sharing module, and a dual-stream residual guidance module, the problem of structural consistency in three-dimensional scene texture reconstruction is solved, efficient and unified texture reconstruction is achieved, and the flexibility and quality of three-dimensional content creation are improved.
Patent Information
- Application Number
- CN202511146073.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies have difficulty maintaining structural consistency in three-dimensional scene texture reconstruction, resulting in distortion of the reconstructed scene geometric structure, discontinuous texture, and inconsistent performance under different perspectives. They lack structural perception capabilities and are unable to meet the requirements of realism and immersion.
A multi-view texture reconstruction method based on structure perception is adopted. Multi-view rendered images are processed in batches through a pre-trained diffusion model. A noise sharing module and a two-stream residual guidance module are introduced. Structural information is used as an anchor point to achieve spatial coordination and visual consistency in texture editing.
It improves the flexibility and accuracy of 3D scene editing, realizes texture consistency editing across perspectives and batches, and significantly improves the quality and efficiency of 3D content generation.
Smart Images

Figure CN120635282A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision processing, and in particular relates to a multi-view texture reconstruction method and system based on structure perception. Background Art
[0002] In fields such as digital content creation, virtual reality, film and television post-production, and interactive 3D applications, the demand for high-quality modeling and flexible editing of 3D scenes is growing. In actual content creation, users often prefer to use existing 3D scenes as templates for personalized modifications, thereby saving modeling costs and improving content generation efficiency. However, traditional methods for texture modification often lack a thorough understanding of 3D structure and texture consistency, resulting in significant deficiencies in the resulting spatial structure and texture consistency.
[0003] In the actual texture reconstruction process, it is usually necessary to render multiple images from different perspectives from the original 3D scene, then perform batch texture reconstruction operations on these images, and finally send the edited results back to the 3D scene for updating and optimization. However, this process faces three key technical challenges: (1) It is difficult to maintain the consistency of the original structure during the texture reconstruction process, which easily leads to distortion or local deformation of the geometric structure of the reconstructed scene, affecting the overall visibility and authenticity; (2) The texture editing results between images from different perspectives in the same batch are uneven, and obvious discontinuities or artifacts are likely to appear in the 3D fusion stage, destroying the unified visual experience under multiple perspectives; (3) There are differences in style or details between texture reconstructions in different batches, resulting in inconsistent performance of the final 3D scene constructed from different perspectives, which is difficult to meet the requirements of realism and immersion. Although some existing methods have tried to introduce consistency constraints or local guidance mechanisms to alleviate the above-mentioned multi-perspective problems, they still show limitations such as insufficient consistency control ability and insufficient structural understanding in complex texture reconstruction tasks. Innovative methods with more structure-aware capabilities are urgently needed to make breakthroughs. Summary of the Invention
[0004] This invention aims to address the existing challenges and provide a structure-aware, multi-view texture reconstruction method and system. This approach leverages the structural information in existing 3D scenes to link texture information across different viewpoints, thereby enhancing the spatial coordination and visual consistency of texture editing. This method not only improves the flexibility and precision of 3D scene editing but also provides a reliable technical foundation for rapid 3D content generation for large-scale, real-world scenes.
[0005] The inventive concept of the present invention is: using multi-perspective rendered images of the three-dimensional original scene as input data, passing the multi-perspective rendered images into the diffusion model and control network in batches, extracting the structural information of each perspective, thereby guiding a batch of latent variables to denoise and reshape the texture. For the rendered images of the same batch, a noise sharing module is introduced to perform controllable weighted mixing of the predicted noise in the batch, so that each latent variable in the batch can obtain the texture information of other latent variables. For the rendered images of different batches, a dual-stream residual guidance module is introduced. By constructing a dual-branch architecture of the reconstruction path and the editing path, the structural information is used as an anchor to guide the optimization process of the edited image, thereby alleviating the problem of information isolation between different batches and providing a stable geometric semantic basis for subsequent unified texture reshaping. After all batches of images have been edited, they can be used to optimize the three-dimensional original scene to reshape its texture. The present invention can efficiently complete customized texture editing of complex three-dimensional scenes without relying on additional training, significantly improving the flexibility and quality of three-dimensional content creation.
[0006] In order to achieve the above-mentioned object of the invention, the present invention specifically adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a multi-view texture reconstruction method based on structure perception, which comprises the following steps:
[0008] S1: using a pre-trained diffusion model combined with a control network to batch process multi-view rendered images of a 3D original scene, encoding the multi-view rendered images using an image encoder in the diffusion model, and encoding the text description of the original scene and the text description of the target scene respectively through a text encoder in the diffusion model;
[0009] S2: extracting the structural information of the multi-view rendered image at each viewpoint, using the extracted structural information as a structural guidance condition, and passing it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network;
[0010] S3: Construct a dual-branch architecture for the reconstruction path and the editing path. A batch of noise is sampled under a standard Gaussian distribution as the initial noise latent variables for the editing and reconstruction paths. The UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. Both paths perform multi-step iterative denoising. During the iterative denoising process, the noise sharing module performs a controllable weighted mixing of the predicted noise corresponding to the same batch of multi-view rendered images. The two-stream residual guidance module corrects the predicted noise generated by the editing path for different batches of multi-view rendered images.
[0011] S4: After all batches of multi-view rendered images have been edited, the resulting edited images are used to optimize the original three-dimensional scene, achieving multi-view consistent texture reshaping.
[0012] Based on the above solution, each step can be implemented in the following preferred specific manner.
[0013] As a preferred embodiment of the first aspect, in step S1, the control network includes a copy of the image encoder in the diffusion model and a zero convolution layer.
[0014] As a preferred embodiment of the first aspect, the specific steps in step S1 include:
[0015] S11: sequentially passing a batch of multi-view rendered images to the image encoder of the diffusion model for processing, and obtaining image codes corresponding to the rendered images of each view;
[0016] S12: Encode the text description of the original scene and the text description of the target scene using the text encoder in the diffusion model respectively, to obtain the text description encoding of the original scene and the text description encoding of the target scene respectively.
[0017] As a preferred embodiment of the first aspect above, in step S2, the specific steps of extracting structural information are: sequentially passing a batch of multi-perspective rendered images into a structural feature extractor selected according to the three-dimensional scene for processing, and obtaining the structural information corresponding to the rendered images of each perspective.
[0018] As a preferred embodiment of the first aspect, in step S3, the specific processing process in the noise sharing module includes:
[0019] AS31: Under the standard Gaussian distribution, a batch of noise is sampled as the initial noise latent variable of the current batch, and the same noise is copied to reconstruct the initial noise latent variable of the path;
[0020] AS32: Performs T-step iterative denoising on the reconstruction path using the initial noise latent variables of the reconstruction path, and performs T-step iterative denoising on the editing path using the sampled initial noise latent variables.
[0021] AS33: For the reconstruction path, the noise latent variable at the t-th time step is used as the reconstruction noise latent variable, and the current time step t and the reconstruction noise latent variable, the structural information of the current batch, and the text description of the original scene are encoded and passed into the UNet network to predict the noise; for the editing path, the noise latent variable at the t-th time step is used as the editing noise latent variable, and the current time step t and the editing noise latent variable, the structural information of the current batch, and the text description of the target scene are encoded and passed into the UNet network to predict the noise;
[0022] AS34: A shared time step threshold range is preset. If the current time step t is not within the shared time step threshold range, noise sharing is not performed, and the predicted noise corresponding to each perspective rendered image in the current batch is directly used. If the current time step t is within the shared time step threshold range, the predicted noise corresponding to each perspective rendered image in the current batch is traversed for weighted sharing.
[0023] As a preferred embodiment of the first aspect, in step AS34, the predicted noise of the rendered image of the vth view at the tth time step in the reconstruction path is used as the first predicted noise, the average of the predicted noises of the rendered images of other views in the current batch is used as the first auxiliary information, and the first auxiliary information and the first predicted noise are weighted to form the reconstructed shared noise of the rendered image.
[0024] The predicted noise of the v-th perspective rendered image at the t-th time step of the editing path is used as the second predicted noise, and the average of the predicted noises of the rendered images of other perspectives in the current batch is used as the second auxiliary information. The second auxiliary information and the second predicted noise are weighted to form the editing shared noise of the rendered image.
[0025] Furthermore, in step AS34, Time step Reconstruction of shared noise from rendered images from different perspectives Share noise with editors The specific calculation method is:
[0026]
[0027]
[0028] in, is the batch size, is the shared noise intensity; Reconstruction path Time step Prediction noise of rendered images from different perspectives; Indicates the editing path Time step Prediction noise of rendered images from different perspectives; Reconstruction path Time step Prediction noise of rendered images from different perspectives; Indicates the editing path Time step Predicted noise for rendered images from different perspectives.
[0029] As a preferred embodiment of the first aspect, in step S3, the specific processing process in the dual-stream residual guidance module includes:
[0030] BS31: For the t-th time step of the reconstruction path, multiply the image code of the rendered image by a preset first scaling factor to obtain a first calculation result, subtract the first calculation result from the reconstruction noise latent variable at that time step to obtain a second calculation result, and multiply the reciprocal of the preset second scaling factor by the second calculation result to obtain the ideal noise at that time step; wherein the first scaling factor is the arithmetic square root of the guided coefficient at the t-th time step of the reconstruction path, and the second scaling factor is the arithmetic square root of the difference between 1 and the guided coefficient at the t-th time step of the reconstruction path;
[0031] BS32: Subtract the ideal noise at the t-th time step from the reconstructed shared noise of the rendered image as the noise residual at that time step;
[0032] BS33: Use the ideal noise at the t-th time step to denoise the reconstruction noise latent variable at that time step in the reconstruction path to generate the reconstruction noise latent variable at the t-1-th time step;
[0033] BS34: For the t-th time step of the editing path, the noise residual is used as the anchor guide and multiplied by the guide strength factor to form a noise correction. The noise correction is then added to the edit shared noise of the rendered image at that time step to obtain the corrected noise of the editing path at that time step.
[0034] BS35: Use the corrected noise at the t-th time step to denoise the editing noise latent variable at that time step in the editing path to generate the editing noise latent variable at the t-1-th time step.
[0035] Further, in step BS31, the path is rebuilt ideal noise at time steps The calculation method is:
[0036]
[0037] in, represents the first proportionality coefficient; is the second proportional coefficient; It is The bootstrap coefficient for each time step; An image encoding representing a rendered image; Indicates the The reconstruction noise latent variable of time steps.
[0038] As a preferred embodiment of the above-mentioned first aspect, the specific processing process of the denoising operation in step BS33 is: multiplying the ideal noise at the t-th time step by the second proportional coefficient to obtain a third calculation result, subtracting the reconstruction noise latent variable at the t-th time step from the third calculation result to obtain a fourth calculation result, dividing the fourth calculation result by the first proportional coefficient to obtain a fifth calculation result, multiplying the fifth calculation result by the preset third proportional coefficient to obtain a sixth calculation result, multiplying the ideal noise at the t-th time step by the preset fourth proportional coefficient to obtain a seventh calculation result, and adding the sixth calculation result and the seventh calculation result as the reconstruction noise latent variable at the t-1-th time step; wherein, the third proportional coefficient is the arithmetic square root of the guide coefficient of the reconstruction path at the t-1-th time step, and the fourth proportional coefficient is the arithmetic square root of the difference between 1 and the guide coefficient of the reconstruction path at the t-1-th time step.
[0039] Further, in step BS33, The reconstruction noise latent variable of time steps The specific calculation method is:
[0040]
[0041] in, represents the third proportional coefficient; is the fourth proportional coefficient; It is The guidance coefficient for each time step.
[0042] As a preferred embodiment of the above-mentioned first aspect, the specific processing process of the denoising operation in step BS35 is: multiplying the corrected noise at the t-th time step by the second proportional coefficient to obtain the eighth calculation result, subtracting the editing noise latent variable at the t-th time step from the eighth calculation result to obtain the ninth calculation result, dividing the ninth calculation result by the first proportional coefficient to obtain the tenth calculation result, multiplying the tenth calculation result by the third proportional coefficient to obtain the eleventh calculation result, multiplying the corrected noise at the t-th time step by the fourth proportional coefficient to obtain the twelfth calculation result, and adding the eleventh calculation result and the twelfth calculation result as the editing noise latent variable at the t-1-th time step.
[0043] Further, in step BS35, The editing noise latent variable of time steps The specific calculation method is:
[0044]
[0045]
[0046] in, Indicates the Edit noise hidden variables at time steps; Indicates that the editing path is in The corrected noise at time steps; is the guidance strength factor, which is used to balance the reconstruction guidance and editing freedom; Indicates the Noise residual at time steps; Indicates the Editing shared noise of rendered images at time steps.
[0047] In a second aspect, the present invention provides a multi-view texture reconstruction system based on structure perception, comprising:
[0048] an encoding module for batch processing multi-view rendered images of a three-dimensional original scene using a pre-trained diffusion model combined with a control network, encoding the multi-view rendered images using an image encoder in the diffusion model, and encoding a text description of the original scene and a text description of the target scene respectively through a text encoder in the diffusion model;
[0049] An information injection module is used to extract structural information of the multi-view rendered image at each viewpoint, use the extracted structural information as a structural guidance condition, and pass it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network;
[0050] The denoising module is used to construct a dual-branch architecture for the reconstruction path and the editing path. A batch of noise is sampled under a standard Gaussian distribution as the initial noise latent variables for the editing and reconstruction paths. The UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. Both paths undergo multi-step iterative denoising. During the iterative denoising process, the noise sharing module performs a controllable weighted blending of the predicted noise corresponding to the same batch of multi-view rendered images. The dual-stream residual guidance module corrects the predicted noise generated by different batches of multi-view rendered images for the editing path.
[0051] The result acquisition module is used to optimize the three-dimensional original scene with the obtained edited images after all batches of multi-view rendering images are edited, so as to achieve multi-view consistent texture reshaping.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] The present invention introduces a structure-aware multi-view texture reconstruction framework, and on the basis of a pre-trained diffusion model, constructs a collaborative mechanism of intra-batch sharing and inter-batch guidance to achieve cross-view and cross-batch texture consistency editing. Specifically, in the texture reconstruction process, the multi-view images are input into the diffusion model in batches, and a noise sharing module is introduced to perform controllable weighted fusion of the prediction noise between views within the same batch, thereby improving the transmission capability of texture features and enhancing the consistency of views in the same batch; between batches, a dual-stream residual guidance module of reconstruction-editing dual paths is designed, which uses structural information as an anchor point to guide the latent variable optimization process in the editing path, thereby alleviating the information discontinuity problem between different batches and ensuring the continuity of geometric structure and semantic expression. Ultimately, the edited image can be fed back to the original three-dimensional scene to achieve multi-view scene reconstruction with consistent structure and unified texture. Based on the present invention, users can efficiently complete customized texture editing of complex three-dimensional scenes without relying on additional training, significantly improving the flexibility and quality of three-dimensional content creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flow chart of the steps of the method of the present invention;
[0055] Figure 2 Schematic diagram of the process of reshaping texture by the method of the present invention;
[0056] Figure 3 Comparison diagram of the editing results of the present invention and other texture reconstruction methods at different viewing angles;
[0057] Figure 4 This is a system block diagram of the present invention. DETAILED DESCRIPTION
[0058] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined accordingly without conflicting with each other.
[0059] In the description of the present invention, it should be understood that the terms "first" and "second" are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or implicitly specifying the number of technical features being described. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of such features.
[0060] In this embodiment, the technical background of the present invention is first described, which can be summarized into the following aspects:
[0061] 1) Texture reshaping: This refers to the process of replacing the style, changing the material, or semantically enhancing the texture of an object's surface while maintaining the stability of the three-dimensional structure. This process usually renders multiple perspective images from an existing three-dimensional scene, then uses an image editing model to partially or completely modify its texture, and finally projects the edited results back into the three-dimensional space to update the scene. This method has the advantages of strong flexibility and high efficiency, but it also places high demands on the alignment between texture and structure during the editing process. If the texture modification exceeds the structural boundary or lacks semantic continuity, it will directly affect the authenticity and scene consistency of the editing results. Therefore, the texture reshaping method must cooperate with structural constraints, accurately control the modification range, and ensure that the modified texture maintains a consistent appearance from different perspectives.
[0062] 2) Structural and texture consistency of multi-view images: In 3D scene editing tasks, multi-view consistency is a key factor affecting the quality of the final reconstruction. Structural consistency requires that objects in all views perfectly match in terms of geometric outline and spatial position; texture consistency requires that texture details and style remain coherent and seamless across different views. Because multi-view images are processed in batches, there is often a lack of connections and constraints within and between batches during the texturing process. The generated results between different views often suffer from structural distortion, texture shifts, or semantic conflicts. To address this issue, texture reconstruction methods need to establish a cross-view alignment mechanism and use structural information as an anchor to guide the synchronization of images from different views during the editing process. Furthermore, transition consistency between batches should be considered to prevent sudden style changes or texture discontinuities during multi-stage processing. Only by achieving high multi-view consistency in both structure and texture can the reconstructed 3D scene be stable and realistic.
[0063] like Figure 1 As shown, in a preferred implementation of the present invention, the multi-view texture reconstruction method based on structure perception includes the following steps S1 to S4. The specific implementation process is described below.
[0064] S1: Use a pre-trained diffusion model combined with a control network to batch process multi-perspective rendered images of a three-dimensional original scene, encode the multi-perspective rendered images using an image encoder in the diffusion model, and encode the text description of the original scene and the text description of the target scene respectively through a text encoder in the diffusion model for subsequent use.
[0065] S2: Extracting the structural information of the multi-view rendered image at each viewpoint, using the extracted structural information as a structural guidance condition, and passing it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network.
[0066] S3: Construct a dual-branch architecture of the reconstruction path and the editing path, sample a batch of noise under the standard Gaussian distribution as the initial noise latent variables of the editing path and the reconstruction path, and the UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. The two paths each perform multi-step iterative denoising; during the iterative denoising process, the noise sharing module performs controllable weighted mixing of the predicted noise corresponding to the same batch of multi-view rendered images, and the dual-stream residual guidance module corrects the predicted noise generated by the editing path for different batches of multi-view rendered images.
[0067] S4: After all batches of multi-view rendered images have been edited, the resulting edited images are used to optimize the original three-dimensional scene, achieving multi-view consistent texture reshaping.
[0068] The method of the present invention described in S1 to S4 above can efficiently complete customized texture editing of complex three-dimensional scenes in multi-view texture reconstruction tasks based on structure perception without relying on additional training, significantly improving the flexibility and quality of three-dimensional content creation.
[0069] In this embodiment, the above steps can be implemented in the following specific ways.
[0070] It should be noted that in step S1 of the present invention, the diffusion model includes a text encoder, an image encoder and a U-type network (i.e., UNet). The model is pre-trained on a large amount of text image data and has super image generation capabilities.
[0071] The structure-aware control network aims to provide additional, refined conditional input and control to the diffusion model. It comprises a trainable replica of the image encoder and zero convolutional layers within the diffusion model. It can accept additional structure-guided conditional inputs (such as edge maps, depth maps, or normal maps). This structural information is embedded as a constraint signal into the image generation and editing process, ensuring that the geometric outline and spatial layout of the target are preserved when modifying the texture, thereby achieving precise generative control over the pre-trained diffusion model. Specifically, the network introduces an auxiliary branch outside the main diffusion model that receives the structure-guided map and extracts its high-level semantic and spatial features. These features are then passed to each stage of the main network through a cross-layer fusion mechanism. During training, the main model parameters remain frozen, while the control network fine-tunes the generation process, achieving high-fidelity and highly consistent image editing. In the task of 3D scene texture reconstruction, the control network ensures accurate texture modification while effectively avoiding structural deformation or spatial misalignment, making it a key module for ensuring consistent multi-view editing.
[0072] It should be noted that if Figure 2 As shown, the specific steps of S1 include:
[0073] S11: a batch of multi-perspective rendered images is sequentially passed to the image encoder of the diffusion model for processing, and the image encoding corresponding to each perspective of the rendered image is obtained.
[0074] In this embodiment, for the rendered image of the current perspective in the batch, , and its corresponding image encoding is , It is an image encoder.
[0075] S12: Text description of the original scene and a textual description of the target scene Use the text encoder in the diffusion model respectively Encoding, corresponding to the text description encoding of the original scene and text description encoding of the target scene :
[0076]
[0077]
[0078] The above image encoding stores the structure and texture information of each perspective, while the text description encoding stores the semantic features of the original scene and the target scene, which lays an important foundation for the subsequent accurate and consistent texture reconstruction under multiple perspectives.
[0079] It should be noted that, in this embodiment, step S2 aims to maintain the consistency of the original structure during the texture reconstruction process. The specific steps of extracting structural information are: selecting a suitable structural feature extractor (such as depth, normal, edge, etc.) according to the 3D scene, and sequentially passing a batch of multi-view rendered images to the selected structural feature extractor for processing to obtain the structural information corresponding to each perspective of the rendered image. , and its corresponding structural information is , is the selected structural feature extractor.
[0080] It should be noted that in this embodiment, the noise sharing module in step S3 is designed to prevent inconsistent texture editing results between images rendered from different perspectives within the same batch. It performs a controllable weighted blend of the predicted noise corresponding to the multi-perspective rendered images of the same batch, so that the noise latent variable at each time step within the batch can obtain the texture information of other noise latent variables at that time step. The two-stream residual guidance module aims to alleviate the problem of information isolation between different batches. By constructing a dual-branch architecture of reconstruction and editing paths, structural information is used as an anchor to guide the optimization process of the edited image, providing a stable geometric semantic foundation for subsequent unified texture reconstruction.
[0081] It should be noted that in step S3, the specific processing process in the noise sharing module includes:
[0082] AS31: Under the standard Gaussian distribution, sample a batch of noise as the initial noise latent variable of the current batch , replicate the same noise to reconstruct the initial noise latent variable of the path .in, represents the standard Gaussian distribution, Represents the identity matrix.
[0083] AS32: Using initial noise latent variables On the reconstruction path Step iterative denoising, using the initial noise latent variable On the edit path Iterative denoising. Represents the total time steps of the denoising process.
[0084] AS33: For rebuilt paths, The noise latent variables at time steps are used as the reconstruction noise latent variables , the current time step And reconstruct the noise latent variables, the structural information of the current batch and the text description encoding of the original scene are passed into the UNet network to predict the noise ; For editing paths, change The noise latent variable at each time step is used as the editing noise latent variable , the current time step And edit the noise latent variables, the structural information of the current batch and the text description encoding of the target scene into the UNet network to predict the noise .
[0085] AS34: Preset a shared time step threshold range. If the current time step If the time step is not within the shared time step threshold, the noise is not shared and the predicted noise corresponding to each perspective rendering image of the current batch is directly used; if the current time step is Within the shared time step threshold, the predicted noise corresponding to each perspective rendered image in the current batch is traversed for weighted sharing.
[0086] In the present invention, the path is reconstructed Time step Prediction noise of rendered images from different perspectives As the first predicted noise, the average value of the predicted noise of the rendered images of other perspectives in the current batch is used as the first auxiliary information, and the first auxiliary information and the first predicted noise are weighted to form the reconstructed shared noise of the rendered image.
[0087] In the present invention, the editing path Time step Prediction noise of rendered images from different perspectives As the second predicted noise, the average value of the predicted noise of the rendered images from other perspectives in the current batch is used as the second auxiliary information, and the second auxiliary information and the second predicted noise are weighted to form the edited shared noise of the rendered image.
[0088] In this embodiment, the reconstruction of the rendered image shares noise Share noise with editors The specific calculation method is:
[0089]
[0090]
[0091] in, is the batch size, is the shared noise intensity; Reconstruction path Time step Prediction noise of rendered images from different perspectives; Indicates the editing path Time step Predicted noise for rendered images from different perspectives.
[0092] The noise sharing module introduces noise information from samples from other perspectives in the early stages of diffusion, enabling each sample to have stronger global context awareness, thereby effectively improving texture consistency. By stopping noise sharing in the later stages, the unique details of each perspective are retained, achieving a balance between global consistency and local fidelity.
[0093] It should be noted that in step S3, the specific processing process in the dual-stream residual guidance module includes:
[0094] BS31: For the reconstruction path time steps, encode the rendered image and the preset first proportional coefficient Multiply them to get the first calculation result, and reconstruct the noise latent variable at this time step Subtract the first calculation result to obtain the second calculation result, and multiply the reciprocal of the preset second proportional coefficient by the second calculation result as the ideal noise at this time step. ; Among them, the first proportional coefficient is the reconstruction path The square root of the guidance coefficient of the time step, the second proportional coefficient is 1 and the reconstruction path is The square root of the difference between the bootstrap coefficients at each time step.
[0095] In this embodiment, the reconstruction path ideal noise at time steps The calculation method is:
[0096]
[0097] in, represents the first proportionality coefficient; is the second proportional coefficient; It is The guidance coefficient for each time step.
[0098] BS32: The ideal noise at time steps is shared with the reconstructed noise of the rendered image Subtract, as the noise residual at this time step .
[0099] BS33: Use ideal noise at time steps The reconstruction noise latent variable of this time step in the reconstruction path Perform denoising operation to generate the The reconstruction noise latent variable of time steps .
[0100] It should be noted that in step BS33 of the present invention, the specific processing process of the denoising operation is: ideal noise at time steps Multiply it by the second proportional coefficient to get the third calculation result. The reconstruction noise latent variable at time steps Subtract the third calculation result from the fourth calculation result to obtain the fourth calculation result, divide the fourth calculation result by the first proportional coefficient to obtain the fifth calculation result, multiply the fifth calculation result by the preset third proportional coefficient to obtain the sixth calculation result, and multiply the fifth calculation result by the preset third proportional coefficient to obtain the sixth calculation result. ideal noise at time steps The seventh calculation result is obtained by multiplying the sixth calculation result and the seventh calculation result by the preset fourth proportional coefficient. The reconstruction noise latent variable of time steps ; Among them, the third proportional coefficient is the reconstruction path The square root of the guidance coefficient of the time step, the fourth proportional coefficient is 1 and the reconstruction path is The square root of the difference between the bootstrap coefficients at each time step.
[0101] In this embodiment, the The reconstruction noise latent variable of time steps The specific calculation method is:
[0102]
[0103] in, represents the third proportional coefficient; is the fourth proportional coefficient; It is The guidance coefficient for each time step.
[0104] BS34: For the first time steps, the noise residual As an anchor guide and multiplied by the guide strength factor to form a noise modifier, the noise modifier is shared with the edit of the rendered image at this time step Add together to get the noise of the edit path modified at this time step .
[0105] In this embodiment, the editing path is The corrected noise at time steps The calculation is as follows:
[0106]
[0107] in, is the guidance strength factor, which is used to balance the reconstruction guidance and editing freedom.
[0108] BS35: Use The corrected noise at time steps The edit noise latent variable for this time step in the edit path Perform denoising operation to generate the The editing noise latent variable of time steps .
[0109] It should be noted that in step BS35 of the present invention, the specific processing process of the denoising operation is: The corrected noise at time steps Multiply it by the second proportional coefficient to get the eighth calculation result. Editing noise latent variables at time steps Subtract the eighth calculation result from the ninth calculation result, divide the ninth calculation result by the first proportional coefficient to obtain the tenth calculation result, multiply the tenth calculation result by the third proportional coefficient to obtain the eleventh calculation result, and multiply the tenth calculation result by the third proportional coefficient to obtain the eleventh calculation result. The corrected noise at time steps The twelfth calculation result is obtained by multiplying the eleventh calculation result and the twelfth calculation result by the sum of the eleventh calculation result and the twelfth calculation result. The editing noise latent variable of time steps .
[0110] In this embodiment, the The editing noise latent variable of time steps The specific calculation method is:
[0111]
[0112] The above-mentioned dual-stream residual guidance module effectively transfers the deviation signal of UNet's understanding of different perspective structures and textures from the reconstruction path to the editing path, alleviating the problems of texture offset and uneven editing intensity caused by perspective differences during the editing process, thereby improving the consistency and coordination of cross-perspective editing results in texture expression and structure preservation.
[0113] The following tests apply the structure-aware multi-view texture reconstruction method described in steps S1 through S4 of the preferred implementation to a specific dataset. The specific steps are as described in S1 through S4 and will not be repeated here. The focus will be on demonstrating the specific parameters and technical results.
[0114] Example
[0115] Following the implementation process of steps S1 to S4 above, a 3D scene dataset collected from the real world is first acquired for 3D scene reconstruction, followed by rendering of multi-view images. During the texture reshaping process, the multi-view rendered images are fed into a diffusion model in batches, and a noise sharing module is introduced to perform controllable weighted fusion of the prediction noise between views within the same batch, improving the transferability of texture features and enhancing the consistency of views within the same batch. Between batches, a two-stream residual guidance module for the reconstruction-editing dual path is designed. Using structural information as an anchor, it guides the latent variable optimization process in the editing path, thereby alleviating the information gap between different batches and ensuring the continuity of geometric structure and semantic expression. Ultimately, the edited image can be fed back into the original 3D scene, achieving multi-view scene reconstruction with consistent structure and uniform texture.
[0116] As a quantitative indicator, this embodiment compares the method of the present invention with four existing methods: IN2N (GS), GaussianEditor, GaussCtrl, and DGE. Three indicators are selected: CLIP text-image direction similarity, CLIP direction consistency, and user preference. The test results of the above methods on the IN2N dataset are shown in Table 1. The editing effects of the method of the present invention and the above four existing methods are compared. Figure 3 shown.
[0117] Table 1 Quantitative evaluation table of editing effect
[0118] It should also be noted that the multi-view texture reconstruction method based on structure perception in the above embodiment can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a multi-view texture reconstruction system based on structure perception corresponding to the multi-view texture reconstruction method based on structure perception provided in the above embodiment, such as Figure 4 As shown, it includes:
[0119] an encoding module for batch processing multi-view rendered images of a three-dimensional original scene using a pre-trained diffusion model combined with a control network, encoding the multi-view rendered images using an image encoder in the diffusion model, and encoding a text description of the original scene and a text description of the target scene respectively through a text encoder in the diffusion model;
[0120] An information injection module is used to extract structural information of the multi-view rendered image at each viewpoint, use the extracted structural information as a structural guidance condition, and pass it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network;
[0121] The denoising module is used to construct a dual-branch architecture for the reconstruction path and the editing path. A batch of noise is sampled under a standard Gaussian distribution as the initial noise latent variables for the editing and reconstruction paths. The UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. Both paths undergo multi-step iterative denoising. During the iterative denoising process, the noise sharing module performs a controllable weighted blending of the predicted noise corresponding to the same batch of multi-view rendered images. The dual-stream residual guidance module corrects the predicted noise generated by different batches of multi-view rendered images for the editing path.
[0122] The result acquisition module is used to optimize the three-dimensional original scene with the obtained edited images after all batches of multi-view rendering images are edited, so as to achieve multi-view consistent texture reshaping.
[0123] It should also be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods, for example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.
[0124] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.
Claims
1. A multi-view texture reconstruction method based on structure perception, characterized in that: The following steps are involved: S1: using a pre-trained diffusion model combined with a control network to batch process multi-view rendered images of a 3D original scene, encoding the multi-view rendered images using an image encoder in the diffusion model, and encoding the text description of the original scene and the text description of the target scene respectively through a text encoder in the diffusion model; S2: extracting the structural information of the multi-view rendered image at each viewpoint, using the extracted structural information as a structural guidance condition, and passing it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network; S3: Construct a dual-branch architecture for the reconstruction path and the editing path. A batch of noise is sampled under a standard Gaussian distribution as the initial noise latent variables for the editing and reconstruction paths. The UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. Both paths perform multi-step iterative denoising. During the iterative denoising process, the noise sharing module performs a controllable weighted mixing of the predicted noise corresponding to the same batch of multi-view rendered images. The two-stream residual guidance module corrects the predicted noise generated by the editing path for different batches of multi-view rendered images. S4: After all batches of multi-view rendered images have been edited, the resulting edited images are used to optimize the original three-dimensional scene, achieving multi-view consistent texture reshaping.
2. The multi-view texture reconstruction method based on structure perception according to claim 1, characterized in that: In step S1, the control network includes a copy of the image encoder in the diffusion model and zero convolutional layers.
3. The multi-view texture reconstruction method based on structure perception according to claim 1, characterized in that: The specific steps in step S1 include: S11: sequentially passing a batch of multi-view rendered images to the image encoder of the diffusion model for processing, and obtaining image codes corresponding to the rendered images of each view; S12: Encode the text description of the original scene and the text description of the target scene using the text encoder in the diffusion model respectively, to obtain the text description encoding of the original scene and the text description encoding of the target scene respectively.
4. The multi-view texture reconstruction method based on structure perception according to claim 1, characterized in that: The specific steps of extracting the structural information in step S2 are: sequentially inputting a batch of multi-view rendered images into a structural feature extractor selected according to the three-dimensional scene for processing, and obtaining the structural information corresponding to the rendered images of each view.
5. The multi-view texture reconstruction method based on structure perception according to claim 1, characterized in that: In step S3, the specific processing process in the noise sharing module includes: AS31: Under the standard Gaussian distribution, a batch of noise is sampled as the initial noise latent variable of the current batch, and the same noise is copied to reconstruct the initial noise latent variable of the path; AS32: Performs T-step iterative denoising on the reconstruction path using the initial noise latent variables of the reconstruction path, and performs T-step iterative denoising on the editing path using the sampled initial noise latent variables. AS33: For the reconstruction path, the noise latent variable at the t-th time step is used as the reconstruction noise latent variable, and the current time step t and the reconstruction noise latent variable, the structural information of the current batch, and the text description of the original scene are encoded and passed into the UNet network to predict the noise; for the editing path, the noise latent variable at the t-th time step is used as the editing noise latent variable, and the current time step t and the editing noise latent variable, the structural information of the current batch, and the text description of the target scene are encoded and passed into the UNet network to predict the noise; AS34: A shared time step threshold range is preset. If the current time step t is not within the shared time step threshold range, noise sharing is not performed, and the predicted noise corresponding to each perspective rendered image in the current batch is directly used. If the current time step t is within the shared time step threshold range, the predicted noise corresponding to each perspective rendered image in the current batch is traversed for weighted sharing.
6. The multi-view texture reconstruction method based on structure perception according to claim 5, characterized in that: In step AS34, the predicted noise of the rendered image of the vth view at the tth time step in the reconstruction path is used as the first predicted noise, and the average of the predicted noises of the rendered images of other views in the current batch is used as the first auxiliary information. The first auxiliary information and the first predicted noise are weighted to form the reconstruction shared noise of the rendered image. The predicted noise of the v-th perspective rendered image at the t-th time step of the editing path is used as the second predicted noise, and the average of the predicted noises of the rendered images of other perspectives in the current batch is used as the second auxiliary information. The second auxiliary information and the second predicted noise are weighted to form the editing shared noise of the rendered image.
7. The multi-view texture reconstruction method based on structure perception according to claim 6, characterized in that: In step S3, the specific processing process in the dual-stream residual guidance module includes: BS31: For the t-th time step of the reconstruction path, multiply the image code of the rendered image by a preset first scaling factor to obtain a first calculation result, subtract the first calculation result from the reconstruction noise latent variable at that time step to obtain a second calculation result, and multiply the reciprocal of the preset second scaling factor by the second calculation result to obtain the ideal noise at that time step; wherein the first scaling factor is the arithmetic square root of the guided coefficient at the t-th time step of the reconstruction path, and the second scaling factor is the arithmetic square root of the difference between 1 and the guided coefficient at the t-th time step of the reconstruction path; BS32: Subtract the ideal noise at the t-th time step from the reconstructed shared noise of the rendered image as the noise residual at that time step; BS33: Use the ideal noise at the t-th time step to denoise the reconstruction noise latent variable at that time step in the reconstruction path to generate the reconstruction noise latent variable at the t-1-th time step; BS34: For the t-th time step of the editing path, the noise residual is used as the anchor guide and multiplied by the guide strength factor to form a noise correction. The noise correction is then added to the edit shared noise of the rendered image at that time step to obtain the corrected noise of the editing path at that time step. BS35: Use the corrected noise at the t-th time step to denoise the editing noise latent variable at that time step in the editing path to generate the editing noise latent variable at the t-1-th time step.
8. The multi-view texture reconstruction method based on structure perception according to claim 7, characterized in that: The specific processing process of the denoising operation in step BS33 is: multiplying the ideal noise at the t-th time step by the second proportional coefficient to obtain a third calculation result, subtracting the reconstruction noise latent variable at the t-th time step from the third calculation result to obtain a fourth calculation result, dividing the fourth calculation result by the first proportional coefficient to obtain a fifth calculation result, multiplying the fifth calculation result by the preset third proportional coefficient to obtain a sixth calculation result, multiplying the ideal noise at the t-th time step by the preset fourth proportional coefficient to obtain a seventh calculation result, and adding the sixth calculation result and the seventh calculation result as the reconstruction noise latent variable at the t-1-th time step; wherein the third proportional coefficient is the arithmetic square root of the guide coefficient of the reconstruction path at the t-1-th time step, and the fourth proportional coefficient is the arithmetic square root of the difference between 1 and the guide coefficient of the reconstruction path at the t-1-th time step.
9. The multi-view texture reconstruction method based on structure perception according to claim 7, characterized in that: The specific processing process of the denoising operation in step BS35 is: multiplying the corrected noise at the t-th time step by the second proportional coefficient to obtain the eighth calculation result, subtracting the editing noise latent variable at the t-th time step from the eighth calculation result to obtain the ninth calculation result, dividing the ninth calculation result by the first proportional coefficient to obtain the tenth calculation result, multiplying the tenth calculation result by the third proportional coefficient to obtain the eleventh calculation result, multiplying the corrected noise at the t-th time step by the fourth proportional coefficient to obtain the twelfth calculation result, and adding the eleventh calculation result and the twelfth calculation result as the editing noise latent variable at the t-1-th time step.
10. A multi-view texture reconstruction system based on structure perception, characterized in that: include: an encoding module for batch processing multi-view rendered images of a three-dimensional original scene using a pre-trained diffusion model combined with a control network, encoding the multi-view rendered images using an image encoder in the diffusion model, and encoding a text description of the original scene and a text description of the target scene respectively through a text encoder in the diffusion model; An information injection module is used to extract structural information of the multi-view rendered image at each viewpoint, use the extracted structural information as a structural guidance condition, and pass it to each stage of the UNet network in the diffusion model through the cross-attention mechanism by the control network; The denoising module is used to construct a dual-branch architecture for the reconstruction path and the editing path. A batch of noise is sampled under a standard Gaussian distribution as the initial noise latent variables for the editing and reconstruction paths. The UNet network of the reconstruction path takes the initial noise latent variables, the text description encoding of the original scene, and the extracted structural information as input. The UNet network of the editing path takes the initial noise latent variables, the text description encoding of the target scene, and the extracted structural information as input. Both paths undergo multi-step iterative denoising. During the iterative denoising process, the noise sharing module performs a controllable weighted blending of the predicted noise corresponding to the same batch of multi-view rendered images. The dual-stream residual guidance module corrects the predicted noise generated by different batches of multi-view rendered images for the editing path. The result acquisition module is used to optimize the three-dimensional original scene with the obtained edited images after all batches of multi-view rendering images are edited, so as to achieve multi-view consistent texture reshaping.
Citation Information
Patent Citations
CG image detection method based on double-flow neural network channel combination and soft pooling
CN115410029A
Noise point suppression method and device based on neural radiation field three-dimensional reconstruction and electronic equipment
CN117391990A
Construction method and system of three-dimensional model and image rendering method
CN118762123A
Three-dimensional texture reconstruction method and system guided by multi-view semantic segmentation information
CN118864713A
Orchard three-dimensional reconstruction and fruit tree semantic segmentation method based on neural radiation field
CN119516195A
Cited By
Three-dimensional scene texture generation method and system based on visual guidance
CN121564177A