A single-view 3D scene reconstruction method, device and medium based on diffusion model

Through the diffusion model and offset camera strategy, combined with MLLM filling and iterative noise reduction processing, the problems of view consistency and clarity in single-view 3D scene reconstruction are solved, and high-quality and realistic 3D scenes are generated.

CN120259571BActive Publication Date: 2025-08-29SHANDONG UNIV OF FINANCE & ECONOMICS +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510747860.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-29
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, single-view three-dimensional scene reconstruction method is difficult to ensure the consistency and clarity of multiple views at the same time when ensuring the diversity and rationality of new perspectives.

Method used

Using a diffusion model-based approach, through offset camera strategy analysis, MLLM fill area, noise addition strategy and iterative noise reduction processing, combined with prediction-refinement loop optimization, visually consistent 3D scenes are generated, and object replacement and addition of single-views are supported.

Benefits of technology

It improves the diversity, consistency and consistency of three-dimensional scene reconstruction, generates high-quality and realistic 3D scenes, and solves the problem that view consistency and clarity cannot be met at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259571B_ABST
    Figure CN120259571B_ABST
Patent Text Reader

Abstract

The present application discloses a single-view 3D scene reconstruction method, device, and medium based on a diffusion model, which relates to the technical field of three-dimensional scene reconstruction. The method includes: obtaining a camera trajectory, and based on the camera trajectory, obtaining a set of views to be filled through an offset camera strategy analysis; filling the blank areas of the set of views to be filled to obtain a filled view, and performing a refinement loop optimization on the filled view to obtain the current reconstructed 3D scene; according to the current 3D scene, determining a noisy rendered image group through a diffusion model forward processing with a noise addition strategy; performing an iterative denoising reverse processing of the diffusion model on the noisy rendered image group to obtain a visually consistent 3D scene; inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object. The present application solves the technical problem in the prior art that the consistency and clarity of the three-dimensional scene reconstruction views cannot be simultaneously met through the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of three-dimensional scene reconstruction technology, in particular to a single-view method based on a diffusion model. Figure 3 D scene reconstruction method, equipment and medium. Background Art

[0002] Three-dimensional scene reconstruction is a key task in fields such as augmented reality (AR), virtual reality (VR), and computer graphics. With the development of 3DGS technology, 3D reconstruction has been significantly promoted, enabling efficient use of abundant data for scene reconstruction. However, reconstruction quality is highly dependent on large-scale, high-precision 3D data, but the acquisition of this data is time-consuming and costly. Therefore, single-view-based 3D reconstruction methods have emerged to overcome the limitations of data acquisition and improve reconstruction efficiency.

[0003] However, a single view contains limited information and it is difficult to fully present the hidden structure of an object. In recent years, the diffusion model, as a generative technology, has achieved remarkable results in image generation and completion tasks. It can generate potential perspectives based on a single view, effectively making up for the lack of single view information. These methods still face challenges in large-scale 3D reconstruction, especially in terms of repeated, unreasonable and inconsistent content in multiple views. Previous methods such as "Single Image to Novel Views with Semantic-Preserving Generative Warping" and "Text-driven 3d scene generation withinpainting and depth diffusion. In International Confer ence on 3D Vision (3DV)" focus on single views and newly generated views. Figure 1 consistency, but ignores global consistency; recent methods such as VistaDream: Sampling multiview consistent images for single-view scenereconstruction can ensure multi-view Figure 1 However, it will lead to blurring and reduce the reconstruction quality. Therefore, how to ensure the diversity and rationality of new perspectives while ensuring the multi-view Figure 1 Consistency and clarity are the core issues in 3D scene reconstruction. Summary of the Invention

[0004] The embodiment of the present application provides a single-view Figure 3 3D scene reconstruction method, device and medium solve the existing problems of 3D scene reconstruction Figure 1The technical problem is that consistency and clarity cannot be met at the same time.

[0005] In the first aspect, the embodiment of the present application provides a single-view method based on a diffusion model. Figure 3 A 3D scene reconstruction method is characterized in that the method includes: obtaining a camera trajectory, and based on the camera trajectory, obtaining a view set to be filled by offset camera strategy analysis; filling blank areas of the view set to be filled to obtain filled views, and performing refinement and loop optimization on the filled views to obtain a reconstructed current 3D scene; according to the current 3D scene, determining a noisy rendered image group through forward processing of a diffusion model with a noise addition strategy; performing iterative denoising reverse processing of the diffusion model on the noisy rendered image group to obtain a visually consistent 3D scene; and inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene in which the changed object is replaced.

[0006] In one implementation of the present application, based on the camera trajectory, an offset camera strategy analysis is performed to obtain a set of views to be filled, specifically including: based on the camera trajectory, an updated extrinsic parameter is obtained by controlling the camera offset angle; based on the updated extrinsic parameter, an updated internal parameter is obtained by adjusting the rendering view; based on the updated extrinsic parameter and the updated internal parameter, an offset camera is determined, and based on the offset camera, a set of views to be filled is obtained by sampling the 3D scene.

[0007] In one implementation of the present application, a blank area of ​​a view set to be filled is filled to obtain a filled view, specifically including: obtaining MLLM extraction parameters, and based on the MLLM extraction parameters, predicting the filling area of ​​the view set to be filled to obtain an object of the area to be filled; wherein the MLLM extraction parameters include: global hints and local hints; according to the object of the area to be filled, a filling view is obtained through a filling area mask threshold analysis.

[0008] In one implementation of the present application, a refinement loop optimization is performed on the filled view to obtain the reconstructed current 3D scene, specifically including: performing a refinement loop optimization based on the filled view, and obtaining view evaluation data through MLLM consistency evaluation; wherein, the MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, and view style consistency evaluation; quality screening is performed on the view evaluation data to determine the current optimal view, and based on the current optimal view, the MLLM descriptive prompt loop update is performed to obtain the reconstructed current 3D scene.

[0009] In one implementation of the present application, according to the current 3D scene, a noisy rendered image group is determined by forward processing of a diffusion model of a noise adding strategy, specifically comprising: generating a rendered image group based on the current 3D scene by image rendering; performing cosine similarity analysis of adjacent images of the rendered image group with adjacent image influence intensity control to obtain cosine similarity of adjacent images; and determining the cosine similarity of adjacent images by visually comparing the cosine similarity of adjacent images. Figure 1 Add consistent random noise to determine the group of noisy rendered images.

[0010] In one implementation of the present application, based on the cosine similarity of adjacent images, Figure 1 After adding consistent random noise to determine the noisy rendered image group, the method also includes: determining an attention parameter based on the noisy rendered image group through image potential feature analysis; according to the attention parameter, obtaining an update component through random sampling of the Batch dimension; wherein the update component includes: a key and a value; splicing the update component with the corresponding key and value, and performing attention analysis on the spliced ​​update component to determine the correlation-aware attention mechanism.

[0011] In one implementation of the present application, the diffusion model of iterative denoising is reversely processed on the noisy rendered image group to obtain a visually consistent 3D scene, specifically including: inputting the noisy rendered image group into a preset denoising network, and correcting the denoising direction of the denoising network through a 3D Gaussian field to obtain the denoising result of the current time step; based on the denoising results of the noisy rendered image group and the current time step, determining the deviation of the anti-denoising direction from the origin; according to the deviation of the anti-denoising direction from the origin, pulling back through the denoising direction to obtain a visually consistent 3D scene.

[0012] In one implementation of the present application, a visually consistent 3D scene is input into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object, specifically including: rendering the target to be edited of the visually consistent 3D scene to determine the target to be edited; based on the target to be edited, performing gradient descent optimization on the scene parameters of the target to be edited to obtain scene optimization parameters; wherein the constraints of the gradient descent optimization include: depth loss constraint; according to the scene optimization parameters, performing mask area extraction on the single-view replacement target of the target to be edited to obtain the mask area of ​​the source object, and performing replacement changes on the mask area to determine the 3D scene that replaces the changed object.

[0013] In the second aspect, the embodiment of the present application also provides a single-view method based on a diffusion model. Figure 3A 3D scene reconstruction device, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to: obtain a camera trajectory, and based on the camera trajectory, obtain a set of views to be filled by analyzing an offset camera strategy; fill blank areas of the set of views to be filled to obtain filled views, and perform refinement loop optimization on the filled views to obtain a reconstructed current 3D scene; according to the current 3D scene, determine a noisy rendered image group by forward processing of a diffusion model with a noise addition strategy; perform iterative noise reduction on the noisy rendered image group by reverse processing of a diffusion model with iterative noise reduction to obtain a visually consistent 3D scene; and input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object.

[0014] In a third aspect, the embodiment of the present application also provides a single-view method based on a diffusion model. Figure 3 A non-volatile computer storage medium for 3D scene reconstruction stores computer-executable instructions, wherein the computer-executable instructions are configured to: obtain a camera trajectory, and based on the camera trajectory, obtain a set of views to be filled by performing an offset camera strategy analysis; fill blank areas in the set of views to be filled to obtain filled views, and perform a refinement loop optimization on the filled views to obtain a reconstructed current 3D scene; determine a noisy rendered image group based on the current 3D scene by performing a diffusion model forward processing with a noise addition strategy; perform an iterative denoising diffusion model reverse processing on the noisy rendered image group to obtain a visually consistent 3D scene; and input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object.

[0015] The embodiment of the present application provides a single-view Figure 3 The 3D scene reconstruction method, device and medium use a 2D diffusion model to expand a single view to generate a rough 3D scene, and combine it with a prediction-refinement cycle method to dynamically fill in the content of adjacent views and the original view to improve scene diversity and consistency. At the same time, an offset camera strategy is proposed to deflect the current view posture and use global prompt words to fill in, gradually iterating to generate a coherent 3D scene. To further improve multi-view Figure 1 In order to ensure the consistency of the image, this paper designs a correlation-aware attention and correlation noise addition strategy to enhance the correlation between adjacent views, and introduces an origin constraint to correct the denoising direction to avoid scene blur. In terms of scene editing, based on the replacement and addition of objects in a single view, by extracting masks and filling them, combined with the depth loss constraint to ensure the accuracy and consistency of the edited objects, it solves the existing problems in 3D scene reconstruction. Figure 1 The technical problem is that consistency and clarity cannot be met at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 A single-view method based on a diffusion model is provided in the embodiment of the present application. Figure 3 D. Flowchart of scene reconstruction method;

[0018] Figure 2 A schematic diagram of object replacement and addition in a 3D scene provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of 3D reconstruction of a different camera trajectory provided in an embodiment of the present application is shown. Figure 3 (a) is a schematic diagram of 3D reconstruction when the camera trajectory moves backward, and (b) is a schematic diagram of 3D reconstruction when the camera trajectory moves forward;

[0020] Figure 4 A graph showing a noise addition strategy ablation study provided in an embodiment of the present application;

[0021] Figure 5 A single-view method based on a diffusion model is provided in the embodiment of the present application. Figure 3 Schematic diagram of the internal structure of the D scene device. DETAILED DESCRIPTION

[0022] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] The embodiment of the present application provides a single-view Figure 3 The 3D scene reconstruction method, device and medium use a 2D diffusion model to expand a single view to generate a rough 3D scene, and combine it with a prediction-refinement cycle method to dynamically fill in the content of adjacent views and the original view to improve scene diversity and consistency. At the same time, an offset camera strategy is proposed to deflect the current view posture and use global prompt words to fill in, gradually iterating to generate a coherent 3D scene. To further improve multi-view Figure 1In order to ensure consistency, this paper designs correlation-aware attention and correlation noise addition strategies to enhance the correlation between adjacent views, and introduces origin constraints to correct the denoising direction to avoid scene blur. In terms of scene editing, this application also supports object replacement and addition based on a single view. By extracting masks and filling them, combined with depth loss constraints, the edited objects are ensured to be accurate and consistent, solving the existing problems in 3D scene reconstruction. Figure 1 The technical problem is that consistency and clarity cannot be met at the same time.

[0024] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0025] Figure 1 A single-view method based on a diffusion model is provided in the embodiment of the present application. Figure 3 D scene reconstruction method flow chart. Figure 1 As shown, the embodiment of the present application provides a single-view method based on a diffusion model. Figure 3 The D scene reconstruction method specifically includes the following steps:

[0026] Step 101: Obtain a camera trajectory, and based on the camera trajectory, obtain a view set to be filled by analyzing the offset camera strategy.

[0027] For example, if all views sampled from the camera trajectory are filled with predicted objects, while this introduces a certain degree of diversity, it can easily lead to scene discontinuity and even deviation from the original scene. Furthermore, since the blank area between adjacent views typically accounts for 50%, the filled objects may be inconsistent, resulting in incomplete object generation and, in turn, visual distortion and other issues. By analyzing the offset camera strategy, the filling results for the set of views to be filled are complete, improving the consistency of the filled objects.

[0028] Specifically, based on the camera trajectory, the offset camera strategy analysis is performed to obtain the view set to be filled, including: based on the camera trajectory, the camera offset angle control is used to obtain updated extrinsic parameters; according to the updated extrinsic parameters, the updated intrinsic parameters are obtained by adjusting the rendering view; based on the updated extrinsic parameters and the updated intrinsic parameters, the offset camera is determined, and according to the offset camera, the view set to be filled is obtained by sampling the 3D scene.

[0029] In one embodiment, if all views sampled by the camera trajectory are filled with predicted objects, although a certain degree of diversity is introduced, it is easy to cause scene discontinuity and may even deviate from the original scene. Moreover, the blank area of ​​adjacent views usually occupies 50%, and the filling objects may be inconsistent, resulting in incomplete object generation, which in turn causes problems such as visual distortion. In order to effectively enhance the scene coherence and ensure the complete generation of objects, we propose an offset camera strategy. The preset camera trajectory is , where the camera Including intrinsic and extrinsic parameters, which can render new views To simplify, the camera The height and width of the internal parameters To change the width and height of the render view, camera position and target location Controlling the Camera Based on this camera parameter, we adjust or Parameters, and introduce an offset angle to construct an offset camera. The above offset camera construction process is explained by the following formula:

[0030] (1)

[0031] in, 、 、 is a predefined deflection angle;

[0032] Indicates winding Axis rotation, Indicates that around Axis rotation, Indicates that around Axis rotation;

[0033] 、 、 represents a unit vector;

[0034] Quaternion multiplication is defined, for The conjugate of ; where the quaternion is and .

[0035] By controlling the deflection angle, the new position coordinates can be calculated , and combined with the camera position pos to form a new external parameter, while adjusting the height of the rendering view and width The parameters change the intrinsic parameters of the camera, thus generating a deflection camera.

[0036] Using two sets of parameters and Construct two deflection cameras respectively, denoted as and ;in, The zoom factor is used to control the zooming in and out of the render view.

[0037] for ,The deflection angle is generally set to negative to capture more known scene information.,When filling the view rendered by the deflected camera, a global prompt is,used to ensure the coherence of the scene.

[0038] It should be noted that The deflection angle is generally set to positive, and a larger and value, so as to complete the predicted object while ensuring that the object is reasonably integrated into the current scene.

[0039] Furthermore, after determining the two offsets, the 3D frame is sampled through the camera trajectory to obtain the view set to be filled .

[0040] Step 102 : Fill the blank areas of the to-be-filled view set to obtain filled views, and perform refinement and loop optimization on the filled views to obtain the reconstructed current 3D scene.

[0041] For example, existing techniques leverage large multimodal language models (MLLMs) to understand the original view, but use the same cues for iterative infilling, which can easily lead to scene homogeneity when generating complex perspectives (such as 360-degree surround views). To prevent this problem and enhance the potential of MLLMs, we leverage the original view information, breaking through its content limitations and building a self-optimizing framework to enhance scene diversity.

[0042] Specifically, the blank area of ​​the view set to be filled is filled to obtain a filled view, including: obtaining MLLM extraction parameters, and predicting the filling area of ​​the view set to be filled based on the MLLM extraction parameters to obtain the area object to be filled; wherein the MLLM extraction parameters include: global prompts and local prompts; according to the area object to be filled, the filling view is obtained through the filling area mask threshold analysis.

[0043] Furthermore, the filled view is refined and cyclically optimized to obtain the reconstructed current 3D scene, specifically including: performing refined cyclic optimization based on the filled view, and obtaining view evaluation data through MLLM consistency evaluation; wherein, MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, and view style consistency evaluation; quality screening of the view evaluation data to determine the current optimal view, and according to the current optimal view, cyclically updating through MLLM descriptive prompts to obtain the reconstructed current 3D scene.

[0044] Exemplarily, for a given original view, MLLM is first used to extract global cues and local cues; the global cues include the overall information and style of the scene, while the local cues are a detailed description of the original view.

[0045] The original view is expanded and filled using local hints to generate a preliminary 3D frame. This 3D frame is then sampled using the camera trajectory to obtain a set of views, and a diffusion model is used to fill in the blank areas in each view.

[0046] For each view When filling, we combine global cues and information from adjacent views to predict the possible objects. The mask of the filled area is usually set to around 50% to meet the generation requirements of most objects. When the filled area is less than 10%, only global hints are used for filling.

[0047] After generating a complete view, we use MLLM to evaluate its quality, content and style consistency with the original view and global hints. Finally, we perform quality screening on the view evaluation data and select the best view to generate more descriptive hints using MLLM to further optimize the quality. , to produce higher quality views.

[0048] Through the above user-friendly self-improvement process, this application can gradually generate images with high-quality and diverse content and styles, and reconstruct the entire 3D scene through continuous cycles.

[0049] Furthermore, the specific algorithm logic of the 3D reconstruction method of offset camera strategy analysis and fusion prediction-refinement loop optimization is expressed as follows.

[0050] Input: Filling the model , number of iterations ,Camera trajectory collection , global prompt , deflection angle , scaling factor ;

[0051] Output: 3D scene.

[0052] cycle :

[0053] Select from a collection of camera tracks , ;

[0054] Equal to the offset camera strategy ;

[0055] Equal to the offset camera strategy ;

[0056] pass Sampling a view from a 3D scene ;

[0057] ;

[0058] use Update the 3D scene;

[0059] prediction-refinement cycle approach;

[0060] pass Sampling a view from a 3D scene ;

[0061] ;

[0062] use Update the 3D scene;

[0063] Finish;

[0064] Return to the 3D scene.

[0065] Step 103: According to the current 3D scene, a noisy rendered image group is determined by forward processing of a diffusion model using a noise addition strategy.

[0066] For example, the 3D scene reconstructed based on the offset camera strategy analysis and fusion prediction-refinement loop optimization still has some defects, mainly manifested in that the camera trajectory cannot cover all view angles, resulting in tiny holes in some rendering results, and the image has multi-view inconsistency problems. When sampling along a predefined camera trajectory, if the trajectory coverage is small, there is usually a strong correlation between the views. However, with the complexity of the trajectory sampling method (such as straight forward, backward or 360° rotation), the A view is only strongly correlated with the adjacent views before and after. In order to enhance the correlation between views, Figure 1 A consistent diffusion model is used to process 3D scenes to solve the above problems and improve the consistency and fidelity of 3D scene reconstruction. The diffusion model consists of two stages. In the forward process, noise needs to be added to the image.

[0067] Specifically, according to the current 3D scene, a noisy rendered image group is determined by forward processing of a diffusion model of a noise adding strategy, including: based on the current 3D scene, generating a rendered image group by image rendering; performing cosine similarity analysis of adjacent images of the rendered image group with adjacent image influence intensity control to obtain cosine similarity of adjacent images; according to the cosine similarity of adjacent images, Figure 1 Add consistent random noise to determine the group of noisy rendered images.

[0068] In one embodiment, a set of images is first rendered from the 3D scene according to the camera trajectory, and then a 2D diffusion model is applied to enforce consistency constraints between these views. Finally, the 3D scene is updated with the improved views, thereby improving the consistency and fidelity of the reconstruction.

[0069] In the forward process of the diffusion model, given rendered images These images need to be gradually added with noise. However, directly introducing random noise may destroy the multi-viewing between images. Figure 1 This makes it difficult for the denoising network to establish consistency between views within a given time step. To this end, noise is added using a related noise addition strategy, which is explained in detail in the following formula.

[0070] (2)

[0071] (3)

[0072] (4)

[0073] (5)

[0074] in, is a predefined random noise constant;

[0075] is an independent random noise, is the cosine similarity of adjacent images;

[0076] It is used to control the influence of adjacent images on the current image, and is also a constant.

[0077] when When , the strategy degenerates into introducing independent random noise to all images; It is set to 0.2 to add appropriate randomness while maintaining consistency between views. The main goal of this strategy is to provide better initial conditions for the inverse generation process of the diffusion model, thereby achieving multi-view in fewer time steps. Figure 1 Consistent construction.

[0078] Furthermore, according to the cosine similarity of adjacent images, Figure 1After adding consistent random noise to determine the noisy rendered image group, the method also includes: determining an attention parameter based on the noisy rendered image group through image potential feature analysis; according to the attention parameter, obtaining an update component through random sampling of the Batch dimension; wherein the update component includes: a key and a value; splicing the update component with the corresponding key and value, and performing attention analysis on the spliced ​​update component to determine the correlation-aware attention mechanism.

[0079] In one embodiment, based on the characteristics of video consistency processing, the present application improves a correlation-aware attention to replace self-attention, thereby enhancing the correlation between the current image and its adjacent images.

[0080] Formally, given an image latent feature ,in 、 and Represents the batch size, the number of image tags and the number of channels respectively. The attention layer is implemented by the function To calculate self-attention. In this process, the query (Q), key (K), and value (V) are mapped from the latent features through linear transformations, which can be explained by the following formula.

[0081] (6)

[0082] In order to establish interactions between adjacent images and thus maintain multi-view consistency, adjacent images are randomly sampled in the Batch dimension to generate new keys (K) and values ​​(V), which are then concatenated with the corresponding K and V for attention calculation, which is explained by the following formula.

[0083] , (7)

[0084] (8)

[0085] Step 104: Perform inverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene.

[0086] Exemplarily, the diffusion model includes two stages. In the reverse (denoising) stage, it is necessary to achieve 3D scene reconstruction by removing noise. By correcting the denoising direction in the next step, the denoising direction will shift as the time step increases. To solve this problem, the present application further corrects the denoising direction through the reverse processing of the diffusion model of iterative denoising, thereby enhancing the stability and consistency of the denoising process. It not only ensures the visual consistency between the reconstructed 3D scenes under multiple perspectives, but also ensures that the reconstructed scene is highly semantically consistent with the original scene.

[0087] Specifically, the diffusion model of iterative denoising is performed on the noisy rendered image group to obtain a visually consistent 3D scene, including: inputting the noisy rendered image group into a preset denoising network, and correcting the denoising direction of the denoising network through a 3D Gaussian field to obtain the denoising result of the current time step; based on the denoising results of the noisy rendered image group and the current time step, determining the deviation of the anti-denoising direction from the origin; according to the deviation of the anti-denoising direction from the origin, pulling back through the denoising direction to obtain a visually consistent 3D scene.

[0088] In one embodiment, during the iterative denoising process, this application first adopts a similar setup as in "Single Image to NovelViews with Semantic-Preserving Generative Warping" (hereinafter referred to as VistaDream by the author) to correct the denoising direction by training a 3D Gaussian field. The difference is that this application introduces an additional origin To prevent the denoising direction from deviating from the original image, it aims to ensure that the generated image remains consistent with the original scene during the denoising process, while reducing the visual distortion caused by the deviation of the denoising direction, which is explained by the following formula.

[0089] (9)

[0090] (10)

[0091] (11)

[0092] in , is a constant, is the denoising network;

[0093] and is the preset weight coefficient;

[0094] Used to prevent overexposure.

[0095] VistaDream Pass To correct the denoising direction of the next step, but as the time step increases, once the denoising direction shifts, this shift will gradually accumulate, eventually leading to blurred distortion of the image.

[0096] In order to solve the above problems, this application introduces , which is calculated by the original image and the denoising result of the current time step. If the denoising direction gradually deviates from the original scene, This pulls the denoising direction back, preventing it from deviating from the original scene, significantly reducing image blur. By introducing this new mechanism, we significantly enhance the stability and consistency of the denoising process, ensuring not only visual consistency between reconstructed 3D scenes from multiple perspectives but also a close semantic fit between the reconstructed scenes and the original.

[0097] Step 105: Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object.

[0098] For example, since the reconstruction framework constructed by the present application has higher flexibility, for single-view 3D editing, objects in the 3D scene can be replaced or added without sampling multiple views of the same object, and only a single view of the object needs to be edited.

[0099] Specifically, a visually consistent 3D scene is input into a single-view 3D editing algorithm to determine a 3D scene for replacing the changed object, specifically including: rendering the target to be edited of the visually consistent 3D scene to determine the target to be edited; based on the target to be edited, performing gradient descent optimization on the scene parameters of the target to be edited to obtain scene optimization parameters; wherein the constraints of the gradient descent optimization include: depth loss constraint; according to the scene optimization parameters, performing mask area extraction on the single-view replacement target of the target to be edited to obtain the mask area of ​​the source object, and performing replacement changes on the mask area to determine the 3D scene for replacing the changed object.

[0100] In one embodiment, first, the Grounded DINO and SAM models are used to Extract the masked region of the source object.

[0101] Then, the Flux model is used to fill the area with the hint of the replacement object, thereby achieving object replacement.

[0102] For object addition, a specified mask area and object hint are required in advance. The Flux model supplements the specified area based on the input hint to complete the object addition.

[0103] Directly use the gradient descent method to optimize the parameters of the 3D scene, and use the pixel space in the optimization process. loss and structural similarity (SSIM) loss as basic constraints.

[0104] It should be noted that since editing is limited to a single view, relying solely on the above loss function for optimization may lead to drift. To this end, this application introduces an additional depth loss as a constraint to improve the stability of the optimization. The optimization objective function is explained by the following formula.

[0105] (12)

[0106] (13)

[0107] Where DPT represents the depth prediction model, and the global depth information is realigned through L1 loss to solve the problem that the edited object may interfere with the surrounding depth distribution. The weight coefficient 𝜆 is set to 0.05.

[0108] Figure 2 This is a schematic diagram of object replacement and addition in a 3D scene provided by an embodiment of this application. In a comparison of the experimental verification results of this application, quantitative evaluation results in 11 different scenes achieved significant improvements over existing technologies in five key dimensions: noise level, edge clarity, structural accuracy, detail recovery, and overall visual quality.

[0109] For qualitative analysis, the prediction-refinement cycle method and the offset camera strategy are introduced to fill the scene, which significantly enhances the diversity and continuity of the scene. Figure 1 When the image is consistent, the method adopted in this application can not only maintain the style characteristics of the original image, but also improve the consistency between images, thereby generating high-quality and realistic 3D scenes.

[0110] For object replacement and addition in 3D scenes, this application can accurately replace and add objects in 3D scenes, allowing flexible manipulation of objects while ensuring that the background remains unchanged, thereby maintaining the integrity of the original scene.

[0111] Figure 3 The schematic diagrams of 3D reconstruction using different camera trajectories provided for the embodiments of this application show the 3D scenes reconstructed from various trajectories, highlighting the effectiveness of this application. Even with complex trajectories, this method can consistently produce high-quality and accurate 3D reconstructions, maintaining the coherence and rationality of the scene.

[0112] Figure 4 This is a graph showing an ablation study of a noise addition strategy provided in an embodiment of the present application. The absence of a prediction-refinement loop causes the reconstructed scene to have high repeatability of elements in the original image. This problem becomes particularly evident when reconstructing large-scale and highly complex scenes.

[0113] Furthermore, without an offset camera strategy, the alignment and depth information between different views will be severely mismatched, making it difficult to form a coherent 3D scene with precise spatial relationships. This application effectively addresses these issues, improving the diversity, coherence, and rationality of the reconstructed scenes, ultimately achieving more realistic and visually appealing reconstruction results than existing technologies.

[0114] The introduction of random noise significantly reduces the correlation between views, leading to suboptimal multi-view denoising. Figure 1 The correlated noise strategy provides a more effective starting point, and the integration of correlation-aware attention further enhances the consistency between views, thereby improving the overall quality of reconstruction.

[0115] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a single-view method based on a diffusion model. Figure 3 D scene reconstruction equipment, its structure is as follows Figure 5 shown.

[0116] Figure 5 A single-view method based on a diffusion model is provided in the embodiment of the present application. Figure 3 D Schematic diagram of the internal structure of the scene reconstruction device. Figure 5 As shown, the equipment includes:

[0117] at least one processor 501;

[0118] and, a memory 502 in communication with the at least one processor;

[0119] The memory 502 stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor 501 to enable the at least one processor 501 to:

[0120] The camera trajectory is obtained, and based on the camera trajectory, the set of views to be filled is obtained through offset camera strategy analysis; the blank areas of the view set to be filled are filled to obtain filled views, and the filled views are refined and cyclically optimized to obtain the reconstructed current 3D scene; according to the current 3D scene, a noisy rendered image group is determined through forward processing of a diffusion model with a noise addition strategy; the noisy rendered image group is subjected to reverse processing of a diffusion model with iterative denoising to obtain a visually consistent 3D scene; the visually consistent 3D scene is input into a single-view 3D editing algorithm to determine the 3D scene in which the changed object is replaced.

[0121] Some embodiments of the present application provide corresponding Figure 1 A single-view method based on diffusion model Figure 3 A non-volatile computer storage medium for D scene reconstruction stores computer executable instructions, wherein the computer executable instructions are set to:

[0122] The camera trajectory is obtained, and based on the camera trajectory, the set of views to be filled is obtained through offset camera strategy analysis; the blank areas of the view set to be filled are filled to obtain filled views, and the filled views are refined and cyclically optimized to obtain the reconstructed current 3D scene; according to the current 3D scene, a noisy rendered image group is determined through forward processing of a diffusion model with a noise addition strategy; the noisy rendered image group is subjected to reverse processing of a diffusion model with iterative denoising to obtain a visually consistent 3D scene; the visually consistent 3D scene is input into a single-view 3D editing algorithm to determine the 3D scene in which the changed object is replaced.

[0123] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the IoT device and media embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.

[0124] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.

[0125] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0127] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0129] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0130] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0131] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0132] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0133] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A single-view 3D scene reconstruction method based on a diffusion model, characterized in that: The method comprises: Obtaining a camera trajectory, and based on the camera trajectory, obtaining a view set to be filled by analyzing the offset camera strategy; Filling blank areas of the to-be-filled view set to obtain a filled view, and performing refinement and loop optimization on the filled view to obtain a reconstructed current 3D scene; Determining a noisy rendered image group based on the current 3D scene through forward processing of a diffusion model using a noise addition strategy; performing an iterative denoising diffusion model inverse process on the noisy rendered image group to obtain a visually consistent 3D scene; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object; Based on the camera trajectory, the view set to be filled is obtained through the offset camera strategy analysis, which specifically includes: Based on the camera trajectory, an updated extrinsic parameter is obtained by controlling the camera offset angle; According to the updated external parameters, the updated internal parameters are obtained by adjusting the rendering view; Determining an offset camera based on the updated extrinsic parameters and the updated intrinsic parameters, and obtaining the to-be-filled view set by sampling the 3D scene according to the offset camera; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object, specifically comprising: Rendering the target to be edited in the visually consistent 3D scene to determine the target to be edited; Based on the target to be edited, the scene parameters of the target to be edited are optimized by gradient descent method to obtain scene optimization parameters; wherein the constraint conditions of the gradient descent method optimization include: depth loss constraint; Extracting a mask region of a single-view replacement target of the target to be edited according to the scene optimization parameters to obtain a mask region of a source object, and performing a replacement change on the mask region to determine a 3D scene of the replacement target; Filling the blank area of ​​the to-be-filled view set to obtain a filled view, specifically comprising: Acquiring MLLM extraction parameters, and performing filling area prediction on the to-be-filled view set based on the MLLM extraction parameters to obtain a to-be-filled area object; wherein the MLLM extraction parameters include: global hints and local hints; The filling view is obtained by performing a filling area mask threshold analysis based on the area object to be filled.

2. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, characterized in that: The filled view is subjected to a refinement loop optimization to obtain a reconstructed current 3D scene, specifically comprising: Performing refinement loop optimization based on the filled view and obtaining view evaluation data through MLLM consistency evaluation; wherein the MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, and view style consistency evaluation; The view evaluation data is quality screened to determine a current optimal view, and the reconstructed current 3D scene is obtained by cyclically updating the MLLM descriptive prompt according to the current optimal view.

3. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, characterized in that: According to the current 3D scene, a noisy rendered image group is determined by forward processing of a diffusion model using a noise addition strategy, specifically including: Based on the current 3D scene, a rendering image group is generated by image rendering; performing cosine similarity analysis of adjacent image influence intensity control on adjacent images of the rendered image group to obtain cosine similarity of adjacent images; The noisy rendered image group is determined by adding random noise based on view consistency according to the cosine similarity of the adjacent images.

4. The single-view 3D scene reconstruction method based on a diffusion model according to claim 3, characterized in that: After determining the noisy rendered image group by adding random noise for view consistency according to the cosine similarity of the adjacent images, the method further includes: Determining an attention parameter based on the noisy rendered image group by analyzing image potential features; According to the attention parameter, randomly sampling the batch dimension to obtain an update component, wherein the update component includes: a key and a value; The update component is concatenated with the corresponding key and value, and attention analysis is performed on the concatenated update component to determine the relevance-aware attention mechanism.

5. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, characterized in that: Performing an iterative denoising diffusion model inverse process on the noisy rendered image group to obtain a visually consistent 3D scene, specifically comprising: Inputting the noisy rendered image group into a preset denoising network, and correcting the denoising direction of the denoising network through a 3D Gaussian field to obtain a denoising result for the current time step; Determining, based on the noisy rendered image group and the denoising result of the current time step, whether the anti-denoising direction deviates from the origin; According to the deviation of the anti-denoising direction from the origin, the denoising direction is pulled back to obtain the visually consistent 3D scene.

6. A single-view 3D scene reconstruction device based on a diffusion model, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtaining a camera trajectory, and based on the camera trajectory, obtaining a view set to be filled by analyzing the offset camera strategy; Filling blank areas of the to-be-filled view set to obtain a filled view, and performing refinement and loop optimization on the filled view to obtain a reconstructed current 3D scene; Determining a noisy rendered image group based on the current 3D scene through forward processing of a diffusion model using a noise addition strategy; performing an iterative denoising diffusion model inverse process on the noisy rendered image group to obtain a visually consistent 3D scene; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object; Based on the camera trajectory, the view set to be filled is obtained through the offset camera strategy analysis, which specifically includes: Based on the camera trajectory, an updated extrinsic parameter is obtained by controlling the camera offset angle; According to the updated external parameters, the updated internal parameters are obtained by adjusting the rendering view; Determining an offset camera based on the updated extrinsic parameters and the updated intrinsic parameters, and obtaining the to-be-filled view set by sampling the 3D scene according to the offset camera; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object, specifically comprising: Rendering the target to be edited in the visually consistent 3D scene to determine the target to be edited; Based on the target to be edited, the scene parameters of the target to be edited are optimized by gradient descent method to obtain scene optimization parameters; wherein the constraint conditions of the gradient descent method optimization include: depth loss constraint; Extracting a mask region of a single-view replacement target of the target to be edited according to the scene optimization parameters to obtain a mask region of a source object, and performing a replacement change on the mask region to determine a 3D scene of the replacement target; Filling the blank area of ​​the to-be-filled view set to obtain a filled view, specifically comprising: Acquiring MLLM extraction parameters, and performing filling area prediction on the to-be-filled view set based on the MLLM extraction parameters to obtain a to-be-filled area object; wherein the MLLM extraction parameters include: global hints and local hints; The filling view is obtained by performing a filling area mask threshold analysis based on the area object to be filled.

7. A non-volatile computer storage medium storing computer executable instructions for single-view 3D scene reconstruction based on a diffusion model, characterized in that: The computer executable instructions are configured to: Obtaining a camera trajectory, and based on the camera trajectory, obtaining a view set to be filled by analyzing the offset camera strategy; Filling blank areas of the to-be-filled view set to obtain a filled view, and performing refinement and loop optimization on the filled view to obtain a reconstructed current 3D scene; Determining a noisy rendered image group based on the current 3D scene through forward processing of a diffusion model using a noise addition strategy; performing an iterative denoising diffusion model inverse process on the noisy rendered image group to obtain a visually consistent 3D scene; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object; Based on the camera trajectory, the view set to be filled is obtained through the offset camera strategy analysis, which specifically includes: Based on the camera trajectory, an updated extrinsic parameter is obtained by controlling the camera offset angle; According to the updated external parameters, the updated internal parameters are obtained by adjusting the rendering view; Determining an offset camera based on the updated extrinsic parameters and the updated intrinsic parameters, and obtaining the to-be-filled view set by sampling the 3D scene according to the offset camera; Inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene that replaces the changed object, specifically comprising: Rendering the target to be edited in the visually consistent 3D scene to determine the target to be edited; Based on the target to be edited, the scene parameters of the target to be edited are optimized by gradient descent method to obtain scene optimization parameters; wherein the constraint conditions of the gradient descent method optimization include: depth loss constraint; Extracting a mask region of a single-view replacement target of the target to be edited according to the scene optimization parameters to obtain a mask region of a source object, and performing a replacement change on the mask region to determine a 3D scene of the replacement target; Filling the blank area of ​​the to-be-filled view set to obtain a filled view, specifically comprising: Acquiring MLLM extraction parameters, and performing filling area prediction on the to-be-filled view set based on the MLLM extraction parameters to obtain a to-be-filled area object; wherein the MLLM extraction parameters include: global hints and local hints; The filling view is obtained by performing a filling area mask threshold analysis based on the area object to be filled.

Citation Information

Patent Citations

  • Single-view three-dimensional modeling method and system based on diffusion model

    CN119068144A

  • Three-dimensional human body reconstruction method based on implicit neural network and diffusion model

    CN119991967A