Single-view 3D scene reconstruction method and device based on diffusion model, and medium
Through the combination of diffusion model and offset camera strategy, the problem of view consistency and clarity in three-dimensional scene reconstruction is solved, and a high-quality 3D scene with diversity and consistency is generated, supporting object replacement and addition.
Patent Information
- Application Number
- CN202510747860.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
In the prior art, three-dimensional scene reconstruction methods are difficult to ensure the consistency and clarity of multi-views while ensuring the diversity and rationality of new perspectives.
A single-view 3D scene reconstruction method based on diffusion model is adopted to generate a visually consistent 3D scene through offset camera strategy analysis, prediction-refinement cycle optimization, noise addition strategy and correlation perceptual attention, combined with depth loss constraints.
Improves the diversity, coherence and consistency of 3D scene reconstruction, and generates high-quality and realistic 3D scenes, supporting object replacement and addition.
Smart Images

Figure CN120259571A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional scene reconstruction technology, and in particular to a single-view Figure 3 3D scene reconstruction method, device and medium based on a diffusion model. Background Technique
[0002] Three-dimensional scene reconstruction is a key task in fields such as augmented reality (AR), virtual reality (VR), and computer graphics. With the development of 3DGS technology, three-dimensional reconstruction has been significantly promoted, enabling efficient use of rich data for scene reconstruction. However, the reconstruction quality highly depends on large-scale and high-precision three-dimensional data, but the process of collecting these data is time-consuming and costly. Therefore, single-view based three-dimensional reconstruction methods have emerged, aiming to overcome the limitations of data collection and improve the reconstruction efficiency.
[0003] However, the information contained in a single view is limited and it is difficult to comprehensively present the hidden structure of an object. In recent years, as a generative technology, the diffusion model has achieved remarkable results in image generation and completion tasks, and can generate potential perspectives based on a single view, effectively making up for the lack of single-view information. These methods still face challenges in large-scale three-dimensional reconstruction, especially in terms of multi-view content repetition, unreasonableness, and inconsistency. Previous methods such as "Single Image to Novel Views withSemantic-Preserving Generative Warping", "Text-driven 3d scene generation withinpainting and depth diffusion. In International Confer ence on 3D Vision(3DV)" focus on the consistency between a single view and newly generated views Figure 1 but ignore the global consistency; recent methods such as "VistaDream: Sampling multiview consistent images for single-view scenereconstruction" can ensure multi-view Figure 1 consistency, but will cause blurring phenomena and reduce the reconstruction quality. Therefore, how to ensure multi-view Figure 1 consistency and clarity while ensuring the diversity and reasonableness of new perspectives is the core issue in three-dimensional scene reconstruction. Summary of the Invention
[0004] The embodiments of this application provide a single-view Figure 3 3D scene reconstruction method, device and medium based on a diffusion model, which solves the problem of the view in three-dimensional scene reconstruction in the prior art Figure 1The technical problem that consistency and clarity cannot be satisfied simultaneously.
[0005] In a first aspect, an embodiment of the present application provides a single-view Figure 3 3D scene reconstruction method based on a diffusion model, characterized in that the method includes: obtaining a camera trajectory, and based on the camera trajectory, through offset camera strategy analysis, obtaining a set of views to be filled; filling blank areas in the set of views to be filled to obtain filled views, and performing refined cyclic optimization on the filled views to obtain the reconstructed current 3D scene; according to the current 3D scene, through forward processing of the diffusion model with a noise addition strategy, determining a set of noisy rendered images; performing reverse processing of the diffusion model for iterative denoising on the set of noisy rendered images to obtain a visually consistent 3D scene; inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object.
[0006] In an implementation manner of the present application, based on the camera trajectory, through offset camera strategy analysis, obtaining a set of views to be filled specifically includes: based on the camera trajectory, through camera offset angle control, obtaining updated external parameters; according to the updated external parameters, through rendering view adjustment, obtaining updated internal parameters; based on the updated external parameters and the updated internal parameters, determining an offset camera, and according to the offset camera, through 3D scene sampling, obtaining a set of views to be filled.
[0007] In an implementation manner of the present application, filling blank areas in the set of views to be filled to obtain filled views specifically includes: obtaining MLLM extraction parameters, and based on the MLLM extraction parameters, predicting filling areas in the set of views to be filled to obtain filling area objects; where the MLLM extraction parameters include: global prompts, local prompts; according to the filling area objects, through filling area mask threshold analysis, obtaining filled views.
[0008] In an implementation manner of the present application, performing refined cyclic optimization on the filled views to obtain the reconstructed current 3D scene specifically includes: performing refined cyclic optimization based on the filled views, through MLLM consistency evaluation, obtaining view evaluation data; where the MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, view style consistency evaluation; performing quality screening on the view evaluation data to determine the current optimal view, and according to the current optimal view, through MLLM descriptive prompt cyclic update, obtaining the reconstructed current 3D scene.
[0009] In one implementation of the present application, according to the current 3D scene, a noisy rendered image group is determined by forward processing of a diffusion model of a noise adding strategy, specifically comprising: based on the current 3D scene, generating a rendered image group by image rendering; performing cosine similarity analysis of adjacent images of the rendered image group with adjacent image influence intensity control to obtain cosine similarity of adjacent images; according to the cosine similarity of adjacent images, Figure 1 Consistent random noise is added to determine the group of noisy rendered images.
[0010] In one implementation of the present application, according to the cosine similarity of adjacent images, Figure 1 After adding consistent random noise to determine the noisy rendered image group, the method also includes: determining attention parameters based on the noisy rendered image group through image potential feature analysis; according to the attention parameters, obtaining update components through random sampling of the Batch dimension; wherein the update components include: keys and values; splicing the update components with the corresponding keys and values, and performing attention analysis on the spliced update components to determine the correlation-aware attention mechanism.
[0011] In one implementation of the present application, a diffusion model of iterative denoising is performed on a noisy rendered image group inversely to obtain a visually consistent 3D scene, specifically including: inputting the noisy rendered image group into a preset denoising network, and correcting the denoising direction of the denoising network through a 3D Gaussian field to obtain a denoising result of the current time step; based on the denoising result of the noisy rendered image group and the current time step, determining the deviation of the anti-denoising direction from the origin; according to the deviation of the anti-denoising direction from the origin, pulling back through the denoising direction to obtain a visually consistent 3D scene.
[0012] In one implementation of the present application, a visually consistent 3D scene is input into a single-view 3D editing algorithm to determine a 3D scene for replacing a changed object, specifically comprising: rendering a target to be edited of the visually consistent 3D scene to determine the target to be edited; based on the target to be edited, optimizing the scene parameters of the target to be edited by gradient descent method to obtain scene optimization parameters; wherein the constraints of the gradient descent method optimization include: a depth loss constraint; based on the scene optimization parameters, extracting a mask area of a single-view replacement target of the target to be edited to obtain a mask area of a source object, and performing a replacement change on the mask area to determine the 3D scene for replacing the changed object.
[0013] In a second aspect, the present application also provides a single-view method based on a diffusion model. Figure 3D scene reconstruction device, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: obtain a camera trajectory, and based on the camera trajectory, obtain a view set to be filled through offset camera strategy analysis; fill the blank areas of the view set to be filled to obtain a filled view, and perform refined loop optimization on the filled view to obtain the reconstructed current 3D scene; according to the current 3D scene, determine a group of noisy rendering images through forward processing of the diffusion model with a noise addition strategy; perform reverse processing of the diffusion model for iterative denoising on the group of noisy rendering images to obtain a visually consistent 3D scene; input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene for replacing changed objects.
[0014] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for single-view Figure 3 D scene reconstruction based on a diffusion model, storing computer-executable instructions, characterized in that the computer-executable instructions are set to: obtain a camera trajectory, and based on the camera trajectory, obtain a view set to be filled through offset camera strategy analysis; fill the blank areas of the view set to be filled to obtain a filled view, and perform refined loop optimization on the filled view to obtain the reconstructed current 3D scene; according to the current 3D scene, determine a group of noisy rendering images through forward processing of the diffusion model with a noise addition strategy; perform reverse processing of the diffusion model for iterative denoising on the group of noisy rendering images to obtain a visually consistent 3D scene; input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene for replacing changed objects.
[0015] An embodiment of the present application provides a single-view Figure 3 D scene reconstruction method, device and medium. A rough three-dimensional scene is generated by expanding a single view through a two-dimensional diffusion model, and combined with a prediction-refinement loop method, the scene diversity and consistency are improved through dynamic filling of the content of adjacent views and the original view. At the same time, an offset camera strategy is proposed to deflect the current view pose and use global prompt words for filling, and a coherent three-dimensional scene is gradually generated through iteration. To further improve multi-view Figure 1 consistency, this paper designs correlation-aware attention and related noise addition strategies to enhance the correlation of adjacent views, and corrects the denoising direction by introducing origin constraints to avoid scene blurring. In terms of scene editing, based on object replacement and addition in a single view, by extracting masks and filling them, combined with depth loss constraints to ensure the accuracy and consistency of the edited objects, the technical problem that the visual Figure 1 consistency and clarity of three-dimensional scene reconstruction in the prior art cannot be satisfied at the same time is solved. Description of the Drawings
[0016] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 A single-view Figure 3 D scene reconstruction method flowchart provided by an embodiment of the present application; Figure 2 A schematic diagram of object replacement and addition in a 3D scene provided by an embodiment of the present application; Figure 3 A 3D reconstruction schematic diagram of different camera trajectories provided by an embodiment of the present application, Figure 3 in which (a) is a 3D reconstruction schematic diagram of the camera trajectory moving backward, and (b) is a 3D reconstruction schematic diagram of the camera trajectory moving forward; Figure 4 A curve graph of ablation study on noise addition strategy provided by an embodiment of the present application; Figure 5 A single-view Figure 3 D scene device internal structure schematic diagram provided by an embodiment of the present application. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0018] An embodiment of the present application provides a single-view Figure 3 D scene reconstruction method, device, and medium. By using a two-dimensional diffusion model to expand a single view to generate a rough three-dimensional scene, and combining a prediction-refinement loop method, dynamic filling is performed through the content of adjacent views and the original view to improve scene diversity and consistency. At the same time, an offset camera strategy is proposed to deflect the current view pose and fill it with global prompt words, and a coherent three-dimensional scene is gradually iteratively generated. To further improve multi-view Figure 1 consistency, a correlation-aware attention and a correlation noise addition strategy are designed in this paper to enhance the correlation of adjacent views, and the denoising direction is corrected by introducing an origin constraint to avoid scene blurring. In terms of scene editing, the present application also supports object replacement and addition based on a single view. By extracting masks and filling them, and combining depth loss constraints to ensure the accuracy and consistency of the edited objects, the problem of three-dimensional scene reconstruction vision in the prior art is solved. Figure 1The technical problem that consistency and clarity cannot be satisfied simultaneously.
[0019] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 A single-view Figure 3 D scene reconstruction method flowchart provided for the embodiments of the present application. As Figure 1 shown, a single-view Figure 3 D scene reconstruction method provided for the embodiments of the present application specifically includes the following steps: Step 101, obtain the camera trajectory, and based on the camera trajectory, through the analysis of the offset camera strategy, obtain the view set to be filled.
[0021] Exemplarily, if all views sampled from the camera trajectory are filled with predicted objects, although a certain degree of diversity is introduced, it is likely to lead to scene discontinuity and may even deviate from the original scene. Moreover, the blank areas of adjacent views usually account for 50%, and the filling objects may be inconsistent, resulting in incomplete object generation and further causing problems such as visual distortion. Through the analysis of the offset camera strategy, the filling results of the view set to be filled are made complete, and the consistency of the filling objects is improved.
[0022] Specifically, based on the camera trajectory, through the analysis of the offset camera strategy, obtaining the view set to be filled includes: based on the camera trajectory, through the control of the camera offset angle, obtaining the updated external parameters; according to the updated external parameters, through the adjustment of the rendered view, obtaining the updated internal parameters; based on the updated external parameters and updated internal parameters, determining the offset camera, and according to the offset camera, through 3D scene sampling, obtaining the view set to be filled.
[0023] In one embodiment, if all views sampled from the camera trajectory are filled with predicted objects, although a certain degree of diversity is introduced, it is likely to lead to scene discontinuity and may even deviate from the original scene. Moreover, the blank areas of adjacent views usually account for 50%, and the filling objects may be inconsistent, resulting in incomplete object generation and further causing problems such as visual distortion. To effectively enhance scene coherence and ensure the complete generation of objects, we propose an offset camera strategy. The preset camera trajectory is where the camera includes intrinsic and extrinsic parameters and can render new views . For simplicity, the height and width in the intrinsic parameters of the camera are used to change the width and height of the rendered view, and the camera position and the target position control the camera The external parameters. Based on these camera parameters, we construct an offset camera by adjusting or parameters and introducing an offset angle. The process of constructing the above offset camera is explained by the following formula: (1) Wherein, , , are predefined deflection angles; represents rotation around the axis, represents rotation around the axis, represents rotation around the axis; , , represent unit vectors; defines quaternion multiplication, is the conjugate; wherein, the quaternion is and .
[0024] By controlling the deflection angle, the new position coordinates can be calculated and combined with the camera position pos to form new external parameters. At the same time, by adjusting the height and width parameters of the rendering view, the internal parameters of the camera are changed, thereby generating a deflected camera.
[0025] Two sets of parameters and are used to construct two deflected cameras, denoted as and ; wherein, is a scaling factor used to control the zoom in and out of the rendering view.
[0026] For , its deflection angle is generally set to negative to capture more known scene information. When filling the view rendered by the deflected camera, a global prompt is used to ensure the coherence of the scene.
[0027] It should be noted that 's deflection angle is generally set to positive, and larger and values are used to complete the predicted object while ensuring that the object is reasonably integrated into the current scene.
[0028] Further, after determining the two offsets, sample the 3D framework through the camera trajectory to obtain a set of views to be filled. .
[0029] Step 102: Fill the blank areas of the set of views to be filled to obtain filled views, and perform refined cyclic optimization on the filled views to obtain the reconstructed current 3D scene.
[0030] Exemplarily, in the prior art, a multi-modal large language model (MLLM) is fully utilized to understand the original views, but the same prompts are used during iterative filling, resulting in a problem of scene singularity when generating complex perspectives (such as 360-degree surround views). To prevent the problem of scene singularity and enhance the potential of the MLLM, the content limitation is broken through by leveraging the original view information, and a self-optimizing framework is constructed to enhance the diversity of the scene.
[0031] Specifically, filling the blank areas of the set of views to be filled to obtain filled views includes: obtaining MLLM extraction parameters, and based on the MLLM extraction parameters, predicting the filling areas of the set of views to be filled to obtain objects of areas to be filled; wherein, the MLLM extraction parameters include: global prompts, local prompts; according to the objects of areas to be filled, obtain filled views through filling area mask threshold analysis.
[0032] Further, performing refined cyclic optimization on the filled views to obtain the reconstructed current 3D scene specifically includes: performing refined cyclic optimization based on the filled views, and obtaining view evaluation data through MLLM consistency evaluation; wherein, the MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, view style consistency evaluation; performing quality screening on the view evaluation data to determine the current optimal view, and based on the current optimal view, obtaining the reconstructed current 3D scene through cyclic update of MLLM descriptive prompts.
[0033] Exemplarily, for a given original view, first use the MLLM to extract global prompts and local prompts; wherein, the global prompts include the overall information, style, etc. of the scene, while the local prompts are detailed descriptions of the original view.
[0034] Through the local prompts, expand and fill the original view to generate a preliminary 3D framework. Then, sample the 3D framework through the camera trajectory to obtain a set of views, and use a diffusion model to fill the blank areas in each view.
[0035] When filling each view we combine the global prompts and the information of adjacent views to predict the Possible objects. The mask of the filled area is usually set to about 50% to meet the generation requirements of most objects. When the filled area is less than 10%, only the global hint is used for filling.
[0036] After generating the complete view, we use MLLM to evaluate its consistency with the original view and the global hint in terms of quality, content, and style. Finally, the view evaluation data is screened for quality, and the optimal view is selected. Using MLLM, more descriptive hints are generated to further optimize , to generate a higher-quality view.
[0037] Through the above user-friendly self-improving process, this application can gradually generate images with high quality and diverse content and styles, and reconstruct the entire 3D scene through continuous cycling.
[0038] Furthermore, the specific algorithm logic of the 3D reconstruction method that combines offset camera strategy analysis with fusion prediction-refinement loop optimization is shown as follows.
[0039] Input: Filling model , Number of iterations , Set of camera trajectories , Global hint , Deflection angle , Scaling factor ; Output: 3D scene.
[0040] Loop : Select from the set of camera trajectories , ; Equal to the offset camera strategy ; Equal to the offset camera strategy ; Through Sample a view from the 3D scene ; ; Use To update the 3D scene; Prediction-refinement loop method; Through Sample a view from the 3D scene ; ; Use To update the 3D scene; End; Return to the 3D scene.
[0041] Step 103: According to the current 3D scene, perform forward processing through a diffusion model with a noise addition strategy to determine a group of noisy rendered images.
[0042] Exemplarily, there are still some defects in the 3D scene based on the offset camera strategy analysis and fusion prediction - refinement loop optimization reconstruction. It is mainly manifested that the camera trajectory cannot cover all perspectives, resulting in slight holes in some rendering results, and there are problems of multi - view inconsistency in the images. When sampling along the predefined camera trajectory, if the trajectory coverage is small, there is usually a strong correlation between views. However, as the trajectory sampling method becomes more complex (such as moving straight forward, backward, or rotating 360°), the nth view only has a strong correlation with the adjacent views before and after. To enhance the correlation between views, the 3D scene is processed through a diffusion model for multi - view Figure 1 consistency to solve the above problems and improve the consistency and fidelity of 3D scene reconstruction; among them, the diffusion model includes two stages. In the forward (positive) process, noise needs to be added to the image.
[0043] Specifically, according to the current 3D scene, performing forward processing through a diffusion model with a noise addition strategy to determine a group of noisy rendered images includes: based on the current 3D scene, generating through image rendering to obtain a group of rendered images; performing cosine similarity analysis on adjacent images of the group of rendered images to control the influence intensity of adjacent images to obtain the cosine similarity of adjacent images; according to the cosine similarity of adjacent images, determining a group of noisy rendered images through Figure 1 multi - view consistency random noise addition.
[0044] In one embodiment, first render a group of images from the 3D scene according to the camera trajectory, and then apply a 2D diffusion model to enforce the consistency constraint between these views. Finally, use the improved views to update the 3D scene, thereby improving the consistency and fidelity of the reconstruction.
[0045] In the forward process of the diffusion model, given a rendered image These images need to be gradually added noise. However, directly introducing random noise may destroy the multi - view Figure 1 consistency between images, making it difficult for the denoising network to establish the consistency between views within the specified time step. Therefore, noise addition is performed through a related noise addition strategy, and the noise addition strategy is explained in detail by the following formula.
[0046] (2) (3) (4) (5) Wherein, is a predefined random noise constant; is independent random noise, is the cosine similarity of adjacent images; is used to control the influence intensity of the current image by adjacent images and is also a constant.
[0047] When , the strategy degenerates to introducing independent random noise to all images; set to 0.2 to appropriately introduce randomness while maintaining consistency between views. The main goal of this strategy is to provide better initial conditions for the reverse generation process of the diffusion model, so as to achieve multi-view Figure 1 consistency construction in fewer time steps.
[0048] Further, after determining the group of noisy rendered images according to the cosine similarity of adjacent images through view Figure 1 consistency random noise addition, the method further includes: based on the group of noisy rendered images, determining attention parameters through image latent feature analysis; according to the attention parameters, obtaining updated components through random sampling in the Batch dimension; wherein the updated components include: keys, values; splicing the updated components with the corresponding keys and values, and performing attention analysis on the spliced updated components to determine the correlation-aware attention mechanism.
[0049] In one embodiment, based on the characteristics of video consistency processing, the present application improves a correlation-aware attention to replace self-attention, enhancing the correlation between the current image and its adjacent images.
[0050] Formally, given an image latent feature , where , and represent the batch size, the number of image tokens, and the number of channels respectively. The attention layer calculates self-attention through the function . In this process, the query (Q), key (K), and value (V) are all mapped from the latent feature through linear transformation, which is explained by the following formula.
[0051] (6) To establish interactions between adjacent images to maintain multi-view consistency, random sampling is performed on its adjacent images in the Batch dimension to generate new keys (K) and values (V), and they are spliced with the corresponding K and V for attention calculation, which is explained by the following formula.
[0052] , (7) (8) Step 104: Perform reverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene.
[0053] Exemplarily, the diffusion model includes two stages. In the reverse (denoising) stage, 3D scene reconstruction needs to be achieved by removing noise. By correcting the denoising direction of the next step, the denoising direction deviation will occur as the time step increases. To solve this problem, the present application performs reverse processing of the diffusion model for iterative denoising to further correct the denoising direction, enhancing the stability and consistency of the denoising process. It not only ensures the visual consistency between the reconstructed 3D scenes from multiple perspectives but also ensures that the reconstructed scene is highly consistent with the original scene semantically.
[0054] Specifically, performing reverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene includes: inputting the noisy rendered image group into a preset denoising network, and correcting the denoising direction of the denoising network through a 3D Gaussian field to obtain the denoising result at the current time step; determining the prevention of the denoising direction from deviating from the origin based on the noisy rendered image group and the denoising result at the current time step; and obtaining a visually consistent 3D scene by pulling back the denoising direction according to the prevention of the denoising direction from deviating from the origin.
[0055] In one embodiment, during the iterative denoising process, the present application first adopts a setting similar to that in "Single Image to NovelViews with Semantic-Preserving Generative Warping" (hereinafter abbreviated as the author VistaDream) to correct the denoising direction by training a 3D Gaussian field. Different from this, the present application additionally introduces an origin to prevent the denoising direction from deviating from the original image, aiming to ensure that the generated image maintains consistency with the original scene during the denoising process and reduce visual distortion caused by the deviation of the denoising direction, which is explained by the following formula.
[0056] (9) (10) (11) where , is a constant, is the denoising network; and is a preset weight coefficient; for preventing overexposure.
[0057] VistaDream corrects the denoising direction of the next step through However, as the time step increases, once the denoising direction deviates, this deviation will gradually accumulate, ultimately leading to blurred distortion of the image.
[0058] To solve the above problems, the present application introduces , which is obtained by calculating the original image and the denoising result of the current time step. If the denoising direction gradually deviates from the original scene, will play a role in pulling it back, preventing the denoising direction from deviating from the original scene, thereby significantly reducing the blurring phenomenon of the image. By introducing this new mechanism, we have significantly enhanced the stability and consistency of the denoising process, not only ensuring the visual consistency between the reconstructed 3D scenes under multiple viewpoints, but also ensuring that the reconstructed scenes are highly consistent with the original scene semantically.
[0059] Step 105: Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object.
[0060] Exemplarily, since the reconstruction framework constructed in the present application has higher flexibility, for single-view 3D editing, it can replace or add objects in the 3D scene without sampling multiple views of the same object, and only need to edit a single view of the object.
[0061] Specifically, inputting the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object specifically includes: rendering the target to be edited in the visually consistent 3D scene to determine the target to be edited; based on the target to be edited, optimizing the scene parameters of the target to be edited by the gradient descent method to obtain scene optimization parameters; wherein, the constraint conditions for the gradient descent method optimization include: depth loss constraint; according to the scene optimization parameters, extracting the mask region of the single-view replacement target of the target to be edited to obtain the mask region of the source object, and performing replacement and change on the mask region to determine the 3D scene of the replaced and changed object.
[0062] In one embodiment, first, extract the mask region of the source object through the Grounded DINO and SAM models on .
[0063] Then, use the Flux model to fill this region in combination with the prompt of the replacement object, thereby realizing the replacement of the object.
[0064] For object addition operations, a specified mask area and object hints need to be provided in advance. The Flux model supplements the specified area based on the input hints to complete object addition.
[0065] The gradient descent method is directly used to optimize the parameters of the 3D scene. During the optimization process, the loss in pixel space and the structural similarity (SSIM) loss are used as basic constraints.
[0066] It should be noted that since the editing is limited to a single view, simply relying on the above loss function for optimization may lead to drift problems. Therefore, this application additionally introduces depth loss as a constraint to improve the stability of optimization. The optimization objective function is explained by the following formula.
[0067] (12) (13) where DPT represents the depth prediction model, and the global depth information is realigned through the L1 loss to solve the problem that the edited object may interfere with the surrounding depth distribution. The weight coefficient 𝜆 is set to 0.05.
[0068] Figure 2 This is a schematic diagram of object replacement and addition in a 3D scene provided by an embodiment of this application. In the comparison of the experimental verification results of this application, the quantitative evaluation results in 11 different scenarios show significant improvements compared with the prior art in five key dimensions (including noise level, edge sharpness, structural accuracy, detail recovery, and overall visual quality).
[0069] For qualitative analysis, by introducing a prediction-refinement loop method and an offset camera strategy to fill the scene, the diversity and continuity of the scene are significantly enhanced. When constructing multi-view Figure 1 consistency, the method adopted in this application can not only maintain the style characteristics of the original image but also improve the consistency between images, thus generating high-quality and realistic 3D scenes.
[0070] For object replacement and addition in a 3D scene, this application can accurately replace and add objects in the 3D scene, allowing flexible manipulation of objects while ensuring that the background remains unchanged, thus maintaining the integrity of the original scene.
[0071] Figure 3 This is a schematic diagram of 3D reconstruction with different camera trajectories provided by an embodiment of this application, showing 3D scenes reconstructed from various trajectories, highlighting the effectiveness of this application. Even for complex trajectories, this method can stably generate high-quality and accurate 3D reconstructions, always maintaining the coherence and rationality of the scene.
[0072] Figure 4 A curve graph for ablation study of a noise addition strategy provided by an embodiment of the present application. The absence of the prediction-refinement loop results in a high repeatability of elements in the original image in the reconstructed scene, and this problem becomes particularly obvious when reconstructing large-scale and highly complex scenes.
[0073] In addition, without the offset camera strategy, the alignment and depth information between different views will be severely mismatched, making it difficult to form a coherent 3D scene with precise spatial relationships. The present application effectively solves the above problems, improves the diversity, coherence, and rationality of the reconstructed scene, and finally achieves a more realistic and visually appealing reconstruction result compared with the prior art.
[0074] The introduction of random noise significantly reduces the correlation between views, resulting in suboptimal multi-view Figure 1 consistency during the denoising process. The correlated noise strategy provides a more effective starting point, and the integration of correlation-aware attention further enhances the inter-view consistency, thereby improving the overall quality of the reconstruction.
[0075] The above is the method embodiment proposed by the present application. Based on the same inventive concept, the embodiment of the present application also provides a single-view Figure 3 3D scene reconstruction device based on a diffusion model, and its structure is as Figure 5 shown.
[0076] Figure 5 A schematic diagram of the internal structure of a single-view Figure 3 3D scene reconstruction device based on a diffusion model provided by an embodiment of the present application. As Figure 5 shown, the device includes: At least one processor 501; And a memory 502 communicatively connected to the at least one processor; Wherein, the memory 502 stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor 501 so that the at least one processor 501 can: Obtain the camera trajectory, and based on the camera trajectory, through the analysis of the offset camera strategy, obtain the set of views to be filled; fill the blank areas of the set of views to be filled to obtain filled views, and perform refinement loop optimization on the filled views to obtain the reconstructed current 3D scene; according to the current 3D scene, through the forward processing of the diffusion model with the noise addition strategy, determine the group of noisy rendering images; perform the reverse processing of the diffusion model for iterative denoising on the group of noisy rendering images to obtain a visually consistent 3D scene; input the visually consistent 3D scene into the 3D editing algorithm of a single view to determine the 3D scene for replacing the changed object.
[0077] Corresponding to some embodiments of the present applicationFigure 1 A non-volatile computer storage medium for single-view Figure 3 3D scene reconstruction based on a diffusion model, storing computer-executable instructions, and the computer-executable instructions are set to: Obtain the camera trajectory, and based on the camera trajectory, through the analysis of the offset camera strategy, obtain the view set to be filled; fill the blank areas of the view set to be filled to obtain the filled view, and perform refined loop optimization on the filled view to obtain the reconstructed current 3D scene; according to the current 3D scene, through the forward processing of the diffusion model with the noise addition strategy, determine the noisy rendering image group; perform the reverse processing of the diffusion model for iterative noise reduction on the noisy rendering image group to obtain the visually consistent 3D scene; input the visually consistent 3D scene into the 3D editing algorithm of a single view to determine the 3D scene for replacing the changed object.
[0078] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0079] The systems and media provided by the embodiments of this application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.
[0080] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0081] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.
[0082] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.
[0084] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0085] Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.
[0086] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0087] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0088] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A single-view 3D scene reconstruction method based on a diffusion model, characterized in that The method includes: Obtain a camera trajectory, and based on the camera trajectory, obtain a view set to be filled through offset camera strategy analysis; Fill the blank areas of the view set to be filled to obtain a filled view, and perform refined loop optimization on the filled view to obtain the reconstructed current 3D scene; According to the current 3D scene, determine a group of noisy rendering images through forward processing of a diffusion model with a noise addition strategy; Perform reverse processing of the diffusion model for iterative denoising on the group of noisy rendering images to obtain a visually consistent 3D scene; Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine a 3D scene for replacing changed objects.
2. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, wherein Based on the camera trajectory, obtain a view set to be filled through offset camera strategy analysis, specifically including: Based on the camera trajectory, obtain updated external parameters through camera offset angle control; According to the updated external parameters, obtain updated internal parameters through rendering view adjustment; Based on the updated external parameters and the updated internal parameters, determine an offset camera, and based on the offset camera, obtain the view set to be filled through 3D scene sampling.
3. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, wherein Fill the blank areas of the view set to be filled to obtain a filled view, specifically including: Obtain MLLM extraction parameters, and based on the MLLM extraction parameters, predict filling regions for the view set to be filled to obtain filling region objects; wherein the MLLM extraction parameters include: global prompts, local prompts; According to the filling region objects, obtain the filled view through filling region mask threshold analysis.
4. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, wherein, Perform refined loop optimization on the filled view to obtain the reconstructed current 3D scene, specifically including: Perform refined loop optimization based on the filled view, and obtain view evaluation data through MLLM consistency evaluation; wherein the MLLM consistency evaluation includes: global prompt quality consistency evaluation, view content consistency evaluation, view style consistency evaluation; Perform quality screening on the view evaluation data to determine the current optimal view, and based on the current optimal view, obtain the reconstructed current 3D scene through iterative update of MLLM descriptive prompts.
5. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, wherein, According to the current 3D scene, determine a group of noisy rendering images through forward processing of a diffusion model with a noise addition strategy, specifically including: Based on the current 3D scene, obtain a group of rendering images through image rendering generation; Perform cosine similarity analysis for controlling the influence intensity of adjacent images on the adjacent images in the group of rendering images to obtain adjacent image cosine similarities; According to the adjacent image cosine similarities, determine the group of noisy rendering images through random noise addition for view consistency.
6. The single-view 3D scene reconstruction method based on a diffusion model according to claim 5, wherein, After determining the group of noisy rendering images through random noise addition for view consistency according to the adjacent image cosine similarities, the method further includes: Based on the group of noisy rendering images, determine attention parameters through image latent feature analysis; According to the attention parameters, obtain updated components through random sampling in the Batch dimension; wherein the updated components include: keys, values; Concatenate the updated component with the corresponding key and value, and perform attention analysis on the concatenated updated component to determine the relevance-aware attention mechanism.
7. A single-view 3D scene reconstruction method based on a diffusion model according to claim 1, characterized in that, Perform reverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene, specifically including: Input the noisy rendered image group into a preset denoising network, and correct the denoising direction of the denoising network through a 3D Gaussian field to obtain the denoising result at the current time step; Based on the noisy rendered image group and the denoising result at the current time step, determine the prevention of the denoising direction from deviating from the origin; According to the prevention of the denoising direction from deviating from the origin, pull back the denoising direction to obtain the visually consistent 3D scene.
8. The single-view 3D scene reconstruction method based on a diffusion model according to claim 1, wherein, Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object, specifically including: Render the target to be edited in the visually consistent 3D scene to determine the target to be edited; Based on the target to be edited, optimize the scene parameters of the target to be edited by the gradient descent method to obtain scene optimization parameters; wherein, the constraint conditions for the gradient descent method optimization include: depth loss constraint; According to the scene optimization parameters, extract the mask region of the single-view replacement target of the target to be edited to obtain the mask region of the source object, and perform replacement and change on the mask region to determine the 3D scene of the replaced and changed object.
9. A single-view 3D scene reconstruction device based on a diffusion model, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Obtain the camera trajectory, and based on the camera trajectory, analyze through the offset camera strategy to obtain the set of views to be filled; Fill the blank areas in the set of views to be filled to obtain filled views, and perform refined loop optimization on the filled views to obtain the reconstructed current 3D scene; According to the current 3D scene, perform forward processing of the diffusion model with a noise addition strategy to determine the noisy rendered image group; Perform reverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene; Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object.
10. A non-volatile computer storage medium for single-view 3D scene reconstruction based on a diffusion model, storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: Obtain the camera trajectory, and based on the camera trajectory, analyze through the offset camera strategy to obtain the set of views to be filled; Fill the blank areas in the set of views to be filled to obtain filled views, and perform refined loop optimization on the filled views to obtain the reconstructed current 3D scene; According to the current 3D scene, perform forward processing of the diffusion model with a noise addition strategy to determine the noisy rendered image group; Perform reverse processing of the diffusion model for iterative denoising on the noisy rendered image group to obtain a visually consistent 3D scene; Input the visually consistent 3D scene into a single-view 3D editing algorithm to determine the 3D scene for replacing the changed object.
Citation Information
Patent Citations
Three-dimensional scene reconstruction method and system
CN109242959A
Single-view three-dimensional modeling method and system based on diffusion model
CN119068144A
Three-dimensional human body reconstruction method based on implicit neural network and diffusion model
CN119991967A
Geometry-aware three-dimensional synthesis in all angles
US20240265628A1
Cited By
Image generation method, image synthesis method, computing device and electronic device
CN121329804A