A sparse perspective complex outdoor scene 3D reconstruction method, electronic device and storage medium

By acquiring camera pose and initial 3D point cloud information, and combining self-attention mechanism and occlusion processing module, the 3D Gaussian radiation field is optimized, solving the problems of transient occlusion and illumination changes under sparse viewpoints in complex outdoor scenes, and achieving high-quality 3D reconstruction and immersive experience.

CN119832161BActive Publication Date: 2025-10-28SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510032590.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-10-28
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle transient occlusion and dynamic color changes caused by varying lighting conditions in complex outdoor scenes, especially under sparse viewing conditions, resulting in poor reconstruction performance.

Method used

The camera pose information and initial 3D point cloud information are obtained by the initialization module. A pseudo-supervision signal is generated by combining the self-attention mechanism. The transient occlusion is removed by the occlusion processing module. The 3D Gaussian radiation field is optimized by the course learning strategy to reconstruct the 3D image of the complex outdoor scene.

Benefits of technology

It achieves realistic reproduction of outdoor scenes from a sparse perspective, providing an immersive roaming experience and improving reconstruction results and visual consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832161B_ABST
    Figure CN119832161B_ABST
Patent Text Reader

Abstract

This invention discloses a method, electronic device, and storage medium for 3D reconstruction of complex outdoor scenes with sparse perspectives. The method includes: acquiring camera pose information and initial 3D point cloud information with rich geometric priors through an initialization module; generating a first target image through a new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, using the first target image as a pseudo-supervision signal to supervise the original rendered image; acquiring a high-fidelity second target image with transient occlusion removed through an occlusion processing module based on the camera pose information and the initial 3D point cloud information; and optimizing the 3D Gaussian radiation field through a learning strategy based on the first and second target images to reconstruct a 3D image of the complex outdoor scene. The embodiments of this invention can realistically reproduce landscape scenes, provide an immersive roaming experience, improve reconstruction results and visual consistency, and can be widely applied in the field of computer technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, electronic device and storage medium for 3D reconstruction of complex outdoor scenes with sparse perspective. Background Technology

[0002] Common methods for sparse viewpoint scene reconstruction include regularization-based methods and depth-prior-based methods. The core idea of ​​regularization-based methods is to improve modeling accuracy by introducing different regularization strategies to guide model training. For example, the RegNeRF model, proposed in related technologies, designs a patch-based regularization strategy, reducing artifacts and improving the model's geometric modeling capabilities. Another related technology proposes the CoR-GS model, which introduces 3D Gaussian splashing as a scene representation method and improves reconstruction quality through pruning and artifact-based regularization. Depth-prior-based methods typically introduce additional depth information to improve model performance. Another related technology proposes the DS-NeRF model based on depth information supervision. Yet another related technology proposes the DNGaussian model based on 3D Gaussian splashing, introducing a pre-trained depth estimation model to obtain more accurate depth information, thus achieving more accurate modeling results. However, these methods face challenges when applied to complex outdoor scenes because they assume the scene is static and do not consider transient occlusion (such as pedestrians, vehicles, etc.) and dynamic appearance colors caused by different lighting conditions.

[0003] To address complex outdoor scenes, the NeRF-W model, the first to apply implicit radiative fields to such scenes, was proposed. This model optimizes the appearance vector of each image to characterize different lighting conditions and trains a multilayer perceptron to simulate transient radiative fields, thus removing transient occlusion. The CR-NeRF model, proposed by the related researchers, handles transient occlusion by optimizing a 2D visibility map and introduces a CNN-based appearance encoder to model dynamic appearances. With the emergence and development of 3D Gaussian splashing techniques, the GS-W model based on 3D Gaussian splashing was proposed, introducing intrinsic and dynamic appearance features for each Gaussian point to better simulate different lighting conditions. The WildGaussians model was also proposed, using a learnable multilayer perceptron to extract global appearance vectors from reference images and introducing DINO feature priors to handle occlusion. However, these methods inherently rely on neural networks based on multiple views of static objects. Figure 1 Using consistent methods to model complex scenes makes it difficult to achieve good reconstruction results under sparse viewpoint input conditions. Summary of the Invention

[0004] The main objective of this invention is to propose a 3D reconstruction method, electronic device, and storage medium for complex outdoor scenes with sparse perspectives, which can realistically reproduce landscape scenes, provide an immersive roaming experience, and improve reconstruction results and visual consistency.

[0005] To achieve the above objectives, one aspect of this invention proposes a method for 3D reconstruction of complex outdoor scenes with sparse perspectives, comprising the following steps:

[0006] The camera pose information and initial 3D point cloud information with rich geometric priors are obtained through the initialization module;

[0007] Based on the camera pose information and the initial 3D point cloud information, combined with the self-attention mechanism, a first target image is generated through the new view enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image.

[0008] Based on the camera pose information and the initial 3D point cloud information, a high-fidelity second target image with transient occlusion removed is obtained through the occlusion processing module;

[0009] Based on the first target image and the second target image, the three-dimensional Gaussian radiation field is optimized through a course learning strategy to reconstruct a three-dimensional image of a complex outdoor scene.

[0010] In some embodiments, obtaining camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module includes the following steps:

[0011] Acquire multiple images of complex outdoor scenes;

[0012] The DUSt3R multi-view stereo vision model is used to obtain camera pose and 3D point cloud with rich geometric priors from complex outdoor scene images.

[0013] A pre-trained segmentation model, EVF-SAM, is used to generate corresponding occlusion masks based on predefined text guidance.

[0014] After replacing the color of the occluded area with Gaussian noise, the camera pose information and the initial 3D point cloud information are obtained through a new formula for acquiring point cloud and camera pose.

[0015] In some embodiments, the expression for the process of replacing the color of the occluded area with Gaussian noise is:

[0016]

[0017] in, M represents a collection of complex outdoor scene images with replaced occluded areas; i Images representing complex outdoor scenes The masking layer; ⊙ represents element-wise product; Represents the mean μ and variance σ 2 Gaussian noise; Represents a complex outdoor scene image; 1 indicates the relationship with M. i Matrices of the same shape, but all elements are 1;

[0018] The new formula for acquiring point cloud and camera pose is:

[0019]

[0020] in, Represents a 3D point cloud with rich geometric priors; C represents the camera pose. This represents the DUSt3R, a multi-view stereoscopic vision model.

[0021] In some embodiments, generating the first target image through a new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, includes the following steps:

[0022] First, new viewpoints are sampled from the training camera pose C; for each new viewpoint, an image I is rendered. n The DDIM inversion method is applied to transform it into a standard Gaussian distribution x. T ;

[0023] x T The input is fed into two branches of the reverse diffusion process;

[0024] The reconstruction of branch roads gradually affected x T Denoising was used to reconstruct the original rendered image; a fine-tuned diffusion model was employed by enhancing the branches. Generate high-quality images

[0025] In some embodiments, the step of generating the first target image through the new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, further includes the following steps:

[0026] The structural features of the original rendered image are injected into the enhancement branch, specifically:

[0027] In the self-attention mechanism, the query and key in the reconstruction branch are represented by Q. r and K r This indicates that the query in the enhanced branch is represented by Q, where the key and value are represented by Q. e ,K e and V e express;

[0028] By Q e and Ke Replace with Q r and K r Implement self-attention injection operation;

[0029] The expression for the self-attention injection operation is:

[0030] Among them, the output feature F e quilt Used to predict noise, where d represents the vector dimension;

[0031] After T-step denoising, a high-quality image is obtained.

[0032] AdaIn is used as the post-processor to control the image based on a user-provided reference image. The global appearance; ultimately, this image serves as a pseudo-supervision signal. To supervise the original rendered image I n .

[0033] In some embodiments, obtaining a high-fidelity second target image with transient occlusion removed by an occlusion processing module based on the camera pose information and the initial 3D point cloud information includes the following steps:

[0034] A novel strategy is employed to introduce a pre-trained segmentation network, EVF-SAM, to obtain occlusion masks.

[0035] Based on the similarity between the original information in the occluded region and the information in the adjacent regions, this similarity is used to fill in the occlusion, specifically:

[0036] Through diffusion model Inherent self-attention mechanisms and constrained priors enhance the occluded area;

[0037] Render image I from a given training perspective t and its monitoring signals Use DDIM inversion to obtain the corresponding Gaussian distribution x′. T and

[0038] Use the corresponding mask M i For x′ T and By performing weighted fusion, we obtain The expression for this process is:

[0039] right Perform a two-branch reverse diffusion process, using the corresponding mask M. iWeighted self-attention injection is performed, and the expression for this process is:

[0040] Q OH =M i ⊙Q′ e +(1-M i )⊙Q r ′,

[0041] K OH =M i ⊙K e ′+(1-M i )⊙K r ′,

[0042]

[0043] Wherein, the output feature F e 'quilt Used to predict noise; the output characteristic F e 'quilt Used to predict noise, and the characteristic F of the output is... e ′ as a supervisory signal to optimize the Gaussian radiation field; Q OH Queries representing enhanced branches after fusion; K OH Q′ represents the key in the enhanced branch after fusion. e Represents the query for the enhanced branch in the occlusion processing module; K e ′ represents the key of the enhancement branch in the occlusion processing module; Q r ′ represents the query for reconstructing branches in the occlusion processing module; K r ′ represents the key for rebuilding the branch in the occlusion handling module; V e ′ represents the value of the enhanced branch in the occlusion processing module.

[0044] In some embodiments, optimizing the three-dimensional Gaussian radiation field through a course learning strategy includes the following steps:

[0045] First, train the Gaussian radiation field with simple samples, and then gradually move on to more complex samples;

[0046] The Gaussian radiation field is trained by randomly selecting a training perspective until it is fitted. Then, based on the complexity of the new perspective, the sampled new perspective is gradually incorporated into the training process in three stages.

[0047] These new perspectives are categorized into three levels based on their Euclidean distance from the training perspective: easy, medium, and difficult.

[0048] By selecting new perspectives for training with probability β, a high-quality and multi-perspective consistent Gaussian radiation field can ultimately be obtained through this course learning strategy.

[0049] In some embodiments, the method further includes the following steps:

[0050] For rendering images from a new perspective I n Using its pseudo-monitoring signal Supervision: Where L1 represents the L1 loss function, L SSIM L represents the SSIM loss function; c Represents luminosity loss;

[0051] For training viewpoint rendering image I t When the number of iterations is less than τ, the corresponding mask M is used. i Masking transient occlusions; using pseudo-monitoring signals when the number of iterations is greater than or equal to τ. Supervise the entire image:

[0052]

[0053] in, L o Represents occlusion loss; I t The image is rendered from the training perspective;

[0054] The total loss function is expressed as: L = L o +λ3L c ;

[0055] Where λ1, λ2, λ3, and τ are hyperparameters.

[0056] Another aspect of this invention provides a 3D reconstruction system for complex outdoor scenes with sparse perspectives, comprising:

[0057] The first module is used to obtain camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module;

[0058] The second module is used to generate a first target image based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, through a new perspective enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image.

[0059] The third module is used to obtain a high-fidelity second target image with transient occlusion removed by the occlusion processing module based on the camera pose information and the initial three-dimensional point cloud information.

[0060] The fourth module is used to optimize the three-dimensional Gaussian radiation field based on the first target image and the second target image through a course learning strategy, and reconstruct a three-dimensional image of a complex outdoor scene.

[0061] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0062] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0063] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.

[0064] The embodiments of this invention include at least the following beneficial effects: This invention provides a method, electronic device, and storage medium for 3D reconstruction of complex outdoor scenes with sparse perspectives. This scheme acquires camera pose information and initial 3D point cloud information with rich geometric priors through an initialization module; based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, a first target image is generated through a new perspective enhancement module. This first target image is used as a pseudo-supervision signal to supervise the original rendered image; based on the camera pose information and the initial 3D point cloud information, a high-fidelity second target image with transient occlusion removed is acquired through an occlusion processing module; based on the first target image and the second target image, a 3D Gaussian radiation field is optimized through a course learning strategy to reconstruct a 3D image of the complex outdoor scene. The embodiments of this invention can realistically reproduce landscape scenes, provide an immersive roaming experience, and improve reconstruction effects and visual consistency. Attached Figure Description

[0065] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0066] Figure 2 This is a flowchart of the overall steps provided in the embodiments of the present invention;

[0067] Figure 3 This is a schematic diagram of the overall structure of SparseGS-W provided in an embodiment of the present invention;

[0068] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.

[0070] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”

[0071] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.

[0073] The sparse-viewpoint complex outdoor scene 3D reconstruction method, electronic device, and storage medium provided in this invention relate to the field of computer technology. The sparse-viewpoint complex outdoor scene 3D reconstruction method provided in this invention can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited thereto; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the sparse-viewpoint complex outdoor scene 3D reconstruction method, but is not limited to the above forms.

[0074] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0075] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.

[0076] Server 101 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0077] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0078] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc. It can also be a vehicle-mounted terminal of the various device types described above, but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.

[0079] Exemplary based on Figure 1 The implementation environment shown in this embodiment of the invention provides a 3D reconstruction method for complex outdoor scenes with sparse perspective. The following description uses the application of this 3D reconstruction method for complex outdoor scenes with sparse perspective in server 101 as an example. It can be understood that this method can also be applied to terminal 102.

[0080] Reference Figure 2 , Figure 2 This is a flowchart illustrating a sparse-view, complex outdoor scene 3D reconstruction method applied to a server, provided as an embodiment of the present invention. The execution subject of this method can be any of the aforementioned computer devices (including servers or terminals). (Refer to...) Figure 2 The method may include the following steps:

[0081] The camera pose information and initial 3D point cloud information with rich geometric priors are obtained through the initialization module;

[0082] Based on the camera pose information and the initial 3D point cloud information, combined with the self-attention mechanism, a first target image is generated through the new view enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image.

[0083] Based on the camera pose information and the initial 3D point cloud information, a high-fidelity second target image with transient occlusion removed is obtained through the occlusion processing module;

[0084] Based on the first target image and the second target image, the three-dimensional Gaussian radiation field is optimized through a course learning strategy to reconstruct a three-dimensional image of a complex outdoor scene.

[0085] In some embodiments, obtaining camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module includes the following steps:

[0086] Acquire multiple images of complex outdoor scenes;

[0087] The DUSt3R multi-view stereo vision model is used to obtain camera pose and 3D point cloud with rich geometric priors from complex outdoor scene images.

[0088] A pre-trained segmentation model, EVF-SAM, is used to generate corresponding occlusion masks based on predefined text guidance.

[0089] After replacing the color of the occluded area with Gaussian noise, the camera pose information and the initial 3D point cloud information are obtained through a new formula for acquiring point cloud and camera pose.

[0090] In some embodiments, the expression for the process of replacing the color of the occluded area with Gaussian noise is:

[0091]

[0092] in, M represents a collection of complex outdoor scene images with replaced occluded areas; i Images representing complex outdoor scenes The masking layer; ⊙ represents element-wise product; Represents the mean μ and variance σ 2 Gaussian noise; Represents a complex outdoor scene image; 1 indicates the relationship with M. i Matrices of the same shape, but all elements are 1;

[0093] The new formula for acquiring point cloud and camera pose is:

[0094]

[0095] in, Represents a 3D point cloud with rich geometric priors; C represents the camera pose. This represents the DUSt3R, a multi-view stereoscopic vision model.

[0096] In some embodiments, generating the first target image through a new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, includes the following steps:

[0097] First, new viewpoints are sampled from the training camera pose C; for each new viewpoint, an image I is rendered. n The DDIM inversion method is applied to transform it into a standard Gaussian distribution x. T ;

[0098] x T The input is fed into two branches of the reverse diffusion process;

[0099] The reconstruction of branch roads gradually affected x T Denoising was used to reconstruct the original rendered image; a fine-tuned diffusion model was employed by enhancing the branches. Generate high-quality images

[0100] In some embodiments, the step of generating the first target image through the new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, further includes the following steps:

[0101] The structural features of the original rendered image are injected into the enhancement branch, specifically:

[0102] In the self-attention mechanism, the query and key in the reconstruction branch are represented by Q. r and K r This indicates that the query in the enhanced branch is represented by Q, where the key and value are represented by Q. e ,K e and V e express;

[0103] By Q e and K e Replace with Q r and K r Implement self-attention injection operation;

[0104] The expression for the self-attention injection operation is:

[0105] Among them, the output feature F e quilt Used to predict noise, where d represents the vector dimension;

[0106] After T-step denoising, a high-quality image is obtained.

[0107] AdaIn is used as the post-processor to control the image based on a user-provided reference image. The global appearance; ultimately, this image serves as a pseudo-supervision signal. To supervise the original rendered image I n .

[0108] In some embodiments, obtaining a high-fidelity second target image with transient occlusion removed by an occlusion processing module based on the camera pose information and the initial 3D point cloud information includes the following steps:

[0109] A novel strategy is employed to introduce a pre-trained segmentation network, EVF-SAM, to obtain the occlusion mask M.

[0110] Based on the similarity between the original information in the occluded region and the information in the adjacent regions, this similarity is used to fill in the occlusion, specifically:

[0111] Through diffusion model Inherent self-attention mechanisms and constrained priors enhance the occluded area;

[0112] Render image I from a given training perspective t and its monitoring signals Use DDIM inversion to obtain the corresponding Gaussian distribution x′. T and

[0113] Use the corresponding mask M i For x′ T and By performing weighted fusion, we obtain The expression for this process is:

[0114] right Perform a two-branch reverse diffusion process, using the corresponding mask M. i Weighted self-attention injection is performed, and the expression for this process is:

[0115] Q OH =M i ⊙Q′ e +(1-M i )⊙Q r ′,

[0116] K OH =M i ⊙K e ′+(1-M i )⊙K r ′,

[0117]

[0118] Wherein, the output feature F e 'quilt Used to predict noise; the output characteristic F e 'quilt Used to predict noise, and the characteristic F of the output is... e′ as a supervisory signal to optimize the Gaussian radiation field; Q OH Queries representing enhanced branches after fusion; K OH Q′ represents the key in the enhanced branch after fusion. e Represents the query for the enhanced branch in the occlusion processing module; K e ′ represents the key of the enhancement branch in the occlusion processing module; Q r ′ represents the query for reconstructing branches in the occlusion processing module; K r V' represents the key for the reconstructed branch in the occlusion handling module; V' e This represents the value of the enhanced branch in the occlusion processing module.

[0119] In some embodiments, optimizing the three-dimensional Gaussian radiation field through a course learning strategy includes the following steps:

[0120] First, train the Gaussian radiation field with simple samples, and then gradually move on to more complex samples;

[0121] The Gaussian radiation field is trained by randomly selecting a training perspective until it is fitted. Then, based on the complexity of the new perspective, the sampled new perspective is gradually incorporated into the training process in three stages.

[0122] These new perspectives are categorized into three levels based on their Euclidean distance from the training perspective: easy, medium, and difficult.

[0123] By selecting new perspectives for training with probability β, a high-quality and multi-perspective consistent Gaussian radiation field can ultimately be obtained through this course learning strategy.

[0124] In some embodiments, the method further includes the following steps:

[0125] For rendering images from a new perspective I n Using its pseudo-monitoring signal Supervision: Where L1 represents the L1 loss function, L SSIM L represents the SSIM loss function; c Represents luminosity loss;

[0126] For training viewpoint rendering image I t When the number of iterations is less than τ, the corresponding mask M is used. i Masking transient occlusions; using pseudo-monitoring signals when the number of iterations is greater than or equal to τ. Supervise the entire image:

[0127]

[0128] in, L o Represents occlusion loss; It The image is rendered from the training perspective;

[0129] The total loss function is expressed as: L = L o +λ3L c ;

[0130] Where λ1, λ2, λ3, and τ are hyperparameters.

[0131] The implementation process of the embodiments of the present invention in specific application scenarios will be described in detail below with reference to the accompanying drawings:

[0132] Specifically, the overall structure of SparseGS-W is as follows: Figure 3 As shown, this method consists of three modules: an initialization module, a new view enhancement module, and an occlusion handling module. Given a complex outdoor scene with sparse viewpoints, the initialization module is responsible for estimating the camera pose and generating an initial point cloud. The new view enhancement module utilizes a diffusion model prior to enhance low-quality new viewpoints, ensuring that the generated images have high quality and 3D consistency. The occlusion handling module is responsible for removing transient occlusions during training. Furthermore, this embodiment introduces a curriculum learning strategy to accelerate model convergence and improve reconstruction accuracy.

[0133] The following is a detailed description of the specific content of each module:

[0134] 1. Initialize the module:

[0135] 3D Gaussian splashing is an explicit representation method for static scenes. Its ability to model the scene largely depends on the quality of the initial point cloud and the supervision signal. Given several images of a complex outdoor scene... Traditional motion reconstruction methods cannot provide sufficient geometric priors, which affects the training quality of the 3D Gaussian radiation field. Therefore, this embodiment of the invention employs the multi-view stereo vision model DUSt3R, using... Let C be the camera pose and P be the 3D point cloud with rich geometric priors.

[0136]

[0137] However, this embodiment of the invention observes that the DUSt3R model assumes the absence of transient occlusions in the image, causing it to mistakenly identify transient occlusions as textures on the surface of static objects, thus affecting subsequent training. To address this issue, this embodiment of the invention uses a pre-trained segmentation model, EVF-SAM, to generate corresponding occlusion masks based on user-provided text guidance. Then, in this embodiment of the invention, the color of the occluded area is replaced with Gaussian noise:

[0138]

[0139] Where 1 represents M i Matrices of the same shape, but all elements are 1. Represents the mean μ and variance σ 2 The Gaussian noise, with mean and variance derived from the unoccluded region, is used in this embodiment of the invention. The formulas for obtaining point clouds and camera poses are modified as follows:

[0140]

[0141] 2. New Perspective Enhancement Module:

[0142] The objective of this invention is to enhance low-quality images rendered during the training process using a 3D Gaussian radiation field. To avoid identity shift during enhancement, this invention uses sparse training views as anchor points to fine-tune the diffusion model ε. θ (x t This restricts the generation space to a clean subspace. To distinguish it from the original diffusion model, the fine-tuned diffusion model is expressed as...

[0143] This invention first samples new viewpoints from the training camera pose C. For each new viewpoint, an image I is rendered. n In this embodiment of the invention, DDIM inversion is applied to convert it into a standard Gaussian distribution x. T Then, x T The input is fed into two backdiffusion process branches. The reconstruction branch gradually affects x. T Denoising is used to reconstruct the original rendered image, while the enhancement branch focuses on using... Generate high-quality images

[0144] This invention observes that while fine-tuning the diffusion model can effectively constrain the generation space to produce high-quality images with the same identity, it struggles to accurately preserve the image structure, resulting in poor 3D consistency. Therefore, this invention injects the structural features of the original rendered image into the enhancement branch. Specifically, in the self-attention mechanism, the query and key in the reconstruction branch are represented by Q... r and K r This indicates that in the enhanced branch query, the key and value are represented by Q. e ,K e and V e Indicated by Q. e and K e Replace with Q r and K r Self-attention injection operation was implemented:

[0145]

[0146] Among them, the output feature F e quilt Used to predict noise, where d represents the vector dimension. After T-step denoising, this embodiment of the invention obtains a high-quality image. This image eliminates artifacts while preserving identity information. This embodiment of the invention does not train an additional appearance extraction network; instead, it uses AdaIn as post-processing to control the image based on a user-provided reference image. The overall appearance. Ultimately, this image serves as a pseudo-supervision signal. To supervise the original rendered image I n .

[0147] 3. Obstruction handling module:

[0148] Previous methods have used dense viewpoints to train neural networks to predict transient occlusions, which is difficult to succeed under sparse viewpoint conditions. Furthermore, their unsupervised training paradigm results in poor flexibility, failing to remove specific occlusions according to user needs. Therefore, this invention employs a novel strategy, introducing a pre-trained segmentation network, EVF-SAM, to flexibly obtain occlusion masks. This invention has observed that the original information in an occluded region is often similar to the information in adjacent regions, or similar to certain regions in other training views. Based on this finding, this invention can utilize this similarity to fill in occlusions. Specifically, this invention uses a diffusion model... Inherent self-attention mechanisms and constrained priors are used to enhance occluded regions. Given a training viewpoint, render image I. t and its monitoring signals This invention uses DDIM inversion to obtain the corresponding x′. T and Then, use the corresponding mask M i For x′ T and By performing weighted fusion, we obtain

[0149]

[0150] Similar to the previous section, the embodiments of the present invention address... The reverse diffusion process is executed in two branches. To distinguish the symbols, this embodiment of the invention uses a superscript to distinguish Q′. e K e ′, V e ′, Q r ′, K r ′ and Q e K e , V e Qr and K r The difference is that the embodiments of the present invention use a corresponding mask M. i Perform weighted self-attention injection:

[0151] Q OH =M i ⊙Q′ e +(1-M i )⊙Q r ′,

[0152] K OH =M i ⊙K e ′+(1-M i )⊙K r ′,

[0153]

[0154] Output feature F e 'quilt Used to predict noise. After T-step denoising and post-processing, the embodiments of the present invention obtain a high-fidelity image free from transient occlusion. This is used as a supervisory signal to optimize the Gaussian radiation field. Weighted fusion and injection operations help minimize the information loss inherent in the diffusion process, making the constrained... Focus on enhancing the occlusion area.

[0155] 4. Course Learning Strategies:

[0156] Inspired by the human learning process, which typically progresses from simple to complex, this invention introduces a curriculum-based learning strategy to progressively optimize the 3D Gaussian radiation field. Specifically, this invention first trains the Gaussian radiation field using simple samples, then gradually moves to more complex samples. This invention randomly selects training perspectives to train the Gaussian radiation field until a fit is achieved, and then incorporates the sampled new perspectives into the training process in three stages based on their complexity. These new perspectives are categorized into three levels—easy, medium, and difficult—based on their Euclidean distance from the training perspective. To reduce errors introduced by the diffusion model, this invention selects new perspectives for training with a probability β. Through this curriculum-based learning strategy, this invention ultimately obtains a high-quality Gaussian radiation field that is consistent across multiple perspectives, improving reconstruction results and visual consistency.

[0157] 5. Model loss function:

[0158] For rendering images from a new perspective I n The embodiments of the present invention use its pseudo-monitoring signal. Supervision:

[0159]

[0160] Where L1 represents the L1 loss function, L SSIM This represents the SSIM loss function.

[0161] For training viewpoint rendering image I t When the number of iterations is less than τ, the embodiment of the present invention uses a corresponding mask M. i Masking transient occlusion. When the number of iterations is greater than or equal to τ, a pseudo-monitoring signal is used. Supervise the entire image:

[0162]

[0163] in Combining all loss terms, the total loss function can be expressed as:

[0164] L = L o +λ3L c ,

[0165] Where λ1, λ2, λ3, and τ are hyperparameters, set to 0.8, 0.2, 1.0, and 6500 respectively in this method. This method is trained for 7500 epochs on the dataset. β is set to 0.3 in this method.

[0166] In summary, this invention proposes for the first time a 3D reconstruction method for complex outdoor scenes using sparse perspectives: SparseGS-W. This method utilizes prior information from a pre-trained diffusion model to simplify the problem of high-quality novel perspective synthesis into an image enhancement problem with 3D consistency. The method rapidly fine-tunes the pre-trained diffusion model using a given sparse perspective and introduces a self-attention injection strategy to preserve the unique identity information and local geometry of the original low-quality image during the enhancement process. Based on the similarity between the original texture of the occluded region and observations of similar regions in neighboring regions or other training perspectives, the method gradually removes transient occlusions in the generation space using the fine-tuned diffusion model. To model multiple appearances, the method employs a simple yet effective post-processing operation that can render high-quality images with similar appearances based on any user-provided reference image without incurring additional training costs.

[0167] This invention allows users to virtually tour famous landmarks around the world, such as the Brandenburg Gate in Germany, from the comfort of their homes. Users can achieve this simply by selecting a few photos from their personal albums taken during previous trips or by directly searching and downloading photos uploaded by others from the internet. This invention realistically recreates landscape scenes and, based on user preferences, removes arbitrary transient occlusions from the virtual scene, simulates different seasons and lighting effects, providing an immersive travel experience. Experimental comparisons show that the method of this invention outperforms current state-of-the-art methods by 5.74 dB in PSNR, 0.24 dB in LPIPS, and 0.20 dB in SSIM on the commonly used complex outdoor scene dataset PhotoTourism. Furthermore, the method of this invention also surpasses existing state-of-the-art methods on commonly used unsupervised metrics such as FID, MUSIQ, and ClipIQA, fully demonstrating the effectiveness and superiority of the method of this invention.

[0168] The method based on this invention can be widely applied to the following related fields:

[0169] 1. Virtual roaming of outdoor scenes in the fields of culture and tourism, education, electronic entertainment, and VR. SparseGS-W can be used for virtual roaming of famous landmarks around the world. Users only need to provide a few photos with a sparse perspective to generate realistic 3D scenes, and it can remove transient occlusion and simulate different seasons and lighting effects according to user preferences, making virtual tourism possible.

[0170] 2. The plug-and-play new perspective enhancement module can be applied to any other new perspective compositing task.

[0171] 3. The course learning strategy can be applied to other 3D reconstruction tasks based on 3D Gaussian splashing technology.

[0172] 4. Cultural Heritage Conservation and Digitization. The SparseGS-W method enables high-quality reconstruction under limited data conditions, supporting the digital conservation of cultural heritage.

[0173] Another aspect of this invention provides a 3D reconstruction system for complex outdoor scenes with sparse perspectives, comprising:

[0174] The first module is used to obtain camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module;

[0175] The second module is used to generate a first target image based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, through a new perspective enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image.

[0176] The third module is used to obtain a high-fidelity second target image with transient occlusion removed by the occlusion processing module based on the camera pose information and the initial three-dimensional point cloud information.

[0177] The fourth module is used to optimize the three-dimensional Gaussian radiation field based on the first target image and the second target image through a course learning strategy, and reconstruct a three-dimensional image of a complex outdoor scene.

[0178] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0179] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned 3D reconstruction method for complex outdoor scenes with sparse perspectives. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0180] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0181] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0182] The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.

[0183] The memory 402 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and called by the processor 401 to execute the sparse viewpoint complex outdoor scene 3D reconstruction method of the embodiments of this invention.

[0184] Input / output interface 403 is used to implement information input and output;

[0185] The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0186] Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404);

[0187] The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.

[0188] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described sparse-viewpoint complex outdoor scene 3D reconstruction method.

[0189] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0190] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0191] It should be noted that in various specific embodiments of the present invention, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of the present invention require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data for the normal operation of the embodiments of the present invention acquired.

[0192] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.

[0193] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0195] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0196] The terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0197] It should be understood that in this invention, "at least one (item)" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0198] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0199] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0200] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0201] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0202] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.

Claims

1. A method for 3D reconstruction of complex outdoor scenes with sparse perspective, characterized in that, Includes the following steps: The camera pose information and initial 3D point cloud information with rich geometric priors are obtained through the initialization module; Based on the camera pose information and the initial 3D point cloud information, combined with the self-attention mechanism, a first target image is generated through the new view enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image. Based on the camera pose information and the initial 3D point cloud information, a high-fidelity second target image with transient occlusion removed is obtained through the occlusion processing module; Based on the first target image and the second target image, the three-dimensional Gaussian radiation field is optimized through a course learning strategy to reconstruct a three-dimensional image of a complex outdoor scene; The process of obtaining camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module includes the following steps: Acquire multiple images of complex outdoor scenes; The DUSt3R multi-view stereo vision model is used to obtain camera pose and 3D point cloud with rich geometric priors from complex outdoor scene images. A pre-trained segmentation model, EVF-SAM, is used to generate corresponding occlusion masks based on predefined text guidance. After replacing the color of the occluded area with Gaussian noise, the camera pose information and the initial 3D point cloud information are obtained by using the formula for obtaining the point cloud and camera pose. The step of generating a first target image based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, through a new perspective enhancement module, includes the following steps: First, let's start with training camera pose. Medium-sample new perspectives; render the image for each new perspective. The DDIM inversion method is applied to convert it into a standard Gaussian distribution. ; Will The input is fed into two branches of the reverse diffusion process; Reconstructing branch roads gradually Denoising was used to reconstruct the original rendered image; a high-quality image was generated using a fine-tuned diffusion model by enhancing branches. ; The step of generating the first target image through the new perspective enhancement module based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, further includes the following steps: The structural features of the original rendered image are injected into the enhancement branch, specifically: In the self-attention mechanism, the query and key in the reconstruction branch are respectively used as and This indicates that the query in the enhanced branch is represented by the key and value respectively. , and express; By and Replace with and Implement self-attention injection operation; The expression for the self-attention injection operation is: Among them, output features The diffused model is used to predict noise, where d represents the vector dimension; After H-step denoising, a high-quality image is obtained. ; AdaIn is used as the post-processor to control the image based on a user-provided reference image. The global appearance; ultimately, this image serves as a pseudo-supervision signal. To supervise the original rendered image .

2. The 3D reconstruction method for complex outdoor scenes with sparse perspectives according to claim 1, characterized in that, The expression for the process of replacing the color of the occluded area with Gaussian noise is: in, This represents a collection of images depicting complex outdoor scenes with replaced obscured areas. Images representing complex outdoor scenes The masking layer; Represents element-wise product; Representative mean and variance Gaussian noise; Images representing complex outdoor scenes; Indicates and Matrices of the same shape, but all elements are 1; The formula for obtaining the point cloud and camera pose is: in, This represents a 3D point cloud with rich geometric priors; Represents camera pose; This represents the DUSt3R, a multi-view stereoscopic vision model.

3. The 3D reconstruction method for complex outdoor scenes with sparse perspectives according to claim 1, characterized in that, The step of obtaining a high-fidelity second target image, free from transient occlusion, through an occlusion processing module based on the camera pose information and the initial 3D point cloud information includes the following steps: A strategy is used to introduce a pre-trained segmentation network, EVF-SAM, to obtain occlusion masks. ; Based on the similarity between the original information in the occluded region and the information in the adjacent regions, this similarity is used to fill in the occlusion, specifically: Enhance the occluded region by leveraging the inherent self-attention mechanism and constrained priors of the diffusion model; Render image given training viewpoint and its monitoring signals The corresponding Gaussian distribution is obtained by using DDIM inversion. and ; Use the appropriate mask right and By performing weighted fusion, we obtain The expression for this process is: right Perform the reverse diffusion process in two branches, using the corresponding masks. Weighted self-attention injection is performed, and the expression for this process is: Among them, the output features The diffuse model is used to predict noise, and the features of the output are... Optimize the Gaussian radiation field as a monitoring signal; Queries representing enhanced branches after fusion; This represents the key in the enhanced branch after fusion; This represents a query for the enhanced branch in the occlusion handling module; The key representing the enhanced branch in the occlusion processing module; This represents a query for reconstructing branches in the occlusion handling module; The key represents the reconstructed branch in the occlusion handling module; This represents the value of the enhancement branch in the occlusion handling module; Represents element-wise product; Indicates and Matrices of the same shape, but all elements are 1; Images representing complex outdoor scenes The masking layer.

4. The 3D reconstruction method for complex outdoor scenes with sparse perspectives according to claim 1, characterized in that, The optimization of the three-dimensional Gaussian radiation field through the course learning strategy includes the following steps: First, train the Gaussian radiation field with simple samples, and then gradually move on to more complex samples; The Gaussian radiation field is trained by randomly selecting a training perspective until it is fitted. Then, based on the complexity of the new perspective, the sampled new perspective is gradually incorporated into the training process in three stages. These new perspectives are categorized into three levels based on their Euclidean distance from the training perspective: easy, medium, and difficult. With probability By selecting new perspectives for training, this course learning strategy can ultimately yield high-quality and consistent Gaussian radiation fields from multiple perspectives.

5. A method for 3D reconstruction of complex outdoor scenes with sparse perspectives according to any one of claims 1-4, characterized in that, The method further includes the following steps: For rendering images from a new perspective Using its pseudo-monitoring signal Supervision: in, Describes the L1 loss function. Represents the SSIM loss function; Represents luminosity loss; For training viewpoint rendering images When the number of iterations is less than When using a mask Masking transient occlusion; When the number of iterations is greater than or equal to When using pseudo-monitoring signals Supervise the entire image: in, ; Represents shading loss; The image is rendered from the training perspective; Represents element-wise product; Indicates and Matrices of the same shape, but all elements are 1; Images representing complex outdoor scenes; Images representing complex outdoor scenes The masking layer; The total loss function is expressed as: ; in, , , , This is a hyperparameter.

6. A system for implementing the 3D reconstruction method for complex outdoor scenes with sparse viewpoints as described in any one of claims 1-5, characterized in that, include: The first module is used to obtain camera pose information and initial 3D point cloud information with rich geometric priors through the initialization module; The second module is used to generate a first target image based on the camera pose information and the initial 3D point cloud information, combined with a self-attention mechanism, through a new perspective enhancement module. The first target image is used as a pseudo-supervision signal to supervise the original rendered image. The third module is used to obtain a high-fidelity second target image with transient occlusion removed by the occlusion processing module based on the camera pose information and the initial three-dimensional point cloud information. The fourth module is used to optimize the three-dimensional Gaussian radiation field based on the first target image and the second target image through a course learning strategy, and reconstruct a three-dimensional image of a complex outdoor scene.

7. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Obstacle removal and three-dimensional reconstruction method based on deep learning

    CN114399814A

  • Sparse visual angle three-dimensional reconstruction method based on depth prior information

    CN118657888A