Method and apparatus for 3d reconstruction based on sparse view intra-prior
By using a method based on sparse view inline priors, the overfitting problem of 3D reconstruction under sparse view is solved. Through initialization, random sampling and optimization regularization techniques, the quality of 3D reconstruction and the consistency of rendered images under sparse view are improved.
Patent Information
- Application Number
- CN202411576212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-11-06
AI Technical Summary
Existing 3D reconstruction methods are prone to overfitting and performance degradation under sparse perspectives, making it difficult to synthesize high-quality new perspective views. Furthermore, directly using diffusion model visual priors requires a large amount of external 3D data and computational resources, making it difficult to effectively extract visual information from sparse perspectives.
By using a method based on sparse view inline priors, sparse point clouds are determined from sparse view images, and pseudo-views are generated through initialization and random sampling. The parameters of the 3D Gaussian point cloud are optimized by combining multiple optimization regularization terms and loss functions. Depth regularization and geometric consistency regularization are introduced to correct the distribution of the rendered image and control the direction of pattern search.
It improves the quality of 3D reconstruction under sparse perspective, reduces dependence on external data and computing resources, and improves the quality and consistency of rendered images.
Smart Images

Figure CN119516110B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computer technology, in particular to a three-dimensional reconstruction method and device based on sparse view in-line prior. BACKGROUND
[0002] Three-dimensional Gaussian Splatting methods based on gradient optimization and differentiable rendering are one of the research hotspots of three-dimensional reconstruction (also known as novel view synthesis) task. Such methods often require a large number of dense sparse views for training. In the case of sparse training views, most three-dimensional Gaussian Splatting methods will have serious overfitting and performance degradation problems, and it is difficult to synthesize high-quality new view views. To solve this problem, most researchers choose to introduce external priors to supervise the optimization of sparse view three-dimensional reconstruction, such as semantic information priors, monocular depth priors, and diffusion model visual priors. For diffusion priors, some researchers use a large amount of computing resources to fine-tune the diffusion model prior or pre-train the image encoder through external three-dimensional data. Some researchers do not fine-tune and pre-train when introducing diffusion model visual priors for guidance, but it is difficult to directly extract diffusion model visual prior knowledge, and thus it is difficult to effectively supplement the missing visual information of sparse views. Although the diffusion model visual prior has shown excellent ability through fractional distillation sampling in text-three-dimensional model generation and image-three-dimensional model generation tasks, it performs poorly in the sparse view three-dimensional reconstruction task. This is due to the essential difference between sparse views and text prompts: for invisible views, unlike text prompts, the ideal rendering image supervision information in sparse views is not completely absent. Due to the consistency of three-dimensional geometry and structure, in-line prior supervision information exists in the given sparse views. Due to the domain shift between specific scenes and diffusion model priors, and the inherent suboptimality of the distribution of rendering images under sparse views, if fractional distillation sampling is directly used, it is easy to deviate from the target mode of the diffusion model prior distribution domain, so a large amount of external three-dimensional labeled data and computing resources are often needed for domain correction, without fine-tuning the diffusion model and training the encoder to facilitate the domain correction of the diffusion model prior, and thus effectively extracting three-dimensional knowledge in the diffusion model.
[0003] The above information disclosed in this Background section is only for the purpose of enhancing the understanding of the background of the present inventive concepts, and therefore, it can contain information that does not form the prior art that is already known in this country to those skilled in the art. SUMMARY
[0004] This Background section is intended to provide a brief overview of concepts that are described in more detail in the detailed description section that follows. This Summary section is not intended to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.
[0005] Some embodiments of the present disclosure propose a three-dimensional reconstruction method based on sparse view intra-link prior, to solve one or more of the technical problems mentioned in the background section.
[0006] In a first aspect, some embodiments of the present disclosure provide a three-dimensional reconstruction method based on sparse view intra-link prior, the method comprising: determining a sparse point cloud of a three-dimensional scene from sparse view images; initializing a three-dimensional Gaussian point cloud according to the sparse point cloud to obtain an initialized three-dimensional Gaussian point cloud; randomly sampling a preset range in each of the sparse views of the sparse view images to generate pseudo views to obtain a set of pseudo views; determining a plurality of optimization regular terms according to the sparse views and the set of pseudo views; determining a loss function according to the plurality of optimization regular terms; and optimizing parameters of the initialized three-dimensional Gaussian point cloud according to the loss function to perform sparse view three-dimensional reconstruction.
[0007] In a second aspect, some embodiments of the present disclosure provide a three-dimensional reconstruction device based on sparse view intra-link prior, the device comprising: a first determining unit configured to determine a sparse point cloud of a three-dimensional scene from sparse view images; an initializing unit configured to initialize a three-dimensional Gaussian point cloud according to the sparse point cloud to obtain an initialized three-dimensional Gaussian point cloud; a sampling unit configured to randomly sample a preset range in each of the sparse views of the sparse view images to generate pseudo views to obtain a set of pseudo views; a second determining unit configured to determine a plurality of optimization regular terms according to the sparse views and the set of pseudo views; a third determining unit configured to determine a loss function according to the plurality of optimization regular terms; and an optimization unit configured to optimize parameters of the initialized three-dimensional Gaussian point cloud according to the loss function to perform sparse view three-dimensional reconstruction.
[0008] The present disclosure corrects the rendering image distribution by intra-link prior, decomposes the optimization target of fractional distillation sampling into two sub-targets, and uses the correction distribution as the intermediate state of the optimization target to control the mode search direction. On this basis, a three-dimensional Gaussian point cloud is used as an explicit three-dimensional representation, and depth regularization is introduced to support intra-link prior and geometric consistency regularization to reduce the difference between the pixel-level rendering image distribution and the correction distribution, thereby suppressing mode shift and improving the quality of three-dimensional reconstruction under sparse views. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and elements and elements are not necessarily drawn to scale.
[0010] Figure 1is a flowchart of some embodiments of a method of three-dimensional reconstruction based on sparse view interlinking prior according to the present disclosure;
[0011] Figure 2 is a structural schematic diagram of some embodiments of an apparatus of three-dimensional reconstruction based on sparse view interlinking prior according to the present disclosure; DETAILED DESCRIPTION
[0012] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0013] It should also be noted that, for the sake of brevity, only the parts of the drawings that are relevant to the present disclosure are shown. The embodiments and features in the present disclosure can be combined with each other in the case of no conflict.
[0014] It should be noted that the terms “first”, “second”, and the like mentioned in the present disclosure are merely used to distinguish different devices, modules, or units, and are not intended to limit the functions of these devices, modules, or units, or the order or interdependence of these functions.
[0015] It should be noted that the terms “one”, “multiple” mentioned in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as “one or more”.
[0016] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of these messages or information.
[0017] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0018] Figure 1 is a flowchart 100 of some embodiments of a method of three-dimensional reconstruction based on sparse view interlinking prior according to the present disclosure. The method of three-dimensional reconstruction based on sparse view interlinking prior includes the following steps:
[0019] Step 101, determining a sparse point cloud of a three-dimensional scene from sparse view images.
[0020] In some embodiments, the subject (e.g., a computing device) of the method of three-dimensional reconstruction based on sparse view interlinking prior can determine a sparse point cloud of a three-dimensional scene from sparse view images.
[0021] Here, the sparse view image can refer to an image covering a part of a view.
[0022] As an example, the execution subject can determine a sparse point cloud of a three-dimensional scene from the sparse view image by feature extraction and matching techniques using a structure from motion method.
[0023] It should be noted that the computing device can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. For example, the computing device can be the target terminal. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.
[0024] At step 102, a three-dimensional Gaussian point cloud is initialized according to the sparse point cloud, to obtain an initialized three-dimensional Gaussian point cloud.
[0025] In some embodiments, the execution subject can initialize a three-dimensional Gaussian point cloud according to the sparse point cloud, to obtain an initialized three-dimensional Gaussian point cloud.
[0026] Here, the three-dimensional Gaussian point cloud can represent a set of Gaussian points. The three-dimensional position of the three-dimensional Gaussian point cloud is represented as:
[0027]
[0028] where μ n represents the position of the three-dimensional Gaussian point cloud, ∑ n represents a covariance matrix, c n represents a color, α n represents an opacity, and n represents the ordinal number of the Gaussian point. x represents a random variable. T represents transposition. represents the inverse of the covariance matrix. The above parameters are all optimized in a training process, where the covariance matrix ∑ n A rotation matrix R n and a scaling matrix S n are used for representation to ensure positive definiteness when optimized: where R n represents the rotation matrix. S n represents the scaling matrix. represents the inverse of the scaling matrix. represents the inverse of the rotation matrix.
[0029] Optionally, after initializing the 3D Gaussian point cloud based on the sparse point cloud to obtain the initialized 3D Gaussian point cloud, the method further includes:
[0030] The first step is to render the initialized 3D Gaussian point cloud to obtain a 2D image.
[0031] As an example, the above-mentioned execution entity can be achieved through ∑′ n =JW∑ n W T J T Based on the viewpoint transformation matrix W and the affine approximation Jacobian matrix J of the projection matrix, the covariance matrix ∑′ under the camera viewpoint is determined. n , as a two-dimensional image. T represents transpose.
[0032] The second step is to sort the Gaussian point clouds corresponding to the above two-dimensional images to obtain sorted Gaussian point clouds.
[0033] As an example, the aforementioned execution entity can sort the Gaussian points in the above two-dimensional image that contribute to the rendering of pixels, and obtain a sorted Gaussian point cloud.
[0034] The third step is to render the sorted Gaussian point cloud to obtain the corresponding colors of the rendered image.
[0035] As an example, the above-mentioned execution entity can render the sorted Gaussian point cloud using the following formula to obtain the corresponding color of the rendered image: in, It is a pseudo-perspective rendered image. At pixel p j The rendered colors on It is obtained by projecting the two-dimensional covariance matrix ∑′ of the Gaussian points in two dimensions. n and the opacity α of the Gaussian point itself n The obtained two-dimensional projected Gaussian point at pixel p j The opacity of the point. The subscript n represents the ordinal number of the Gaussian point, and the subscript m represents the ordinal number less than n. N represents the maximum ordinal number of the Gaussian point. The j above represents the identifier of the pseudo-viewpoint.
[0036] Step 103: Randomly sample within a preset range of each sparse viewpoint in the above sparse viewpoint image to generate a pseudo viewpoint and obtain a pseudo viewpoint set.
[0037] In some embodiments, the execution entity may randomly sample within a preset range of each sparse viewpoint in the sparse viewpoint image to generate a pseudo viewpoint and obtain a pseudo viewpoint set.
[0038] Here, the preset range can refer to a preset circumferential range. For example, the preset range can refer to a preset circumferential range with the sparse view angle as the center and 1 as the radius.
[0039] Optionally, the execution subject can perform random sampling in the preset range of each sparse view angle of the sparse view angle image to generate pseudo view angles, and obtain a pseudo view angle set, by the following steps:
[0040] First, set the training image set corresponding to each sparse view angle.
[0041] Here, the training image set can represent Each sparse view angle can represent v i .
[0042] Second, construct a transformation function.
[0043] Here, the transformation function can represent Where R j→i represents the relative pose transformation between two view angles, D j represents the rendering depth of the pseudo view angle set v j .
[0044] Third, transform the training image set from each sparse view angle to the pseudo view angle set by the transformation function.
[0045] Here, the calculation formula for each pixel p j in the pseudo view angle set is: Where d n is the z-buffer of the Gaussian point n. D j (p j ) represents each pixel p j on the rendering depth D j of the pseudo view angle set v j . In the transformation process, for the pixel p j on the pseudo view angle set v j , the pixel p j on the visible view angle v j→i , the transformation relationship calculation formula is: p j→j ~ KR j→i D j (p j )K -1 p j . Where K is the camera intrinsic parameter. Thus, the training image set of each sparse view angle v i to the pseudo view angle set v j is obtained. The transformation image Each pixel p j of the transformation imageThe calculation formula is:
[0046] wherein, Sampler(·) is a nearest neighbor sampling operator.
[0047] In step 104, a plurality of optimization regularization terms are determined according to the above-mentioned each sparse view and the above-mentioned pseudo view set.
[0048] In some embodiments, the above-mentioned execution subject can determine a plurality of optimization regularization terms according to the above-mentioned each sparse view and the above-mentioned pseudo view set.
[0049] Optionally, the above-mentioned execution subject can determine a plurality of optimization regularization terms according to the above-mentioned each sparse view and the above-mentioned pseudo view set by the following steps:
[0050] First, the pixel loss and the structure consistency loss of the ground truth image corresponding to each sparse view and the rendered image are determined.
[0051] As an example, the above-mentioned execution subject can determine the pixel loss and the structure consistency loss of the ground truth image corresponding to each sparse view and the rendered image by As an example, the above-mentioned execution subject can determine the pixel loss and the structure consistency loss of the ground truth image corresponding to each sparse view and the rendered image by represents the pixel loss. The above-mentioned represents the structure consistency loss. The above-mentioned represents the rendered image. The above-mentioned represents the ground truth image corresponding to each sparse view, i.e. the training image set.
[0052] Second, the sparse view monocular depth estimation of the ground truth image corresponding to each sparse view is determined by an external monocular depth prior.
[0053] As an example, the above-mentioned execution subject can determine the sparse view monocular depth estimation of the ground truth image corresponding to each sparse view by As an example, the above-mentioned execution subject can determine the sparse view monocular depth estimation of the ground truth image corresponding to each sparse view by represents the sparse view monocular depth estimation. The subscript m represents monocular.
[0054] Third, the sparse view depth Pearson correlation loss is determined according to the sparse view monocular depth estimation and the rendered depth corresponding to each sparse view.
[0055] As an example, the above-mentioned execution subject can determine the sparse view depth Pearson correlation loss according to the sparse view monocular depth estimation and the rendered depth corresponding to each sparse view by As an example, the above-mentioned execution subject can determine the sparse view depth Pearson correlation loss according to the sparse view monocular depth estimation and the rendered depth corresponding to each sparse view by represents the sparse view depth Pearson correlation loss. denotes the rendering depth corresponding to each sparse view. Wherein Corr(·) is the Pearson correlation operator, and the calculation formula is:
[0056] The D r denotes the rendering depth. The subscript r denotes render. The D m denotes monocular depth estimation.
[0057] The fourth step is to determine the pseudo-view monocular depth estimation corresponding to the pseudo-view set by the external monocular depth prior.
[0058] As an example, the execution subject can determine the pseudo-view monocular depth estimation corresponding to the pseudo-view set by determining the pseudo-view monocular depth estimation corresponding to the pseudo-view set by the external monocular depth prior. denotes the pseudo-view monocular depth estimation.
[0059] The fifth step is to determine the pseudo-view depth Pearson correlation loss according to the rendering depth corresponding to the pseudo-view set and the pseudo-view monocular depth estimation.
[0060] As an example, the execution subject can determine the pseudo-view depth Pearson correlation loss according to the rendering depth corresponding to the pseudo-view set and the pseudo-view monocular depth estimation by determining the pseudo-view depth Pearson correlation loss according to the rendering depth corresponding to the pseudo-view set and the pseudo-view monocular depth estimation. The D denotes the pseudo-view depth Pearson correlation loss. denotes the rendering depth corresponding to the pseudo-view set.
[0061] The sixth step is to determine the transformation image and the consistency mask of the ground true image corresponding to each sparse view to the pseudo-view set, and obtain the geometric consistency loss.
[0062] As an example, the execution subject can first determine the training image set corresponding to each sparse view to the pseudo-view v j The transformation image and the consistency mask M i→j of the pseudo-view. Then, the 1-norm loss of the mask is determined to obtain the geometric consistency loss, that is,
[0063]
[0064] The seventh step is to add noise to the rendering image according to the time step to obtain the noisy rendering image distribution.
[0065] As an example, the execution subject can add noise according to the time step t to the rendering image to obtain the noisy rendering image distribution The noisy rendering image distribution is subject to a Gaussian distribution: wherein, denotes the degenerate coefficient, I denotes the identity matrix. θ denotes the parameter of the 3D Gaussian point cloud. denotes the rendered image denotes the rendered image with noise added at time step t.
[0066] In the eighth step, a score matching loss guided by the inline prior is determined according to the above-mentioned rendered image distribution with noise.
[0067] As an example, the above-mentioned execution subject can guide the rendered image distribution with noise by introducing the inline prior to be revised as using the transformed images obtained from each sparse view to guide the sampling trajectory of , where φ denotes the parameter introduced by the construction of the revised distribution. denotes the rendered image with noise. The above-mentioned denotes the revised distribution. Therefore, the optimization target of the score matching guided by the inline prior can be expressed as:
[0068]
[0069] where ω(t) is a weight function about t, D KL is the KL divergence, η r is the adjustment coefficient of the two sub-optimization targets. denotes the expectation. In actual operation, the diffusion repair model is used to denote the revised distribution. The diffusion model is used to denote the prior distribution of the noisy diffusion model. denotes the diffusion prior distribution. Therefore, the gradient corresponding to the optimization target of the score matching guided by the inline prior can be obtained as:
[0070]
[0071] Here, the above-mentioned γ(t) denotes the above-mentioned obeys the diffusion prior distribution. Therefore, the score matching loss guided by the inline prior can be expressed as:
[0072]
[0073] In the ninth step, the above-mentioned pixel loss and structure consistency loss, the above-mentioned sparse view depth Pearson correlation loss, the above-mentioned pseudo view depth Pearson correlation loss, the above-mentioned geometric consistency loss, and the above-mentioned score matching loss guided by the inline prior are determined as multiple optimization regularization terms.
[0074] The first step to the ninth step above constructs a modified distribution as an intermediate state between the noisy diffusion model prior distribution and the noisy rendered image distribution, guides the mode search behavior of fractional distillation, controls the three-dimensional representation parameter optimization trajectory, effectively extracts the visual knowledge of the diffusion model prior, and thus improves the sparse view three-dimensional reconstruction quality.
[0075] In step 105, the loss function is determined according to the plurality of optimization regular terms.
[0076] In some embodiments, the execution subject can determine the loss function according to the plurality of optimization regular terms.
[0077] As an example, the execution subject can determine the loss function according to the plurality of optimization regular terms by the following formula: The λ above represents a regularization parameter.
[0078] In step 106, the parameters of the initialized three-dimensional Gaussian point cloud are optimized according to the loss function to perform sparse view three-dimensional reconstruction.
[0079] In some embodiments, the execution subject can optimize the parameters of the initialized three-dimensional Gaussian point cloud according to the loss function to perform sparse view three-dimensional reconstruction.
[0080] As an example, the execution subject can supervise each sparse view and the pseudo view set based on the loss function, optimize and update the three-dimensional Gaussian point cloud parameters through a differentiable rendering and gradient descent algorithm, and thus realize three-dimensional reconstruction under sparse view.
[0081] Further referring to Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a three-dimensional reconstruction method based on sparse view in-line prior, which device embodiments correspond to those method embodiments shown in Figure 1 The device can be applied in various electronic devices.
[0082] As Figure 2As shown, the three-dimensional reconstruction device 200 based on sparse view intra-link prior of some embodiments comprises a first determining unit 201, an initialization unit 202, a sampling unit 203, a second determining unit 204, a third determining unit 205 and an optimization unit 206. Among them, the first determining unit 201 is configured to determine a sparse point cloud of a three-dimensional scene from sparse view images; the initialization unit 202 is configured to initialize a three-dimensional Gaussian point cloud according to the above-mentioned sparse point cloud, to obtain an initialized three-dimensional Gaussian point cloud; the sampling unit 203 is configured to randomly sample within a preset range of each sparse view of the above-mentioned sparse view images to generate pseudo views, to obtain a pseudo view set; the second determining unit 204 is configured to determine a plurality of optimization regular terms according to the above-mentioned each sparse view and the above-mentioned pseudo view set; the third determining unit 205 is configured to determine a loss function according to the above-mentioned plurality of optimization regular terms; and the optimization unit 206 is configured to optimize parameters of the above-mentioned initialized three-dimensional Gaussian point cloud according to the above-mentioned loss function, to perform sparse view three-dimensional reconstruction.
[0083] It can be understood that the units recorded in the device 200 correspond to each step in the method described with reference to Figure 1 The above description of the operations, features and beneficial effects of the method also applies to the three-dimensional reconstruction device 200 based on sparse view intra-link prior and the units contained therein, and will not be repeated here.
Claims
1. A method for sparse-view based 3D reconstruction based on intra-view priors, comprising: determining a sparse point cloud of a 3D scene from sparse-view images; initializing a 3D Gaussian point cloud according to the sparse point cloud to obtain an initialized 3D Gaussian point cloud; randomly sampling a preset range of each of the sparse-view images to generate pseudo-views to obtain a set of pseudo-views; determining a plurality of optimization regular terms according to the sparse-view images and the set of pseudo-views, including: determining pixel loss and structure consistency loss of ground-truth images corresponding to the sparse-view images; determining sparse-view monocular depth estimation of the ground-truth images corresponding to the sparse-view images by an external monocular depth prior; determining sparse-view depth Pearson correlation loss according to the sparse-view monocular depth estimation and rendered depth corresponding to the sparse-view images; determining pseudo-view monocular depth estimation of the set of pseudo-views by an external monocular depth prior; determining pseudo-view depth Pearson correlation loss according to the rendered depth corresponding to the set of pseudo-views and the pseudo-view monocular depth estimation; determining transformation images and consistency masks of the ground-truth images corresponding to the sparse-view images to the set of pseudo-views to obtain geometric consistency loss; adding noise to the rendered images according to a time step to obtain a noisy rendered image distribution; determining score matching loss guided by intra-view priors according to the noisy rendered image distribution; and determining the pixel loss and the structure consistency loss, the sparse-view depth Pearson correlation loss, the pseudo-view depth Pearson correlation loss, the geometric consistency loss and the score matching loss guided by intra-view priors as the plurality of optimization regular terms; determining a loss function according to the plurality of optimization regular terms; and optimizing parameters of the initialized 3D Gaussian point cloud according to the loss function to perform sparse-view 3D reconstruction. The random sampling of the preset range of each of the sparse-view images to generate pseudo-views to obtain a set of pseudo-views comprises: setting a set of training images corresponding to each of the sparse-view images; constructing a transformation function; and inversely transforming the set of training images from each of the sparse-view images to the set of pseudo-views by the transformation function. After the initialization of the 3D Gaussian point cloud according to the sparse point cloud to obtain the initialized 3D Gaussian point cloud, the method further comprises: rendering the initialized 3D Gaussian point cloud to obtain a 2D image; sorting the Gaussian point cloud corresponding to the 2D image to obtain a sorted Gaussian point cloud; and rendering the sorted Gaussian point cloud to obtain a rendered image corresponding color. 4.An apparatus for sparse-view based 3D reconstruction based on intra-view priors, comprising: a first determining unit configured to determine a sparse point cloud of a 3D scene from sparse-view images; an initializing unit configured to initialize a 3D Gaussian point cloud according to the sparse point cloud to obtain an initialized 3D Gaussian point cloud; 2. The method of claim 1, wherein, 3. The method of claim 1, wherein, a sampling unit configured to randomly sample a preset range of each of the sparse view images to generate pseudo views, to obtain a set of pseudo views; a second determining unit configured to determine a plurality of optimization regular terms according to the sparse view images and the set of pseudo views, including: determining pixel loss and structure consistency loss of ground truth images corresponding to the sparse view images and rendered images; determining sparse view monocular depth estimation of the ground truth images corresponding to the sparse view images through external monocular depth prior; determining sparse view depth Pearson correlation loss according to the sparse view monocular depth estimation and rendered depth corresponding to the sparse view images; determining pseudo view monocular depth estimation corresponding to the set of pseudo views through external monocular depth prior; determining pseudo view depth Pearson correlation loss according to rendered depth corresponding to the set of pseudo views and the pseudo view monocular depth estimation; determining geometric consistency loss of transformed images and consistency masks of the ground truth images corresponding to the sparse view images to the set of pseudo views; adding noise to the rendered images according to time steps to obtain a noisy rendered image distribution; determining score matching loss guided by inline prior according to the noisy rendered image distribution; and determining the pixel loss and the structure consistency loss, the sparse view depth Pearson correlation loss, the pseudo view depth Pearson correlation loss, the geometric consistency loss and the score matching loss guided by the inline prior as the plurality of optimization regular terms; a third determining unit configured to determine a loss function according to the plurality of optimization regular terms; an optimization unit configured to optimize parameters of the initialized three-dimensional Gaussian point cloud according to the loss function, to perform sparse view three-dimensional reconstruction. 5.An electronic device, comprising: one or more processors; a storage having one or more programs stored thereon; when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-3.
6. A computer readable medium having stored thereon a computer program, wherein, The program is executed by the processor to implement the method of any one of claims 1-3.
Citation Information
Patent Citations
Three-dimensional scene reconstruction method and electronic equipment
CN118365805A
A limited-angle CT reconstruction method based on anisotropic total variation
US20200402274A1