3D reconstruction method for outdoor boundless scenes based on multi-scale features and deep supervised neural radiance field

Through improved multi-scale features and deeply supervised neural radiation field method, the problem of reconstruction fuzzy in dense boundless scenarios is solved, and a high-quality three-dimensional surface model is generated, which improves reconstruction accuracy and adaptability.

CN120431275BActive Publication Date: 2025-09-02QINGDAO UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510947074.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-02
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

The existing neural radiation field methods are difficult to effectively distinguish the foreground and background information in dense boundless scenarios, resulting in blurred reconstruction geometric structures and large calculations, making it difficult to apply on a large scale.

Method used

Using a neural radiation field method based on multi-scale features and deep supervision, the improved feature pyramid network and Depth Anything V2 model is used, combining sparse point clouds and camera parameters to generate high-quality three-dimensional reconstruction results, and the neural radiation field network is optimized using photometric and depth loss functions, and finally the surface three-dimensional model is generated through the explicit processing of the occupancy field.

Benefits of technology

It improves the reconstruction accuracy and adaptability of NeRF in dense scenes, realizes high-quality three-dimensional surface model generation, and provides technical support for applications such as low-altitude flight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431275B_ABST
    Figure CN120431275B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for three-dimensional reconstruction of outdoor borderless scenes based on multi-scale features and deep supervised neural radiation fields, which belongs to the field of computer vision technology. Specifically, the method comprises the following steps: obtaining an image dataset of an outdoor scene, extracting the corresponding multi-scale feature map using an improved feature pyramid network, and obtaining an estimated depth map using a Depth Anything V2 model; combining the source image, the extracted feature map, and the estimated depth map, inputting the neural radiation field network for training, and finally extracting the occupancy field of the scene from the optimized neural radiation field, and obtaining a curved three-dimensional model through explicit processing. The beneficial effects of the present invention are: guiding NeRF to learn correct geometric information through a high-precision depth map, and utilizing multi-scale features as scene priors, thereby effectively improving the adaptability and geometric robustness of NeRF in a borderless environment, achieving accurate reconstruction of outdoor scenes, and providing technical support for applications such as low-altitude flight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method for outdoor borderless scenes based on multi-scale features and deep supervised neural radiation fields, which is particularly suitable for dense scenes in outdoor borderless environments. Background Art

[0002] 3D reconstruction technology holds significant application value in smart agriculture, smart cities, and autonomous driving. In agricultural scenarios, high-precision 3D reconstruction can intuitively reflect crop growth status and provide crucial support for structured crop detection, phenotyping, and smart agricultural management. Current 3D reconstruction methods primarily rely on depth sensors or 2D image data. Depth sensors rely on LiDAR, structured light, or Time of Flight (ToF) technology to acquire depth information. While highly accurate, these sensors are expensive and difficult to deploy on a large scale, limiting their application in boundless environments such as vast farmlands. Traditional 3D reconstruction methods based on 2D images are exemplified by Multi-View Stereo (MVS). This method primarily computes matching relationships between multiple viewpoints, estimates camera motion, and then reconstructs sparse and dense point clouds to achieve 3D reconstruction of the scene. However, the MVS method suffers from a complex computational process and is susceptible to factors such as viewpoint and lighting. This method is prone to reconstruction errors in highly complex natural environments with highly similar textures (e.g., dense scenes containing numerous similar elements), resulting in reduced model integrity and accuracy.

[0003] In recent years, the Neural Radiance Fields (NeRF) method has gained popularity. NeRF uses a deep learning model to learn an implicit 3D representation of a scene and employs volume rendering technology to generate new perspective images, achieving high-precision 3D reconstruction. NeRF has attracted widespread attention in the field of computer vision due to its end-to-end optimization. However, NeRF still faces many challenges when applied to dense outdoor natural scenes. In such dense scenes, because objects are closely arranged and severely occlude each other, NeRF struggles to effectively distinguish foreground and background information during modeling, resulting in blurred reconstructed geometric structures and poor reconstruction quality. Furthermore, the NeRF model is computationally intensive, limiting its application in large-scale scenes. Summary of the Invention

[0004] In response to the current problems of geometric modeling ambiguity, depth estimation bias and missing explicit models in dense and boundless scenes caused by neural radiation fields, this paper proposes a three-dimensional reconstruction method for outdoor boundless scenes based on multi-scale features and deep supervision of neural radiation fields. A method based on multi-scale feature extraction is used to solve the problem of being unable to distinguish foreground and background information, and achieve high-quality NeRF reconstruction.

[0005] In order to achieve the above-mentioned object of the invention, the present invention provides a 3D reconstruction method for outdoor boundless scenes based on multi-scale features and deep supervised neural radiation fields, the 3D reconstruction method comprising the following steps:

[0006] Step S1: Acquire a multi-view image dataset of an outdoor scene;

[0007] Step S2: Based on the improved feature pyramid network, a preliminary feature map is extracted from the source image in the image dataset, and the Depth Anything V2 model is used to generate an estimated depth map D1;

[0008] Step S3: constructing a feature pyramid containing multiple feature maps of different resolutions based on the preliminary feature map, and upsampling and fusing them layer by layer through a bilinear interpolation method to obtain a fused feature map that is aligned with the pixel space of the source image;

[0009] Step S4: using COLMAP to obtain the camera parameters and sparse point cloud of the source image, and taking the source image, fused feature map, camera parameters and sparse point cloud as input to construct a scene neural radiance field network;

[0010] Step S5: uniformly sampling the rays generated by the camera parameters, encoding the obtained sampling point positions and directions into high-dimensional feature vectors, inputting them together with the fused feature map into the scene neural radiation network, outputting the color value and density value of the sampling point, performing volume rendering integration on the color value and density value of the sampling point along each ray, and calculating the predicted color value of the pixel corresponding to the ray;

[0011] Step S6: sampling the rays generated by the camera parameters in three-dimensional space, and inferring the depth value of each ray through cumulative calculation based on the color value and density value of the voxel, thereby generating a predicted depth map D2, and using the estimated depth map D1 of step S2 to perform geometric constraints on the predicted depth map D2;

[0012] Step S7: Jointly optimizing the scene neural radiance field network based on the photometric loss between the predicted color value obtained in step S5 and the true color value of the source image, and the depth loss between the estimated depth map D1 and the predicted depth map D2;

[0013] Step S8: Generate an implicit scene expression based on the optimized scene neural radiation field network, and obtain a surface three-dimensional model through occupancy field explicit processing.

[0014] The step S1 specifically includes: recording a video around a target scene in an outdoor borderless environment using a free trajectory, and performing interval frame extraction on the video to remove the blurred image and retain the N valid source images to form an image dataset.

[0015] In step S2, the improved feature pyramid network is improved based on the ResNet50 network architecture, which includes a convolutional layer group from Conv1 to Conv5, and does not include the last fully connected layer for image classification, and finally outputs an initial feature map.

[0016] In step S5, the position and direction of the sampling point are encoded into a high-dimensional vector , specifically:

[0017] ;

[0018] in, Indicates the position of the sampling point on a certain axis, Indicates the number of encoding layers, which is used to determine the dimension of the encoded vector. is the original 3D coordinate of the sampling point.

[0019] In step S5, the predicted color value of the pixel corresponding to the ray is calculated. The rendering formula is:

[0020] ;

[0021] Where m and n are the near and far boundaries of the ray, respectively. is the density value of the ray at the sampling point; is the transmittance; is the color value on the ray.

[0022] The photometric loss between the predicted color value and the true color value of the source image in step S7 The calculation function is:

[0023] ;

[0024] in, C is the true color value of the source image, R is the set of rays passing through the image pixels.

[0025] In step S7, the depth loss between the initial depth map D1 and the predicted depth map D2 is The calculation function is:

[0026] ;

[0027] in, For the image The estimated depth map estimated using the Depth Anything V2 model, image The predicted depth map D2.

[0028] In step S7, the scene neural radiance field network is optimized by combining the total loss function of the photometric loss and the difference loss. for:

[0029] ;

[0030] in, is the control factor.

[0031] In step S8, the surface three-dimensional model is obtained by performing occupancy field explicit processing, specifically:

[0032] (1) Extracting the density value of each sampling point in the scene from the scene neural radiance field network, sampling the density value in three-dimensional space and discretizing it into a three-dimensional grid structure, where each grid unit corresponds to a voxel. The density value is mapped to the corresponding voxel according to the spatial position to represent the spatial occupancy distribution of the scene;

[0033] (2) classifying voxels in the three-dimensional grid structure according to a preset density threshold, wherein voxels with a density value greater than the density threshold are marked as "occupied", and voxels with a density value less than or equal to the density threshold are marked as "free";

[0034] (3) Based on the binarized occupancy field, a surface extraction algorithm is used to convert the adjacent boundaries of the "occupied" and "idle" state voxels into triangulated surface meshes to obtain the overall surface three-dimensional model of the object.

[0035] The beneficial effects of the present invention are as follows: Based on NeRF's precise implicit representation of scene information, combined with an improved feature extraction network and feature pyramid structure, the present invention enhances NeRF's ability to express dense scene features, and utilizes the Depth Anything V2 model to estimate scene depth information, thereby effectively improving NeRF's adaptability in unbounded environments. An improved upsampling method improves NeRF's accuracy in detail reconstruction in dense scenes; combined with an occupancy field explicitization algorithm, the implicit representation generated by NeRF is converted into a high-quality three-dimensional surface model, achieving accurate reconstruction of large-scale dense scenes and providing technical support for applications such as low-altitude flight. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flow chart of Example 1 and Example 2 of the present invention.

[0037] Figure 2 This is a network architecture diagram of embodiments 1 and 2 of the present invention.

[0038] Figure 3Schematic diagram of obtaining an image dataset in Example 2 of the present invention.

[0039] Figure 4 This is a schematic diagram of a depth map generated by the Depth Anything V2 model in Example 2 of the present invention.

[0040] Figure 5 This is the ResNet50 structure diagram in Example 2 of the present invention.

[0041] Figure 6 This is a specific flow chart of feature extraction and upsampling in Example 2 of the present invention.

[0042] Figure 7 Schematic diagram of the bilinear interpolation method in Example 2 of the present invention.

[0043] Figure 8 This is a schematic diagram of the result model in Example 2 of the present invention. DETAILED DESCRIPTION

[0044] In order to clearly illustrate the technical features of this solution, this solution is described below through specific implementation methods.

[0045] Example 1: See Figure 1 and Figure 2 , an embodiment of the present invention provides a method for 3D reconstruction of outdoor boundless scenes based on multi-scale features and deep supervised neural radiation fields, specifically:

[0046] Step S1: Obtain a multi-view image dataset of an outdoor scene; specifically, around the target scene in an outdoor boundless environment, use a free trajectory to record the video, perform interval frame extraction on the video, remove the blurred image and retain the N valid source images to form an image dataset.

[0047] Step S2: Based on the improved feature pyramid network, a preliminary feature map is extracted from the source image in the image dataset, and an estimated depth map D1 is generated using the Depth Anything V2 model; wherein the improved feature pyramid network is based on the ResNet50 network architecture, which includes a convolutional layer group from Conv1 to Conv5, and does not include the last fully connected layer for image classification, and finally outputs the initial feature map.

[0048] Step S3: constructing a feature pyramid containing multiple feature maps of different resolutions based on the preliminary feature map, and upsampling and fusing them layer by layer through a bilinear interpolation method to obtain a fused feature map that is aligned with the pixel space of the source image.

[0049] Step S4: Use COLMAP to obtain the camera parameters and sparse point cloud of the source image, and use the source image, fused feature map, camera parameters and sparse point cloud as input to construct a scene neural radiation field network.

[0050] Step S5: uniformly sample the rays generated by the camera parameters, encode the obtained sampling point positions and directions into high-dimensional feature vectors, input them together with the fusion feature map into the scene neural radiation network, output the color value and density value of the sampling point, perform volume rendering integration on the color value and density value of the sampling point along each ray, and calculate the predicted color value of the pixel corresponding to the ray; wherein, the position and direction of the sampling point are encoded as high-dimensional vectors , specifically:

[0051] ;

[0052] in, Indicates the position of the sampling point on a certain axis, Indicates the number of encoding layers, which is used to determine the dimension of the encoded vector. is the original 3D coordinate of the sampling point.

[0053] Among them, it is used to calculate the predicted color value of the pixel corresponding to the ray The rendering formula is:

[0054] ;

[0055] Where m and n are the near and far boundaries of the ray, respectively. is the density value of the ray at the sampling point; is the transmittance; is the color value on the ray.

[0056] Step S6: extracting feature maps of each level from the feature pyramid, and calculating the depth value and depth weighted confidence corresponding to each ray;

[0057] Step S7: jointly optimizing the scene neural radiance field network based on the photometric loss between the predicted color value and the true color value of the source image, and the difference loss between the depth value and the depth weighted reset confidence;

[0058] Among them, the photometric loss between the predicted color value and the true color value of the source image The calculation function is:

[0059] ;

[0060] Where, is the predicted color value, C is the true color value of the source image, R is the set of rays passing through the image pixels.

[0061] The depth loss between the initial depth map D1 and the predicted depth map D2 The calculation function is:

[0062] ;

[0063] in, For the image The estimated depth map estimated using the Depth Anything V2 model, image The predicted depth map D2.

[0064] The scene neural radiance field network is optimized by combining the total loss function of the photometric loss and the difference loss. for:

[0065] ;

[0066] in, is the control factor.

[0067] Step S8: Generate an implicit scene representation based on the optimized scene neural radiance field network, and obtain a curved surface 3D model by explicit processing of the occupancy field. Specifically, the curved surface 3D model obtained by explicit processing of the occupancy field is:

[0068] (1) Extracting the density value of each sampling point in the scene from the scene neural radiance field network, sampling the density value in three-dimensional space and discretizing it into a three-dimensional grid structure, where each grid unit corresponds to a voxel, and mapping the density value to the corresponding voxel according to the spatial position to represent the spatial occupancy distribution of the scene;

[0069] (2) classifying voxels in the three-dimensional grid structure according to a preset density threshold, wherein voxels with a density value greater than the density threshold are marked as "occupied", and voxels with a density value less than or equal to the density threshold are marked as "free";

[0070] (3) Based on the binarized occupancy field, a surface extraction algorithm is used to convert the adjacent boundaries of the "occupied" and "idle" state voxels into triangulated surface meshes to obtain the overall surface three-dimensional model of the object.

[0071] Example 2: Figure 1 and Figure 2 As shown, an embodiment of the present invention provides a method for 3D reconstruction of outdoor boundless scenes based on multi-scale features and deep supervised neural radiation fields, specifically including:

[0072] S1: Make the experimental source image dataset: Figure 3As shown in the figure, a video is recorded with a free trajectory around the target scene in an outdoor boundaryless environment, and the video is processed by interval frame extraction to obtain RGB source images, and the blurred images caused by shooting problems are removed. Each remaining RGB source image is recorded as , forming a training data set, where Indicates the number of source images that are finally retained;

[0073] S2: Input the source image into the pre-trained improved feature pyramid network to generate the initial feature map corresponding to the source image. Then, the extracted series of initial feature maps are used to construct a feature pyramid through step S3 to obtain multi-scale feature information, and the Depth Anything V2 model is used to generate the corresponding estimated depth map D1 (such as Figure 4 ). Among them, see Figure 5 , the feature pyramid network is improved based on the ResNet50 network, and the ResNet50 network is used as the basis of the feature extractor. The original ResNet50 consists of a 50-layer convolutional structure, which includes multiple residual modules, which can effectively capture the feature expressions of different scales in the image. This embodiment adjusts the ResNet50 and deletes the last fully connected layer responsible for image classification in the original network, making it a pure convolutional structure to meet the feature extraction requirements related to spatial position. A feature pyramid structure is added to the modified ResNet50 network to extract multi-scale features. The specific implementation steps are as follows:

[0074] First, the input image i After the initial convolution and pooling, it is fed into a series of convolutional layers (Conv1 to Conv5 in ResNet50), and the formula is:

[0075] ;

[0076] in, Represents the feature map of the layer output. After each convolution stage, feature maps of different scales are obtained and recorded as , the spatial size of each layer decreases successively, and the number of channels increases gradually. When constructing the feature pyramid, the deepest feature map The number of channels is adjusted through convolution, and then the bilinear interpolation method is used to upsample to the previous scale (such as the Conv4 scale), and then fused with the Conv4 convolution feature; recursive upsampling and fusion are performed in this way to finally obtain a multi-scale fused feature map with a resolution close to the original image.

[0077] S3: If Figure 6As shown in Figure 1, feature pyramid construction and NeRF ray generation. The multi-scale feature maps extracted by ResNet50 are constructed into a feature pyramid structure (FPN), and the bilinear interpolation method is used to upsample and fuse layer by layer to obtain a fused feature map with a size close to that of the input image. Figure 7 As shown, the feature extraction network is upsampled to improve the traditional transposed convolution, and bilinear interpolation is used to expand the resolution of the output feature map. Assuming that it is known The values ​​of four points need to be solved for the unknown function At the point The value at Direction, yes and Interpolate two points and get the values ​​at two places:

[0078] ;

[0079] ;

[0080] in, , , and then in the y direction according to and Can be solved point:

[0081] Combining the above results, the final result of bilinear interpolation for:

[0082] ;

[0083] This method uses bilinear interpolation at different resolution levels to increase the resolution of feature maps. The processed images are then fused and stitched together into a large feature map, which corresponds one-to-one with each pixel in the input image, ensuring pixel-by-pixel alignment of image features with the original image. This avoids the checkerboard effect caused by increasing the resolution of the feature map, maximizes the preservation of the original information, and recovers more feature information.

[0084] The upsampled fusion feature map is concatenated with the position-encoded sampling point position and line-of-sight direction information to form an input feature vector. The NeRF network uses the following position encoding method:

[0085] ;

[0086] in, Indicates the position of the sampling point on a certain axis, Indicates the number of encoding layers, which determines the dimension of the encoded vector; therefore, for a position The input, after encoding, the position is represented by multiple dimensions to represent the original three-dimensional coordinate vector;

[0087] Among them, the position of the encoded sampling point is put into the NeRF network to predict the color and density of each sampling point. The prediction formula is:

[0088] ;

[0089] S4: As Figure 2 As shown, when training the NeRF model, a loss function is defined to guide model learning and measure the generated image With real images The differences between:

[0090] ;

[0091] By optimizing the rendering loss, the model can gradually learn to generate rendering results that are closer to the real image, thereby improving the accuracy and realism of scene reconstruction;

[0092] The depth loss is based on the mean square error and is set as:

[0093] ;

[0094] in, Refers to the image , Refers to the image Depth map estimated using the Depth Anything V2 model.

[0095] The total loss function is: ;

[0096] Among them, Control factor.

[0097] S5: Visualize the implicit 3D representation of NeRF output via occupancy field techniques (e.g. Figure 8 ), including the following steps:

[0098] (1) Constructing the occupation field:

[0099] The NeRF model extracts density values ​​(i.e., occupancy probabilities) for each point in the scene and discretizes these values ​​into a 3D grid. Each grid cell corresponds to a voxel, and the density values ​​are mapped to the corresponding cells to represent the spatial distribution of objects.

[0100] (2) Threshold segmentation:

[0101] Set a predefined threshold to classify grid cells. Cells with density values ​​exceeding the threshold are marked as "occupied", while cells below the threshold are considered "vacant";

[0102] (3) Surface extraction:

[0103] A surface extraction algorithm is used to extract the object surface from the thresholded occupancy field. Specifically, the boundaries between "occupied" and "free" cells are converted into a triangular mesh to generate a surface representation of the object.

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A 3D reconstruction method for outdoor boundless scenes based on multi-scale features and deep supervised neural radiation fields, characterized by: The three-dimensional reconstruction method comprises the following steps: Step S1: Acquire a multi-view image dataset of an outdoor scene; Step S2: Based on the improved feature pyramid network, a preliminary feature map is extracted from the source image in the image dataset, and the Depth Anything V2 model is used to generate an estimated depth map D1; Step S3: constructing a feature pyramid containing multiple feature maps of different resolutions based on the preliminary feature map, and upsampling and fusing them layer by layer through a bilinear interpolation method to obtain a fused feature map that is aligned with the pixel space of the source image; Step S4: using COLMAP to obtain the camera parameters and sparse point cloud of the source image, and taking the source image, fused feature map, camera parameters and sparse point cloud as input to construct a scene neural radiance field network; Step S5: uniformly sampling the rays generated by the camera parameters, encoding the obtained sampling point positions and directions into high-dimensional feature vectors, inputting them together with the fused feature map into the scene neural radiation network, outputting the color value and density value of the sampling point, performing volume rendering integration on the color value and density value of the sampling point along each ray, and calculating the predicted color value of the pixel corresponding to the ray; Step S6: sampling the rays generated by the camera parameters in three-dimensional space, and inferring the depth value of each ray through cumulative calculation based on the color value and density value of the voxel, thereby generating a predicted depth map D2, and using the estimated depth map D1 of step S2 to perform geometric constraints on the predicted depth map D2; Step S7: Jointly optimizing the scene neural radiance field network based on the photometric loss between the predicted color value obtained in step S5 and the true color value of the source image, and the depth loss between the estimated depth map D1 and the predicted depth map D2; Step S8: Generate an implicit scene expression based on the optimized scene neural radiation field network, and obtain a surface three-dimensional model through occupancy field explicit processing.

2. The three-dimensional reconstruction method according to claim 1, characterized in that: The step S1 is specifically as follows: around the target scene in the outdoor borderless environment, a video is recorded using a free trajectory, and the video is processed by interval frame extraction, and the blurred image is removed and the remaining N valid source images to form an image dataset.

3. The three-dimensional reconstruction method according to claim 1, wherein: In step S2, the improved feature pyramid network is improved based on the ResNet50 network architecture, which includes a convolutional layer group from Conv1 to Conv5, and does not include the last fully connected layer for image classification, and finally outputs an initial feature map.

4. The three-dimensional reconstruction method according to claim 1, characterized in that: In step S5, the position and direction of the sampling point are encoded into a high-dimensional vector , specifically: ; in, Indicates the position of the sampling point on a certain axis, Indicates the number of encoding layers, which is used to determine the dimension of the encoded vector. is the original 3D coordinate of the sampling point.

5. The three-dimensional reconstruction method according to claim 1, characterized in that: In step S5, the predicted color value of the pixel corresponding to the ray is calculated. The rendering formula is: ; Where m and n are the near and far boundaries of the ray, respectively. is the density value of the ray at the sampling point; is the transmittance; is the color value on the ray.

6. The three-dimensional reconstruction method according to claim 1, characterized in that: The photometric loss between the predicted color value and the true color value of the source image in step S7 The calculation function is: ; in, C is the true color value of the source image, R is the set of rays passing through the image pixels.

7. The three-dimensional reconstruction method according to claim 1, characterized in that: In step S7, the depth loss between the initial depth map D1 and the predicted depth map D2 is The calculation function is: ; in, For the image The estimated depth map estimated using the Depth Anything V2 model, image The predicted depth map D2.

8. The three-dimensional reconstruction method according to claim 6 or 7, characterized in that: In step S7, the scene neural radiance field network is optimized by combining the total loss function of the photometric loss and the difference loss. for: ; in, is the control factor.

9. The three-dimensional reconstruction method according to claim 1, characterized in that: In step S8, the surface three-dimensional model is obtained by performing occupancy field explicit processing, specifically: (1) Extract the density value of each sampling point in the scene from the scene neural radiation field network, sample the density value in three-dimensional space and discretize it into a three-dimensional grid structure; (2) Classifying voxels in the three-dimensional grid structure according to a preset density threshold, wherein voxels with a density value greater than the density threshold are marked as "occupied", and voxels with a density value less than or equal to the density threshold are marked as "free"; (3) Based on the binarized occupancy field, a surface extraction algorithm is used to convert the adjacent boundaries of the "occupied" and "idle" state voxels into triangulated surface meshes to obtain the overall surface three-dimensional model of the object.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method without prior pose input

    CN118196298A

  • Non-feature extraction-based dense SFM three-dimensional reconstruction method

    WO2015154601A1