Monocular depth guided object level NeRF reconstruction method

Through the monocular depth-guided object-level NeRF reconstruction method, using self-supervised depth estimation and geometric enhancement technology, the problem of object-level reconstruction under monocular vision is solved, high-precision and controllable object 3D model generation is achieved, and object posture and shape editing is supported.

CN120655824APending Publication Date: 2025-09-16UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510738825.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing NeRF methods have difficulty in accurately recovering the three-dimensional structure of objects from monocular image sequences, and traditional methods have difficulty in object-level reconstruction under monocular vision, making it difficult to achieve independent modeling and control of specific objects.

Method used

A monocular depth-guided object-level NeRF reconstruction method is adopted. Through self-supervised depth estimation, relative pose estimation, object segmentation and geometric enhancement of NeRF model, combined with multi-resolution hash position encoding and spherical harmonic direction encoding, the NeRF training process is optimized to achieve end-to-end object-level controllable reconstruction.

Benefits of technology

It significantly improves the accuracy and controllability of object-level reconstruction in monocular videos, generates high-quality controllable 3D models, can accurately restore the geometry and appearance of objects, and supports controllable editing of object posture and shape.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655824A_ABST
    Figure CN120655824A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of three-dimensional reconstruction, and discloses a monocular depth guided object-level NeRF reconstruction method, which comprises the following steps: firstly, through monocular video sequence input, generating a frame-by-frame initial depth map by using a depth estimation module, and estimating a relative camera attitude between adjacent frames through a relative attitude estimation module; calculating the absolute attitude of the camera in combination with the initial depth map and the relative camera attitude; then constructing a geometrically enhanced NeRF model, and optimizing scene representation through multi-resolution hash position coding and spherical harmonic direction coding; introducing photometric loss, depth contrast loss and density loss to jointly optimize parameters of the NeRF model, and constraining geometric reconstruction of the object by using depth information; and finally, through a four-stage iterative training strategy, alternately optimizing depth estimation, a camera attitude and a NeRF model, and generating an object-level controllable three-dimensional model. According to the method, depth information is fully utilized to optimize the NeRF training process, so that end-to-end reconstruction from a monocular video to an object-level controllable model is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and in particular to a monocular depth-guided object-level NeRF reconstruction method. Background Art

[0002] With the rapid development of modern science and technology, computer vision, as a crucial component of artificial intelligence, has become a core area of ​​academic research and industrial application. In the field of monocular vision, breakthroughs in depth estimation and 3D reconstruction have opened up new possibilities for deriving spatial information from a single image. These technologies play an irreplaceable role in areas such as autonomous driving and robotic navigation, and also lay the technical foundation for cutting-edge applications such as virtual reality and augmented reality.

[0003] However, in applications such as virtual reality, augmented reality, and game development, simply recovering the overall three-dimensional depth information of the scene is far from sufficient to meet the needs. Users often need to independently reconstruct specific objects in the scene and achieve controllable editing of their posture, shape, or texture. This demand for object-level controllability has prompted research to turn to object-level three-dimensional reconstruction technology based on monocular input. Traditional explicit representation methods (such as point clouds and meshes) face significant challenges in object-level reconstruction: they have difficulty in accurately segmenting specific objects from complex scenes, and the editing process is complex and non-intuitive, which limits controllability.

[0004] In the field of 3D reconstruction, traditional multi-view geometry methods are gradually being replaced by innovative techniques based on deep learning. In recent years, Neural Radiance Field (NeRF), based on implicit neural representations, has attracted widespread attention due to its outstanding performance in synthesizing new viewpoints and high-precision geometric representation. NeRF uses a multi-layer perceptron implicit representation to learn the continuous 3D structure of a scene from 2D images. This not only generates high-quality 3D models but also supports synthetic rendering from new viewpoints, opening up new possibilities for object-level reconstruction and controllable editing. Improved NeRF-based methods have achieved significant progress in rendering efficiency, geometric accuracy, and controllability, enabling users to manipulate viewpoints, lighting, and even object properties. However, most existing NeRF methods reconstruct the entire scene, making it difficult to independently model and control specific objects. Furthermore, traditional NeRF relies on multi-view image input. In monocular video scenarios, the lack of multi-view information significantly complicates object-level reconstruction. Therefore, applying such techniques to monocular object-level spatial computation remains challenging. For example, how to accurately recover the 3D structure of objects from monocular image sequences and how to incorporate depth information to improve reconstruction quality remain pressing issues. Summary of the Invention

[0005] To address the above issues, the present invention aims to provide a monocular depth-guided object-level NeRF reconstruction method. This method, with monocular self-supervised depth estimation at its core, fully utilizes depth information to optimize the NeRF training process, thereby achieving end-to-end reconstruction from monocular video to an object-level controllable model. The technical solution is as follows:

[0006] A monocular depth-guided object-level NeRF reconstruction method includes the following steps:

[0007] Step 1: Using a monocular video sequence as input, the self-supervised depth estimation module generates a frame-by-frame depth map, and the relative pose estimation module estimates the relative camera pose between adjacent frames.

[0008] Step 2: Extract the target object mask based on the object segmentation module, and calculate the absolute pose of the camera by combining the depth map and the relative camera pose;

[0009] Step 3: Construct a geometrically enhanced NeRF model, taking the absolute pose as input and optimizing the scene representation through multi-resolution hash position encoding and spherical harmonic direction encoding;

[0010] Step 4: Introduce photometric loss, depth contrast loss, and density loss to jointly optimize NeRF parameters and use depth information to constrain object geometry reconstruction;

[0011] Step 5: Through a four-stage iterative training strategy, alternately optimize the depth estimation, camera pose and NeRF model to generate an object-level controllable 3D model.

[0012] The beneficial effects of the present invention are:

[0013] This invention takes monocular self-supervised depth estimation as the core, fully utilizes depth information to optimize the NeRF training process, and thus completes end-to-end reconstruction from monocular video to object-level controllable model; by integrating monocular self-supervised depth estimation, object segmentation and geometric enhancement NeRF, using four-stage iterative training and depth contrast consistency regularization, high-quality and controllable object three-dimensional model is reconstructed from monocular video, significantly improving the reconstruction accuracy and completeness. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 An object-level implicit neural representation and reconstruction framework for 3D controllable generation.

[0015] Figure 2 It is an MLP (Multilayer Perceptron) network structure.

[0016] Figure 3 is the classification of the sampling light.

[0017] Figure 4This is an example of the Cube-Diorama dataset.

[0018] Figure 5 This is an example of the Replica dataset.

[0019] Figure 6 Object reconstruction results on the Cube-Diorama dataset.

[0020] Figure 7 Reconstruction results for objects on the Replica dataset.

[0021] Figure 8 The following is a graph showing the in-depth visualization results of the first and third stages on the Replica dataset. DETAILED DESCRIPTION

[0022] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] This paper proposes an object-level NeRF reconstruction method based on monocular depth guidance, which aims to reconstruct the three-dimensional model of a specific object from a monocular video and realize controllable editing of its geometry and appearance. The framework takes monocular self-supervised depth estimation as the core, makes full use of depth information to optimize the NeRF training process, and thus completes the end-to-end reconstruction from monocular video to object-level controllable model. Depth information is particularly critical in monocular scenes. It not only provides geometric constraints for NeRF, but also improves the accuracy of object boundaries and the robustness of geometric shapes through depth contrast consistency regularization. First, the depth map of each frame is estimated from the monocular video to initialize a rough three-dimensional model of the object, and the depth information is integrated into the NeRF training. By introducing the depth contrast loss as a regularization term, the reliability of the relative depth relationship is emphasized, and the detail performance and reconstruction accuracy of the model are significantly enhanced. At the same time, combined with advanced object segmentation technology (such as SAMv2), the framework can focus on the target object and block background interference, further improving the accuracy of object-level modeling. This depth-guided strategy not only optimizes the reconstruction quality, but also lays a solid foundation for subsequent controllable editing (such as adjusting object pose), realizing the complete process from monocular video to high-quality, controllable 3D models.

[0024] 1. Overall framework of the method

[0025] The monocular depth-guided object-level NeRF reconstruction framework proposed in this paper is Figure 1As shown in the figure, the entire framework consists of four key modules: depth estimation, pose estimation, object segmentation, and NeRF-based object-level reconstruction. These modules work together, with depth information as the core constraint, to ensure the accuracy and controllability of the reconstructed model. First, a monocular self-supervised depth estimation network is used to generate a depth map for the input monocular video sequence frame by frame. At the same time, the relative camera pose between adjacent frames is calculated through the pose estimation network. These depth maps and pose information provide key geometric priors for subsequent object-level reconstruction. In order to achieve accurate modeling at the object level, the SegmentAnythingv2 (SAMv2) model is introduced to perform instance segmentation on each frame of the image and generate a mask of the target object. With its advanced video segmentation capabilities, SAMv2 can accurately extract object boundaries and provide clear object area constraints for NeRF training.

[0026] Based on the depth map, camera pose, and object mask, the absolute camera pose is further calculated. Specifically, the estimated relative pose and camera intrinsic parameters are used to determine the absolute position and orientation of the camera in the world coordinate system through triangulation. This process converts the temporal information of the monocular video into globally consistent camera parameters, providing the necessary input for NeRF's object-level reconstruction. Depth information acts as a bridge, not only constraining the optimization of the camera pose but also providing a reliable supervisory signal for NeRF's geometric learning.

[0027] In the NeRF-based object-level reconstruction module, an efficient geometrically enhanced neural radiance field model is designed. This model improves representation capabilities and rendering efficiency through multi-resolution hash position encoding and spherical harmonic direction encoding. Multi-resolution hash position encoding divides the three-dimensional space into multi-scale grids and assigns learnable features to each grid vertex, thereby finely depicting the geometric details of the object; spherical harmonic direction encoding uses spherical harmonic functions to represent the direction of light, enhancing the ability to model the appearance of the object (such as lighting and material). This enhanced NeRF is guided by depth maps and object masks to ensure that the training process focuses on the target object and avoids background interference.

[0028] By organically combining depth estimation, pose estimation, object segmentation, and NeRF-based object-level reconstruction, this framework fully leverages monocular depth information to optimize NeRF training, improving the accuracy and convergence speed of object-level reconstruction. The depth-guided strategy further constrains object geometry through depth contrast consistency regularization, ensuring boundary clarity and shape accuracy of the reconstructed model. Furthermore, the efficient geometrically enhanced NeRF model lays the foundation for controllable editing, allowing users to precisely control the object's pose and shape by adjusting the spatial position encoding or orientation encoding. Figure 1 The overall pipeline is demonstrated, emphasizing the core role of depth guidance in object-level reconstruction and controllability.

[0029] (1) Depth estimation and pose estimation

[0030] The depth estimation module extracts multi-scale features from the input monocular image to generate an initial depth map of the entire scene. The pose estimation module uses two consecutive frames as input to estimate the relative camera pose between adjacent frames.

[0031] The depth estimation module includes an encoder and a decoder. The encoder first downsamples the input image using convolutional layers and maximum pooling layers to obtain a feature map. Multiple residual blocks and multi-query attention modules are then stacked to extract deeper features. Each residual block includes two convolutional layers and a skip connection to enhance the network's nonlinear expression capabilities and gradient propagation efficiency. Each multi-query attention module includes a multi-query attention layer and a feedforward network to capture long-range dependencies in the feature map. The decoder stacks multiple residual blocks and upsampling modules to gradually recover the details of the depth map. The PixelShuffle operation is used for upsampling to ultimately obtain the initial depth map.

[0032] The relative pose estimation module includes an encoder and a decoder. The encoder part of the pose estimation module: the initial depth map is subjected to two average pooling operations to obtain two downsampled versions of the image; the two downsampled versions of the image and the original image are used as the input images of the pose estimation module encoder; the pose estimation module encoder extracts multi-scale feature maps through stacked residual blocks and multi-query attention modules. The feature maps capture local details and global context information in the input image. The decoder part of the pose estimation module: includes a global average pooling layer and a pose output module. The global average pooling layer averages the feature maps output by the pose estimation module encoder in the spatial dimension to obtain a fixed-length feature vector; the pose output module maps the feature vector to the final relative camera pose.

[0033] (2) Training process

[0034] To address the problem in object-level NeRF reconstruction of monocular videos, where the sequential dependency between camera pose estimation and NeRF model optimization makes it difficult to directly backpropagate NeRF loss gradients to the upstream depth estimation network, thereby limiting the effective utilization of depth information and the potential for system joint optimization, this paper proposes an iterative optimization training strategy consisting of four core stages.

[0035] This strategy aims to streamline the gradient flow from the NeRF model loss to the depth estimation network through a specific mechanism, enabling iterative refinement of depth information and joint optimization of the NeRF model. This strategy effectively combines the geometric priors provided by monocular self-supervised depth estimation with NeRF's object-level scene representation capabilities. The following is a detailed description of the four phases:

[0036] Phase 1: Monocular Self-Supervised Depth and Pose Pre-estimation. This phase utilizes a monocular self-supervised depth estimation framework to perform preliminary training on the input monocular video sequence, generating a depth map for each frame and the relative camera pose between adjacent frames. The training objective is to obtain an initial geometric prior to provide depth guidance for subsequent NeRF object-level reconstruction. The depth map is optimized using self-supervised signals to ensure consistency with the video temporal information, while the pose estimation network lays the foundation for camera motion modeling. The results of this phase directly impact the accuracy of subsequent object boundaries.

[0037] Phase 2: Initial NeRF training based on depth guidance. After obtaining the initial depth map and relative pose, this information is used to train a preliminary object-level NeRF model. First, the absolute camera pose is calculated based on the relative pose and camera intrinsic parameters as the input parameters of NeRF. Then, the original RGB image is used as the supervision signal to optimize the color prediction of NeRF through the photometric loss (Equation (6)). At the same time, the depth contrast loss (Equation (7)) is introduced to constrain NeRF to learn object geometry with the depth map of phase 1 as supervision. In addition, the density loss (Equation (9)) is used to accelerate training and reduce background interference. Through depth guidance, NeRF focuses on the target object, preliminarily reconstructs its 3D shape and appearance, and provides a basic representation for controllable editing.

[0038] Phase 3: Depth and pose optimization based on NeRF feedback. The NeRF model trained in Phase 2 is used to further refine the depth map and camera pose to improve geometric consistency. Specifically, the depth map of each frame is rendered by NeRF and fused with the depth estimation results of Phase 1 to generate more accurate depth information. At the same time, the absolute camera pose is recalculated through triangulation based on the rendered depth and camera intrinsic parameters. In this phase, the depth estimation network and pose estimation network are optimized through NeRF feedback, allowing them to capture finer object details. The optimized depth and pose serve as stronger geometric constraints and feed back to subsequent NeRF training.

[0039] Phase 4: NeRF object-level optimization based on refined depth and pose. After obtaining a more accurate depth map and camera pose, the NeRF model is retrained to further improve the quality of object-level reconstruction. Similar to Phase 2, the photometric loss, depth contrast loss, and density loss are jointly optimized. The difference is that the depth information in this phase is more reliable and can more effectively constrain the object geometry, ensuring boundary clarity and shape accuracy. Supported by multi-resolution hash position encoding and spherical harmonic direction encoding, the NeRF model not only reconstructs a high-quality 3D representation of the object, but also has controllability, allowing users to edit the pose or perspective by adjusting the encoding parameters.

[0040] In summary, the four-stage iterative optimization strategy described in the present invention establishes and strengthens an effective gradient path between object reconstruction and deep network optimization, so that monocular depth guidance plays a core role in the entire training process. Among them, stages one and three continuously optimize the quality of geometric constraints through depth estimation and NeRF feedback, while stages two and four use this high-quality depth information to guide NeRF to perform accurate object-level modeling. Through iterative improvement and mutual promotion, this strategy not only effectively overcomes the inherent difficulties of depth ambiguity and posture drift in monocular input, but also successfully generates object models with fine geometric details and accurate appearance, and provides a solid foundation and strong support for downstream controllable editing tasks such as object posture adjustment and new perspective synthesis.

[0041] (3) Efficient geometric enhancement of neural radiation field

[0042] 1) Multi-resolution hash position encoding

[0043] The neural radiation field is passed through a multi-layer perceptron network F Θ :(x,d)→(c,σ) maps the input spatial position x=(x,y,z) and viewing direction d=(θ,φ) to the corresponding color c=(r,g,b) and volume density σ. In order to enable MLP to better capture high-frequency information, the original NeRF uses positional encoding γ(·) to map the input coordinates to a higher-dimensional space:

[0044] γ(p)=(sin(2 0 πp), cos(2 0 πp),...,sin(2 L-1 πp), cos(2 L-1 πp)) (1)

[0045] Where p represents the position or direction of the input, and L is a hyperparameter that controls the number of frequencies. For spatial position x, L = 10; for viewing direction d, L = 4.

[0046] While this positional encoding method effectively improves NeRF's ability to represent high-frequency details, it also has some drawbacks. First, a fixed set of frequencies may not adapt to the complexity of different scenes. Due to its relatively simple structure, learning 3D scenes is extremely slow, leading to slow convergence in many cases. Second, the high-dimensional positional encoding vector significantly increases the computational burden of the MLP, reducing training and rendering efficiency.

[0047] To solve these problems, a lot of work has been done around position encoding to make many improvements, the most notable of which is Instant-NGP. Instant-NGP proposes an efficient multi-resolution hash encoding method as an improvement to the original position encoding. The core idea of ​​Instant-NGP is to divide the scene space into multiple grids of different resolutions and store a learnable feature vector at the vertex of each grid. For any spatial point, Instant-NGP first determines the grid it is in, and then maps the index of the grid vertex to a feature table through a hash function to obtain the corresponding feature vector. Finally, these feature vectors are interpolated to obtain the final feature representation of the point.

[0048] Specifically, Instant-NGP divides the space into L grids of different resolutions, and each resolution level l corresponds to a grid size N l For each resolution level, Instant-NGP maintains a hash table of size T, where each entry stores an F-dimensional feature vector. For a given spatial point x, Instant-NGP first calculates its grid coordinates at each resolution level l:

[0049]

[0050] Then, for each mesh vertex x l,u , Instant-NGP uses a spatial hash function to map its index into a hash table:

[0051]

[0052] Among them, π v are predefined coprime integers, Represents a bitwise exclusive OR operation.

[0053] Through the hash function, the feature vector f corresponding to each vertex can be obtained l,u Finally, Instant-NGP performs trilinear interpolation on these feature vectors to obtain the final feature representation f of the point l .

[0054] For the parameter setting of multi-resolution hash coding, the specific parameters used in this embodiment are as follows: L = 16 resolution levels, grid size N l From N min =16 exponential growth to N max =512, hash table size T=2 14 , the feature vector dimension F = 2. The selection is based on the following:

[0055] a) Resolution level L: L = 16 provides enough levels to represent geometric details without excessive computational overhead or loss of details.

[0056] b) Grid size N l :N min =16 and N max =512 ensures coverage from overall shape to fine texture, which is suitable for object reconstruction in limited space.

[0057] c) Hash table size T: T = 2 14 It maintains a low conflict probability while controlling the number of parameters, which is suitable for object-level reconstruction.

[0058] d) Feature vector dimension F: F = 2 can reduce the computational complexity while still providing good reconstruction results.

[0059] 2) Spherical Harmonic Direction Coding

[0060] In the neural radiation field, in addition to encoding the spatial position, encoding the viewing direction is also crucial because it directly affects the perspective consistency and detail performance of the rendering results. Therefore, in order to achieve controllable 3D generation, it is necessary to precisely control the ray sampling direction of NeRF. The original NeRF uses sine and cosine functions similar to position encoding to encode the direction. Although this method is simple and effective, it has certain limitations in capturing high-frequency information and representing complex lighting effects. In order to better represent the viewing direction and enhance NeRF's ability to model complex lighting and material properties, this method uses spherical harmonics (SH) as the basis for direction encoding. Spherical harmonic direction encoding provides an efficient and accurate way to represent and control the direction of rays. By changing the direction vector d input to the spherical harmonic direction encoding module, rays in any direction can be generated, thereby achieving free control of the rendering perspective.

[0061] Spherical harmonics are a set of orthogonal basis functions defined on a sphere, often used to represent signals on a sphere. They can be seen as a generalization of the Fourier series on a three-dimensional sphere. Spherical harmonics are usually expressed as , where n is the degree, m is the order, and θ and φ are the polar angle and azimuth angle in the spherical coordinate system, respectively. The higher the order of the spherical harmonic function, the stronger its ability to represent high-frequency information.

[0062] In the present invention, the first K-order spherical harmonics are used to encode the viewing direction d. Specifically, for a given viewing direction d = (θ, φ), the corresponding spherical harmonic value is first calculated:

[0063]

[0064] in, Represents the spherical harmonic function of the nth order and mth degree. These spherical harmonic function values ​​constitute a (K) 2 dimensional vector, encoding the viewing direction.

[0065] In this method, K=4 is selected, that is, the first four spherical harmonic functions are used to encode the viewing direction. This produces a 16-dimensional direction encoding vector. This spherical harmonic direction encoding vector is concatenated with the feature vector obtained by multi-resolution hash position encoding and used as the input of the MLP. By introducing spherical harmonic direction encoding and modifying the network structure, the representation ability and controllability of the NeRF model are further improved, making it more suitable for the task of 3D controllable generation. This improved NeRF model can better capture the lighting changes and material properties in the scene and achieve precise control of the rendering perspective, laying the foundation for achieving high-quality object-level 3D controllable generation.

[0066] 3) Neural radiation field structure and reasoning process

[0067] In the present invention, the MLP network structure for predicting color c and volume density σ is relatively simple. Since multi-resolution hash coding already provides rich and discriminative feature information, MLP only needs to act as a decoder to map these features to the final output. Therefore, a fully connected network with two hidden layers is adopted, each hidden layer contains 64 neurons, and uses the ReLU activation function. The output layer of the network is a linear layer, which is used to generate volume density σ and perspective-dependent color c. Specifically, the volume density σ is represented by a single scalar value, and the color c is a three-dimensional vector corresponding to the three RGB channels. This lightweight MLP structure helps to further improve the computational efficiency of the neural radiation field, making it more suitable for real-time application scenarios. As Figure 2 shown.

[0068] In the inference phase, given a spatial point x and an observation direction d, the multi-resolution hash code is first calculated, followed by the spherical harmonic direction code. Finally, these two codes are combined and input into the MLP to obtain the final color and volume density. The text inference process includes the following six steps:

[0069] a) Determine the input: given a spatial point x and an observation direction d.

[0070] b) Calculate the multi-resolution hash code: For each resolution level l: calculate the coordinate x in the grid of that level l For x l Calculate the hash index h(x l ,u). Find the corresponding feature vector f from the hash table according to the hash index l,u Perform trilinear interpolation on the 8 eigenvectors to obtain the final feature representation f at this level l .

[0071] c) Calculate the spherical harmonic direction code: Calculate the spherical harmonic function value SH(d) according to the observation direction d.

[0072] d) Splicing features: Concatenate the feature vectors f at all levels l They are concatenated and combined with the spherical harmonic direction code SH(d) to obtain the final input feature f.

[0073] e) MLP prediction: The feature vector f is input into the MLP. The first MLP outputs the intermediate feature vector f based on the position encoding feature mid and volume density σ. The second MLP combines the intermediate feature vector f mid And the spherical harmonic function value SH(d), output color c.

[0074] f) Volume rendering: Using volume rendering technology, the color and density are accumulated along the light to obtain the final pixel color.

[0075] By using multi-resolution hash coding and spherical harmonic directional encoding, Neural Radiance Fields can more efficiently represent the geometry and appearance of a scene. Compared to the original NeRF, this method significantly reduces training time and memory consumption while maintaining comparable or even better rendering quality. This efficient representation opens up greater possibilities for object-level reconstruction and controllable editing.

[0076] (3) Camera pose refinement based on cumulative error suppression

[0077] In order to accurately estimate the absolute pose of the camera in the world coordinate system (camera-to-world, Tcw) from a monocular video sequence, this paper proposes a camera pose refinement method based on cumulative error suppression. This method uses the relative camera pose as the initial estimate and effectively suppresses the error caused by the recursive accumulation of the relative pose through a series of geometric constraints and global optimization steps, providing high-precision camera pose input for subsequent object-level 3D reconstruction based on neural radiation fields. Directly using the relative pose to recursively calculate the absolute pose will inevitably introduce cumulative errors, causing the camera trajectory to drift and deviate from the true motion trajectory. To address this problem, the core idea of ​​this method is to introduce multi-step geometric constraints and global optimization strategies to disperse the error throughout the video sequence, thereby obtaining a more accurate and robust camera pose estimation result.

[0078] First, calculate the initial pose. Initialize the absolute pose of the camera in the first frame to a unit matrix Then, the relative pose T between adjacent frames output by the relative pose estimation network is i+1,i , calculate the initial absolute posture of each frame by recursion. The recursion formula is as follows:

[0079]

[0080] in, Represents the pose of the camera in the world coordinate system of the i-th frame. Through this formula, a preliminary, unoptimized absolute camera pose sequence can be obtained.

[0081] At the same time, to eliminate the impact of differences in coordinate system definitions, the initial absolute camera pose needs to be normalized. Furthermore, to address the cumulative rotation error, translation error, and scale drift, this method proposes a joint optimization strategy. This strategy gradually refines the camera pose through three steps: global rotation optimization, global translation optimization based on optical center line constraints, and adaptive scale normalization. The specific algorithm flow is as follows:

[0082] 1) Global rotation optimization:

[0083] a) Consensus Upward Direction Estimation: First, from each camera pose Extract the y-axis direction vector in the world coordinate system from the rotation component of . Accumulate all these vectors to get an aggregated up vector.

[0084] b) Normalization and Alignment Matrix Calculation: The aggregated up direction vectors are normalized by their L2 norm to obtain a unit vector representing the average up direction of the scene. A rotation matrix R is then calculated to rotate this average up direction to align with the canonical Z axis of the world coordinate system.

[0085] c) Global rotation transformation application: Expand the alignment rotation matrix R into a homogeneous transformation matrix And multiply it left to each camera pose in the sequence This operation rigidly rotates the entire scene and all camera poses so that the up direction of the entire scene is aligned with the Z axis of the world coordinate system.

[0086] 2) Global translation optimization:

[0087] a) Initialize the accumulator: Initialize a three-dimensional zero vector P for accumulating weighted spatial points, and a scalar W to zero for accumulating weights.

[0088] b) Weighted aggregation of pairwise camera consensus points:

[0089] Traverse all camera pose pairs (s, t) through double iteration; for each camera pose pair (s, t), extract the optical center of camera s and the direction of its main optical axis, and perform the same operation on camera t.

[0090] Call a function cp(o1,d1,o2,d2), which infers a three-dimensional space point p based on the optical center and main optical axis direction of the two cameras (s,t) and an associated weight w (s,t) This point can be considered as the estimated position of the scene feature point observed by the two cameras, and the weight may reflect the confidence of the estimate or the strength of the geometric constraint. The function cp(o1,d1,o2,d2) represents the calculation of the closest point and weight of the two rays (o1,d1) and (o2,d2).

[0091] The weighted point p (s,t) w (s,t) Accumulate to P and weight w (s,t)) Accumulate to scalar W.

[0092] c) Scene centroid calculation: Calculate the weighted average of all paired consensus points through P / W to obtain the geometric centroid of the scene

[0093] d) Global translation transformation application: from each camera pose in the sequence The calculated scene center of mass is subtracted from the translation component of

[0094] This operation translates the entire scene and all camera poses so that the center of mass of the scene is aligned with the origin of the world coordinate system. 3) Adaptive scale normalization:

[0095] a) Average scene radius calculation: Calculate the average Euclidean distance L from the optical center of all cameras to the (new) world coordinate system origin O This L OThe value represents the characteristic scale of the current scene.

[0096] b) Scale normalization transformation application: each camera pose in the sequence The translation component of is multiplied by a scaling factor. This operation uniformly adjusts the overall scale of the scene so that the average distance from the camera to the center of the scene is normalized to a preset value (here 4), which helps eliminate scale ambiguity and ensures scale consistency in subsequent processing.

[0097] The above series of geometric constraints and global optimization steps effectively suppress the cumulative error and improve the accuracy and robustness of camera pose estimation. Compared with the traditional recursive-based camera pose calculation method, this method has significant advantages. First, through global rotation and translation optimization, the global consistency of the camera trajectory is guaranteed, and the accumulation of local errors is avoided. Second, the constraints and optimization based on multi-frame information make this method more robust to noise and outliers. The final optimized camera absolute pose is It will serve as the input of the NeRF model to accurately constrain the scene geometry and camera motion trajectory, thereby achieving high-quality object-level scene reconstruction.

[0098] (4) Loss function

[0099] In order to better handle the occlusion problem, the sampled rays are divided into three types according to their instance masks: rays R that hit the target object o , ray R pointing to the background b and the ray R that hits other occluding objects m .like Figure 3 shown.

[0100] 1) Photometric loss: Photometric loss is the most basic loss function in NeRF. It optimizes the NeRF model based on the pixel-level difference between the rendered image and the real image. For each sampled ray, its color value is first calculated according to the NeRF model, and then compared with the color value of the corresponding pixel in the real image. Only the ray R that hits the target object is compared. o , calculate its color loss:

[0101]

[0102] in, represents the light r predicted by the NeRF model i The color value, C(r i ) represents the color value of the corresponding pixel in the real image. By applying photometric loss only to the target object rays, the interference of occluded areas on the object appearance modeling is avoided.

[0103] 2) Depth contrast loss: In order to better utilize the temporal information in monocular videos and further constrain the geometry of NeRF, depth contrast loss is introduced and relative depth relationship is used as the supervisory signal.

[0104] Similarly, only the ray R that hits the target object o Apply this loss:

[0105]

[0106] Among them, P represents the set of all pixel pairs in the target object area, L p (p 0 ,p 1 ) is defined as follows:

[0107]

[0108] Among them, D nerf (p) represents the depth value of the spatial point corresponding to pixel p predicted by the NeRF model in the camera coordinate system, and contrast is the depth contrast value. When contrast ≠ 0, the cross-entropy loss is used to ensure that the depth difference predicted by NeRF is consistent with the depth contrast value; when contrast = 0, the mean squared error loss is used to minimize the depth difference predicted by NeRF.

[0109] 3) Density loss: Density loss directly constrains the volume density of the sampling point and is used to process different types of light. Different density targets are defined based on the light type:

[0110]

[0111] in, Represents ray r i The volume density of the jth sampling point on σ, N is the number of sampling points. target (M o ) represents the object instance M o The relevant target density value (usually set to 1). For the ray R pointing to the background b , encouraging its body density to approach zero; for rays R that hit other blocking objects m , encouraging its volume density to be close to that of the object instance M o Related target values, which helps to correctly handle occlusion relationships in complex scenes.

[0112] 4) Total loss function: Finally, the above three loss functions are combined to obtain the total loss function:

[0113] L Z =L rgb +λ c Lc +λ density L density (10)

[0114] Among them, λ c and λ density is the weight coefficient used to balance the importance of different loss functions. In the experiment, we set λ c = 0.1 and λ density =0.01.

[0115] Through this ray classification-based, multi-loss function joint training strategy, the method of the present invention fully utilizes the information in monocular video and effectively handles occlusions. The photometric loss ensures accurate reconstruction of the target object's appearance, the depth contrast loss provides geometric shape constraints, and the density loss helps distinguish between the target object, occluded objects, and background areas. This comprehensive optimization method reconstructs 3D models of objects with fine geometric detail and accurate appearance from monocular video input, while correctly handling occlusion relationships in complex scenes.

[0116] 2. Dataset and Evaluation Metrics

[0117] (1) Cube-Diorama dataset

[0118] The Nerf-Cube-Diorama dataset, created by Jad Abou-Chakra and Niko Sünder-hauf of the Robotics Center at Queensland University of Technology, Australia, and Feras Dayoub of the School of Computer Science at the University of Adelaide, is a dataset used to evaluate the performance of neural radiation field-related algorithms. The dataset contains synthetic rendered images of four objects: a bellflower, a book, a cup, and a laptop. Figure 4 shown.

[0119] (2) Replica Dataset

[0120] The Replica dataset was created by Facebook AI Research (FAIR) to promote research in the field of indoor 3D scene understanding, such as 3D reconstruction, semantic segmentation, and object recognition. The dataset contains high-quality, realistically rendered images and provides rich information such as corresponding depth maps, normal maps, semantic segmentation labels, and camera poses. The dataset covers a variety of indoor scenes, including offices, apartments, and hotel rooms, showing different layouts, furniture, and decoration styles, such as Figure 4-5As shown in the figure, Replica uses a professional rendering engine to simulate various lighting conditions and material properties, providing complete and accurate scene information. Compared to real-world data, the Replica dataset effectively avoids problems such as noise, occlusion, and annotation errors, providing researchers with a cleaner and more reliable data foundation.

[0121] (3) Evaluation indicators

[0122] To quantitatively evaluate the performance of the proposed object-level 3D reconstruction method, three widely used metrics can be used: reconstruction accuracy, completeness, and completeness at a given distance. These metrics measure the difference between the reconstructed model and the true model from different perspectives and can comprehensively reflect the reconstruction quality.

[0123] 1) Reconstruction Accuracy: Reconstruction accuracy measures the average distance between points on the reconstructed model surface and the true model surface. Specifically, for each point on the reconstructed model, the distance to the nearest point on the true model surface is calculated. The distances across all points are then averaged to obtain the reconstruction accuracy. Reconstruction accuracy reflects the geometric accuracy of the reconstructed model; smaller values ​​indicate higher reconstruction accuracy.

[0124] Let S rec To reconstruct the model surface, S gt is the real model surface, N rec is the number of points on the reconstructed model surface. For each point p on the reconstructed model i ∈S rec , calculate its closest distance to the true model surface:

[0125]

[0126] Among them, q represents a point on the surface of the true model. Then, the distance of all points is averaged to obtain the reconstruction accuracy:

[0127]

[0128] 2) Completeness: Completeness measures the average distance between points on the ground truth model and the reconstructed model. Specifically, for each point on the ground truth model, the distance to the nearest point on the reconstructed model is calculated. The distances are then averaged across all points to determine the degree of completeness. The degree of completeness reflects the extent to which the reconstructed model covers the ground truth model's geometry, with lower values ​​indicating higher completeness.

[0129] Let N gt is the number of points on the surface of the real model. For each point q on the real model j ∈Sgt , calculate its closest distance to the reconstructed model surface:

[0130]

[0131] Where p represents a point on the reconstructed model surface. Then, the distances of all points are averaged to get the degree of completion:

[0132]

[0133] 3) Completion Rate: The completion rate measures the proportion of the ground truth model surface covered by the reconstructed model within a given distance threshold. Specifically, for each point on the ground truth model, the closest distance from the reconstructed model surface is determined to be less than the given distance threshold. If so, the point is considered covered. The completion rate is then calculated by counting the proportion of covered points to the total number of points. The completion rate reflects the completeness of the reconstructed model under different accuracy requirements, with higher values ​​indicating higher completion rates.

[0134] Let τ be a given distance threshold. For each point q on the true model j ∈S gt , to determine whether it is covered:

[0135]

[0136] Then, calculate the proportion of covered points to the total points to get the completion rate:

[0137]

[0138] In experiments on the Replica dataset, the completion rates were calculated for distance thresholds of 1 cm and 5 cm, while in experiments on the Cube-Diorama dataset, only the completion rates were calculated for distance thresholds of 0.4 cm and 1 cm. This design is designed to better adapt to the characteristics of different datasets and to comprehensively evaluate the performance of the reconstructed model under different accuracy requirements. These four metrics can be used to evaluate the quality of the reconstructed model from different perspectives.

[0139] 3. Analysis of experimental results

[0140] (1) Experimental results of the Cube-Diorama dataset

[0141] First, experiments were conducted on the Cube-Diorama dataset to evaluate the performance of the proposed method on object-level 3D reconstruction. Table 1 presents the quantitative evaluation results. As can be seen, compared with So-slam, which uses real depth information, the proposed method is slightly inferior in terms of reconstruction accuracy. This is mainly because So-slam uses real depth data to more effectively capture object geometric features. However, in other key metrics such as completion and completion rate, the proposed method achieved the best results, significantly outperforming other existing technologies. Figure 6 The results of object reconstruction using the Cube-Diorama dataset are presented for each method. It can be observed that the proposed method is able to generate 3D models with fine geometric details and accurate appearance representation, which are highly consistent with the ground-truth reference model. However, due to the inherent limitations of monocular vision methods, some artifacts still exist in the reconstruction results.

[0142] Table 1 Quantitative evaluation results of object reconstruction on the Cube-Diorama dataset

[0143]

[0144] (2) Experimental results of the Replica dataset

[0145] In the experiments on the Replica dataset, in order to facilitate comparison with existing methods, evaluation was performed on two sequences, room0 and office1, which are mainly composed of smaller objects. The quantitative evaluation results are shown in Table 2. As can be seen from the table, the method of the present invention has basically achieved the best results in all indicators. Specifically, compared with the optimal method, the average improvement in all indicators is 3.12%, and the improvement in accuracy indicators is the most significant, especially in the office-1 scene, where it has increased by 12.33%. The qualitative reconstruction results are shown in Table 2. Figure 7 As shown in the figure above, the reconstruction results of multiple objects and the estimated camera poses are shown. The comparison results with other methods are shown below. It can be seen that the method of the present invention can better reconstruct the geometric shape of objects, with clearer and more complete details.

[0146] Table 2 Quantitative evaluation results of object reconstruction on the Replica dataset

[0147]

[0148] (3) Ablation experiment

[0149] To further validate the effectiveness of our method, we conducted a column ablation experiment on the Replica dataset to analyze the impact of different components on reconstruction performance. Specifically, we removed the depth contrast loss, density loss, and four-stage training method, and then compared the reconstruction results of different components. The experimental results are shown in Table 3. As can be seen from the table, the depth contrast loss, density loss, and training method all significantly improved reconstruction performance, increasing reconstruction accuracy, completeness, and completion rate, respectively. This result further demonstrates the effectiveness of our method, and that the design of each component can effectively improve reconstruction performance.

[0150] Table 3 Ablation experiment results

[0151]

[0152] In addition, we further evaluated the depth estimation performance on the Replica dataset in the first and third phases to verify the potential contribution of the multi-stage training strategy to improving depth estimation accuracy. The quantitative analysis results of the experiment are summarized in Table 4, and the corresponding qualitative comparisons are shown in Figure 8 .

[0153] The data in the table shows that the depth estimation performance of the third stage is significantly better than that of the first stage; this improvement is even more intuitive when combined with the visualization results in the figure. Compared to the first stage, where depth estimation tends to over-focus on the boundary areas of the environment, the third stage model demonstrates a more comprehensive and profound understanding of the entire scene. This phenomenon demonstrates that multi-stage training not only optimizes the model's capture of details but also brings substantial improvements in global perception, effectively enhancing the overall accuracy of depth estimation.

[0154] Table 4 Experimental results of the first and third stage depth estimation on the Replica dataset (↓ indicates that the smaller the value, the better, ↑ indicates that the larger the value, the better)

[0155]

[0156] In summary, this paper proposes a monocular depth-guided object-level NeRF reconstruction method. By integrating monocular self-supervised depth estimation, object segmentation, and geometrically enhanced NeRF, and utilizing four-stage iterative training and depth contrast consistency regularization, this method reconstructs high-quality, controllable 3D object models from monocular videos. Experiments demonstrate that this method significantly improves reconstruction accuracy and completeness on the Cube-Diorama and Replica datasets. Furthermore, the encoding design supports controllable editing of pose and position, validating the effectiveness of depth guidance.

Claims

1. A monocular depth-guided object-level NeRF reconstruction method, characterized in that The following steps are involved: Step 1: Using a monocular video sequence as input, a self-supervised depth estimation module is used to generate a preliminary depth map for each frame, and a relative pose estimation module is used to estimate the relative camera pose between adjacent frames. Step 2: Extract the target object mask based on the object segmentation module, and calculate the absolute pose of the camera by combining the preliminary depth map and the relative camera pose; Step 3: Construct a geometrically enhanced NeRF model, taking the absolute pose as input and optimizing the scene representation through multi-resolution hash position encoding and spherical harmonic direction encoding; Step 4: Introduce photometric loss, depth contrast loss, and density loss to jointly optimize the parameters of the NeRF model, and use depth information to constrain object geometry reconstruction; Step 5: Through a four-stage iterative training strategy, alternately optimize the depth estimation, camera pose and NeRF model to generate an object-level controllable 3D model.

2. The monocular depth-guided object-level NeRF reconstruction method according to claim 1, characterized in that In step 1, The depth estimation module includes an encoder and a decoder; The encoder part of the depth estimation module first uses convolutional layers and maximum pooling layers to downsample the input image to obtain a feature map. Then, multiple residual blocks and multi-query attention modules are stacked to extract deeper features. Each residual block includes two convolutional layers and a skip connection to enhance the nonlinear expression ability of the network and the efficiency of gradient propagation. Each multi-query attention module includes a multi-query attention layer and a feedforward network to capture long-range dependencies in the feature map. Depth estimation module decoder: stacks multiple residual blocks and upsampling modules to gradually recover the details of the depth map, and uses the PixelShuffle operation to upsample to obtain the depth map; The relative posture estimation module includes an encoder and a decoder; Relative pose estimation module encoder part: the initial depth map is subjected to two average pooling operations to obtain two downsampled versions of the image; The two downsampled versions of the image are used together with the original image as the input image of the pose estimation module encoder; The relative pose estimation module encoder extracts multi-scale feature maps through stacked residual blocks and multi-query attention modules, which capture local details and global context information in the input image; The decoder part of the pose estimation module includes a global average pooling layer and a pose output module. The global average pooling layer averages the feature map output by the relative pose estimation module encoder in the spatial dimension to obtain a fixed-length feature vector; the pose output module maps the feature vector to the final relative camera pose.

3. The monocular depth-guided object-level NeRF reconstruction method according to claim 2, characterized in that Step 2 specifically includes: Step 2.1: Calculate the initial pose: Initialize the absolute pose of the first frame camera to a unit matrix Step 2.2: Based on the relative camera pose T between adjacent frames output by the relative pose estimation module i+1,i , the initial absolute pose of each frame is calculated recursively; thus a preliminary, unoptimized absolute camera pose sequence is obtained; the recursive formula is as follows: in, Represents the posture of the camera in the world coordinate system of the i-th frame; Step 2.3: The camera pose refinement algorithm based on cumulative error suppression gradually refines the camera pose through three steps: global rotation optimization, global translation optimization based on optical center connection constraint, and adaptive scale normalization.

4. The monocular depth-guided object-level NeRF reconstruction method according to claim 3, characterized in that Step 2.3 is as follows: Step 2.3.1: Global rotation optimization a) Consensus upper direction estimation: First, from each camera pose Extract the y-axis direction vector in the world coordinate system from the rotation component of Accumulate all y-axis direction vectors to obtain an aggregated upward direction vector; b) Normalization and alignment matrix calculation: The aggregated up direction vectors are normalized by the L2 norm to obtain a unit vector representing the average up direction of the scene; then, an alignment rotation matrix R is calculated; c) Global rotation transformation application: Expand the alignment rotation matrix R into a homogeneous transformation matrix And multiply it left to each camera pose in the sequence Align the upward direction of the entire scene with the Z axis of the world coordinate system; Step 2.3.2: Global translation optimization a) Initialize the accumulator: Initialize a three-dimensional zero vector P for accumulating weighted spatial points, and a scalar W to zero for accumulating weights; b) Weighted aggregation of paired camera consensus points: traverse all camera pose pairs (s, t) through double iteration; for each camera pose pair (s, t), extract the optical center and principal optical axis direction of camera s, and perform the same operation for camera t; Call the function cp(o1, d1, o2, d2) to infer a three-dimensional space point p based on the optical center and main optical axis direction of the two cameras (s,t) and an associated weight w (s,t) ; The weighted point p (s,t) w (s,t) Accumulate to the three-dimensional zero vector P, and weight w (s,t)) Accumulate to scalar W; the function cp(o1, d1, o2, d2) represents the calculation of the closest point and weight of two rays (o1, d1) and (o2, d2); c) Scene centroid calculation: Calculate the weighted average of all paired consensus points through P / W to obtain the geometric centroid of the scene d) Global translation transformation application: from each camera pose in the sequence The calculated scene center of mass is subtracted from the translation component of Align the center of mass of the scene with the origin of the world coordinate system; Step 2.3.3: Adaptive scale normalization a) Average scene radius calculation: Calculate the average Euclidean distance L from the optical center of all cameras to the origin of the world coordinate system O ; b) Scale normalization transformation application: each camera pose in the sequence The translation component of is multiplied by a scaling factor so that the average distance from the camera to the center of the scene is normalized to a preset value to eliminate scale ambiguity and ensure scale consistency in subsequent processing.

5. The monocular depth-guided object-level NeRF reconstruction method according to claim 1, characterized in that In step 3, the geometrically enhanced NeRF model is passed through a multi-layer perceptron network F Θ :(x,d)→(c,σ) maps the input spatial point x and viewing direction d to the corresponding color c and volume density σ; specifically: Step 3.1: Multi-resolution hash position encoding Divide the space into L grids of different resolutions, each resolution level l corresponds to a grid size N l ;For each resolution level, maintain a hash table of size T, where each entry stores an F-dimensional feature vector; For a given spatial point x, first calculate its grid coordinates at each resolution level l: For x l Each mesh vertex around x l,u , use a spatial hash function to hash its index h(x l,u ) is mapped into the hash table to obtain the feature vector f corresponding to each mesh vertex l,u ; Then perform trilinear interpolation on each eigenvector to obtain the position encoding feature representation f of each grid vertex l ; The spatial hash function is expressed as: Among them, π v are predefined coprime integers, Represents a bitwise exclusive OR operation, x l,u,v is the vth coordinate component of the uth vertex at resolution level l; v is the coordinate dimension, which is x, y, and z respectively; mod is the modulo operation; Step 3.2: Spherical Harmonic Direction Encoding The first K-order spherical harmonics are used to encode the viewing direction d. For a given viewing direction d = (θ, φ), the corresponding spherical harmonic value SH(d) is calculated: Among them, Y n m (θ, φ) represents the spherical harmonic function of the nth order and mth degree; θ and φ are the polar angle and azimuth angle in the spherical coordinate system respectively; Step 3.3: The spherical harmonic function value SH(d) obtained by spherical harmonic direction encoding is combined with the position encoding feature representation f obtained by multi-resolution hash position encoding l Spliced ​​together as input to the multi-layer perceptron; It includes two multi-layer perceptron networks; the first multi-layer perceptron network represents f according to the position encoding feature l Output intermediate feature vector f mid and volume density σ; the second multi-layer perceptron network combines the intermediate feature vector f mid And the spherical harmonic function value SH(d), the output color c; the volume density σ is represented by a single scalar value, and the color c is a three-dimensional vector corresponding to the three RGB channels; using volume rendering technology, the color and density are accumulated along the light to obtain the final predicted color value.

6. The monocular depth-guided object-level NeRF reconstruction method according to claim 1, characterized in that In step 4, 1) Calculate the luminosity loss: For each sampled ray, its color value is first calculated according to the NeRF model, and then compared with the color value of the corresponding pixel in the real image. o , calculate its luminosity loss L rgb : in, represents the light r predicted by the NeRF model i The color value, C(r i ) represents the color value of the corresponding pixel in the real image; 2) Calculate the depth contrast loss: Only the ray R that hits the target object o Apply depth contrast loss L c : Among them, P represents the set of all pixel pairs in the target object area; pixel pair p 0 and p 1 The loss function L p (p 0 ,p 1 ) is defined as follows: Among them, D nerf (p 0 ) and D nerf (p 1 ) represents the pixel pair p predicted by the NeRF model 0 and p 1 The depth value of the corresponding spatial point in the camera coordinate system, contrast is the depth contrast value; 3) Calculate density loss: The volume density of the sampling point is directly constrained by density loss to handle different types of light; different density targets are defined according to the light type: in, Represents ray r i The volume density of the jth sampling point, N is the number of sampling points; σ target (M o ) represents the object instance M o The relevant target density value; R b is the ray pointing to the background; R m For rays that hit other blocking objects; 4) Calculate the total loss function: Combining the above three loss functions, we get the total loss function: L Z =L rgb +λ c L c +λ density L density (9) Among them, λ c and λ density is the corresponding weight coefficient.

7. The monocular depth-guided object-level NeRF reconstruction method according to claim 1, characterized in that In step 5, The four-stage iterative training strategy is as follows: Stage 1: Monocular self-supervised depth and pose estimation; This stage uses a monocular self-supervised depth estimation framework to perform preliminary training on the input monocular video sequence to generate the initial depth map of each frame and the relative camera pose between adjacent frames; The training goal is to obtain the initial geometric prior to provide deep guidance for the subsequent object-level reconstruction of the NeRF model; Phase 2: Initial NeRF model training based on depth guidance; After obtaining the initial depth map and relative camera pose, a preliminary object-level NeRF model is trained. First, the absolute camera pose is calculated based on the relative camera pose and camera intrinsic parameters, which is used as the input parameter of the NeRF model. Then, the original RGB image is used as the supervision signal, and the color prediction of the NeRF model is optimized through the photometric loss. At the same time, the depth contrast loss is introduced to constrain the NeRF model to learn object geometry using the depth map of stage 1 as supervision. In addition, the density loss is used to accelerate training and reduce background interference. Phase 3: Depth and pose optimization based on NeRF model feedback; The NeRF model trained in stage 2 is used to further refine the depth map and camera pose to improve geometric consistency. Specifically, the NeRF model is used to render the depth map of each frame and fuse it with the depth estimation result of the first stage to generate more accurate depth information. At the same time, the rendered depth and the camera intrinsic parameters are combined to recalculate the absolute camera pose through triangulation. In this phase, the depth estimation module and relative pose estimation module are optimized through feedback from the NeRF model, enabling them to capture finer object details. The optimized depth map and camera pose are used as geometric constraints to feed back into subsequent NeRF model training. Stage 4: NeRF object-level optimization based on refined depth and pose; After obtaining more accurate depth maps and camera poses, the NeRF model is trained again to further improve the object-level reconstruction quality; the model is jointly optimized using photometric loss, depth contrast loss, and density loss; and with the support of multi-resolution hash position encoding and spherical harmonic direction encoding, the final NeRF model reconstructs the three-dimensional representation of the object.

Citation Information

Cited By

  • High-fidelity three-dimensional reconstruction method fusing attitude prior and geometric constraint

    CN120931839A

  • A high-fidelity three-dimensional reconstruction method fusing pose prior and geometric constraint

    CN120931839B

  • Three-dimensional scene reconstruction method and system based on monocular depth estimation

    CN121482285A

  • Bridge crack three-dimensional reconstruction method and system based on depth feature fusion and storage medium

    CN122336154A