Real-time dynamic three-dimensional reconstruction method and system based on fusion of patch matching and lightweight monocular depth estimation
By combining the PatchMatch algorithm and lightweight monocular depth estimation network, the problems of large computing overhead and long inference time in the prior art are solved, real-time three-dimensional reconstruction and high-quality rendering in dynamic scenarios are realized.
Patent Information
- Application Number
- CN202510150059.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing multi-view 3D reconstruction technology has a large calculation overhead and a long inference time in dynamic scenarios or real-time applications, making it difficult to take into account both rendering quality and real-time.
The PatchMatch algorithm is used to combine it with a lightweight monocular depth estimation network to replace the traditional iterative multi-view stereo matching process, and real-time dynamic three-dimensional reconstruction is achieved through an end-to-end differentiable rendering framework.
It significantly improves the reconstruction speed and rendering quality of 3D Gaussian point cloud under sparse viewing conditions, achieves the purpose of real-time three-dimensional reconstruction, and improves the geometric reconstruction accuracy of scenes or characters.
Smart Images

Figure CN120070758A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and three-dimensional reconstruction, and specifically discloses a real-time dynamic three-dimensional reconstruction method and system based on the fusion of patch matching and lightweight monocular depth estimation. Background Art
[0002] Multi-view three-dimensional reconstruction has wide application value in the fields of computer vision, virtual reality, and film production. In recent years, Neural Radiance Field (NeRF) and its extended methods have made remarkable progress in novel view synthesis. However, NeRF usually requires a large amount of computation and optimization of sampling points in three-dimensional space, making it difficult to balance rendering quality and real-time performance. For this reason, explicit representation methods such as point cloud rendering have been favored again, and methods represented by 3DGS have significantly improved the rendering efficiency while maintaining visual quality. By representing the scene as learnable three-dimensional Gaussian distribution particles, 3DGS demonstrates excellent performance and scalability in real-time rendering and three-dimensional reconstruction tasks in complex scenes.
[0003] In the field of novel view synthesis, the GPS-Gaussian method proposes the idea of "direct regression" for multi-view input, no longer relying on continuous fine-tuning of a single scene or person, but directly outputting 3DGS point clouds frame by frame by learning generalizable human body (or other targets) priors on large-scale multi-view data, thereby achieving fast three-dimensional reconstruction. This method uses multi-view depth estimation to "lift" pixel-level Gaussian parameters from a two-dimensional plane to three-dimensional space. However, the multi-view stereo matching (MVS) used in GPS-Gaussian often adopts an iterative depth estimation process, and for targets with complex self-occlusions such as the human body, it is necessary to establish a relatively large-scale feature or cost volume for repeated optimization. This approach will face excessive computational overhead and long inference time in some dynamic scenarios or real-time applications.
[0004] To reduce the computational cost while maintaining high depth estimation accuracy, a feasible improvement idea is to adopt a strategy of combining the PatchMatch algorithm with a lightweight monocular depth estimation network to replace the original iterative MVS process. PatchMatch is a commonly used algorithm for fast matching based on local image patches (patches). Through mechanisms such as neighborhood propagation and random search, PatchMatch can search for similar local patches between multi-view images in an approximate manner and quickly converge to reasonable matching results. Compared with stereo matching methods based on global cost volumes, PatchMatch is more efficient in most texture-rich regions.
[0005] However, PatchMatch may exhibit unstable matching in regions with missing textures or a large number of repetitive textures, and the depth range or absolute scale information cannot be fully constrained solely by local matching. On the other hand, monocular depth estimation technology has developed rapidly in recent years. Through supervised or self-supervised learning with a large amount of data, monocular models can learn a large amount of prior knowledge and output reasonable depth distributions in the vast majority of scenarios. However, due to the lack of geometric constraints between multiple perspectives, monocular depth estimation usually suffers from the "scale ambiguity" problem (i.e., it cannot guarantee that the predicted depth is strictly consistent with the true physical scale). Summary of the Invention
[0006] Object of the Invention: The technical problem to be solved by the present invention is to provide a real-time dynamic three-dimensional reconstruction method and system based on the fusion of PatchMatch and lightweight monocular depth estimation, which combines pixel-level Gaussian parameter regression of source images and an end-to-end differentiable rendering framework, significantly improving the reconstruction speed and rendering quality of 3DGS under sparse view conditions and achieving the purpose of real-time reconstruction. The method includes the following steps:
[0007] Step 1: Prepare a monocular depth feature extraction model Initialize the convolutional neural network for PatchMatch and monocular depth feature fusion Prepare the encoder E required for the per-pixel Gaussian parameter prediction model depth And the decoder for predicting the rotation, scale, and opacity of Gaussian points respectively
[0008] Step 2: Input the training image pair (I 1 , I 2 ) and the corresponding pose matrices (T 1 , T 2 ), as well as the internal parameter matrix K of the camera; where I 1 , I 2 represent the first and second input training images respectively, and T 1 , T 2 represent the pose matrices corresponding to the first training image I 1 and the second training image I 2 respectively;
[0009] Step 3: Perform PatchMatch in the original pixel space of the image pair (I 1 , I 2 ) to obtain the sparse depth maps D 1 , D 2 of the matchable regions of I r1 , D r2, and the matching confidence S r1 , S r2 ;
[0010] Step 4: Input the image pairs (I 1 , I 2 ) into the monocular depth prediction model respectively, to obtain the monocular depth feature maps 1 of I and the monocular depth feature maps 2 of I where S is the dimension of the monocular depth feature map;
[0011] Step 5: Concatenate the sparse depth maps D r1 , D r2 , the matching confidence S r1 , S r2 , and the monocular depth feature maps into a multi-channel input to form a fused feature of (S + 2)×H×W dimension, and perform confidence-based feature fusion on the fused feature through a convolutional neural network to output the absolute depth maps D 1 , D 2 of the globally scale-aligned images I 1 , D 2 ; H and W are the height and width of the feature map respectively. The sparse depth maps D r1 , D r2 and the matching confidence S r1 , S r2 need to be bilinearly interpolated to the height and width of the monocular depth feature map for merging;
[0012] Step 6: Use the absolute depth maps D 1 , D 2 to lift the two-dimensional parameters to the three-dimensional space, construct a preliminary 3D Gaussian distribution, and obtain the per-pixel position matrix M p , color matrix M c , rotation matrix M r , scaling matrix M s , and opacity matrix M a through the encoder E depth and the decoder ; the color matrix of the Gaussian points is obtained from the training image pairs (I 1 , I 2 );
[0013] Step 7: Obtain the images C 1 , C 2 under the viewpoints with the camera pose matrices T 1 , C 2 respectively through volume rendering;
[0014] Step 8, repeat Steps 2 to 7, traverse image pairs at two or more moments, and update the parameters of the convolutional neural network and the Gaussian parameter prediction model until the optimization converges;
[0015] After the training is completed, an end-to-end real-time dynamic 3D reconstruction model is obtained, which has the generalization ability to perform Gaussian distribution regression on any new scene image pair.
[0016] In Step 4, the monocular depth feature map and are calculated by the monocular depth prediction model , and the formula is:
[0017]
[0018] In Step 6, the following formula is used to lift the two-dimensional parameters to the three-dimensional space:
[0019]
[0020] M c (i,j) = I(i,j)
[0021] where M p (i,j) represents the i-th row and j-th column of the position matrix M p ; M c (i,j) represents the i-th row and j-th column of the color matrix; D(i,j) represents the i-th row and j-th column of the absolute depth map D; I(i,j) represents the i-th row and j-th column of the image I; I represents the image I 1 or I 2 ; D represents the absolute depth map D 1 or D 2 .
[0022] In Step 6, the expression of the Gaussian radiation sphere is:
[0023]
[0024] where G i (X) is the distribution of the i-th Gaussian radiation sphere at the spatial position X, μ i is the center position of the i-th Gaussian radiation sphere, e is the natural constant, and Σ i is determined by the rotation quaternion r i and the scaling s i : Σ i = RGG T R T ; where G is the parameter obtained by diagonalizing the scale s i , and R is the rotation by the quaternion r iThe obtained rotation matrix; μ i , r i , s i respectively represent the i-th element disassembled from matrix M p , M r , M s ;
[0025] Among them, the parameters r i , s i and opacity α i of the Gaussian radiation sphere are obtained by the encoder E depth and the decoder :
[0026]
[0027]
[0028] where I is the input image, are the features of different scales for monocular depth extraction, D is the calculated monocular depth; Normalize is the normalization operation, and ClampMax is the maximum truncation operation.
[0029] Step 7 includes:
[0030] Obtaining the color C(i, j) through volume rendering:
[0031]
[0032] where M is the subset of Gaussian points involved in the pixel (i, j), and the elements of the subset are sorted by depth; c k represents the k-th element disassembled from matrix M c ; a j , a k respectively represent the j-th element and the k-th element disassembled from matrix M a .
[0033] In step 8, the following optimization loss function L is established:
[0034] L = αL color + βL ssim + γL d
[0035] where α is the color L 1 loss weight, β is the color structure similarity loss weight, γ is the depth L 1 loss, L color represents the color L 1 loss function, L ssim represents the structure similarity loss function, L d is the depth L1 Loss
[0036] The optimized loss function L is used to calculate the loss, and backpropagation is performed to obtain the gradients of other parameters to be updated, and the parameters to be updated are updated using the gradient descent algorithm.
[0037] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.
[0038] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.
[0039] The present invention also provides a three-dimensional scene real-time reconstruction system, which is characterized by including:
[0040] An input module for receiving training image pairs (I 1 , I 2 ), pose matrices (T 1 , T 2 ) and camera intrinsic matrix K;
[0041] A PatchMatch module for performing PatchMatch in the original pixel space of the image pair (I 1 , I 2 ) to obtain a sparse depth map D r1 , D r2 ;
[0042] A monocular depth prediction module for inputting the image pair (I 1 , I 2 ) into a monocular depth prediction model to obtain a monocular depth feature map
[0043] A feature fusion module for splicing the sparse depth map and the monocular depth feature map and performing feature fusion through a convolutional neural network to output an absolute depth map D 1 , D 2 ;
[0044] A three-dimensional Gaussian distribution construction module for lifting two-dimensional parameters to three-dimensional space using the absolute depth map D 1 , D 2 to construct a preliminary 3D Gaussian distribution;
[0045] A rendering module for rendering images C 1 , C 2 from the perspective of the camera pose matrix T 1 , C2 ;
[0046] An optimization module for traversing more than two image pairs and updating the weights of the convolutional neural network and the Gaussian parameter prediction network until the optimization converges.
[0047] The system further includes a training module for enabling the model to have the generalization ability to perform novel view rendering on dynamic multi-view image pairs of any new scene after the training is completed.
[0048] By introducing a multi-view depth acquisition strategy that combines PatchMatch with monocular depth estimation, the present invention significantly improves the reconstruction accuracy and rendering efficiency of 3D Gaussian point clouds (3DGS) in dynamic scenes. Its advantages and beneficial effects are mainly reflected in the following aspects:
[0049] (1) Real-time performance and efficiency improvement: Compared with traditional iterative MVS methods, the present invention uses PatchMatch for fast matching of local features and provides a complete but possibly scale-biased depth distribution with monocular depth estimation. The two complement each other, enabling the entire depth estimation process to be completed in a short time, greatly reducing the dependence on large-scale cost volumes or long-time iterations, and providing a high-efficiency solution for dynamic scenes or real-time applications.
[0050] (2) Geometric accuracy and robustness: By using the sparse depth obtained by PatchMatch as the global correction basis, the inherent scale ambiguity problem of monocular depth estimation is effectively overcome. Even in complex scenes such as humans with severe texture loss or self-occlusion, the fusion strategy of the present invention can still obtain relatively accurate depth information, thereby significantly improving the geometric reconstruction accuracy of the scene or the person in the 3DGS representation.
[0051] (3) End-to-end differentiable training framework: The complete process of the present invention from multi-view depth to 3DGS to multi-view rendering maintains differentiable properties, supporting the end-to-end optimization of the pixel-level Gaussian parameter prediction network and the feature fusion CNN. During the training process, all network modules are updated collaboratively, gradually improving the synthesis quality of 3D Gaussian point clouds from novel views, providing a solid generalization basis for subsequent fast inference on any new scene or person.
[0052] In summary, by comprehensively utilizing the advantages of PatchMatch and monocular depth estimation, the present invention realizes the three-dimensional reconstruction and novel view synthesis of objects in dynamic scenes or sparse views in an end-to-end differentiable manner, effectively improving the reconstruction speed and geometric accuracy, and providing an effective solution for real-time high-quality 3DGS reconstruction. Description of the Drawings
[0053] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0054] Figure 1 It is a schematic diagram of the method of the present invention.
[0055] Figure 2 It is the implementation result of the present invention under 2K resolution rendering. Specific Embodiments
[0056] An embodiment of the present invention provides a real-time dynamic 3D reconstruction method based on the fusion of patch matching and lightweight monocular depth estimation, which is applied to dynamic scenes or real-time 3D reconstruction tasks. Through steps such as preprocessing of image data, depth estimation, and construction and rendering of 3D Gaussian point clouds, this embodiment realizes efficient and accurate 3D reconstruction. The present invention specifically includes the following steps:
[0057] Step 1: Input data preparation and preprocessing;
[0058] First, input training image pairs (I 1 , I 2 ) and the corresponding pose matrices (T 1 , T 2 ), as well as the internal parameter matrix K of the camera. The key to this step is to collect image pairs with sufficient perspective changes to ensure rich information sources for subsequent 3D reconstruction and depth estimation. In addition, prepare the monocular depth feature extraction model Initialize the PatchMatch for patch matching and the convolutional neural network for monocular depth feature fusion for subsequent depth estimation and feature fusion. Finally, initialize the encoder E depth required for the per-pixel Gaussian parameter prediction model and the decoders
[0059] Step 2: Execute the PatchMatch algorithm;
[0060] As shown in the module "PatchMatch" in Figure 1 , for the image pair (I 1 , I 2 ), execute the PatchMatch algorithm in the original pixel space to calculate the sparse depth maps D r1 , D r2 of each image, as well as the matching confidence S r1 , S r2。The PatchMatch algorithm can efficiently provide depth estimation results for regions with obvious textures through the fast matching of local image patches. This step is particularly suitable for texture-rich regions in complex scenes and can quickly obtain relatively accurate local depth information.
[0061] Step 3: Monocular depth estimation;
[0062] Next, each image I 1 , I 2 is respectively input into the lightweight monocular depth feature extraction model to obtain the monocular depth feature map as shown in the "Monocular Depth Prediction" module in Figure 1 . This model extracts spatial information from a single image through deep learning methods. Although there is scale ambiguity, it can provide useful depth priors in most cases. Training needs to be fully executed for both images, and only the training for one of the images is shown in the figure.
[0063] Step 4: Feature fusion and depth correction;
[0064] In this step, the sparse depth maps D r1 , D r2 , the matching confidence S r1 , S r2 , and the monocular depth feature map are concatenated to form a multi-channel input feature map with the dimension of (S + 2) × H × W. Then, through feature fusion, the absolute depth map D 1 , D 2 after global scale alignment is obtained. This process effectively corrects the scale ambiguity of monocular depth and further improves the accuracy of depth estimation.
[0065] Step 5: Constructing a three-dimensional Gaussian point cloud;
[0066] Using the depth map obtained in Step 4, the two-dimensional pixel depth information is "lifted" into the three-dimensional space to construct a preliminary 3D Gaussian point cloud. The depth information of each pixel is mapped to the Gaussian particle position M p (i, j) and color M c (i, j) in the three-dimensional space, and the formula is as follows:
[0067]
[0068] M c (i, j) = I(i, j)
[0069] Step 6: Gaussian distribution modeling;
[0070] Figure 1 The "3DGS parameters" in it are the parametric modeling of a three-dimensional Gaussian radiation sphere, and the expression of the three-dimensional Gaussian radiation sphere is:
[0071]
[0072] where G i (X) is the distribution of the i-th Gaussian radiation sphere at the spatial position X, and μ i is the center position of the i-th Gaussian radiation sphere, e is the natural constant, and Σ i is determined by the rotation quaternion r i and the scaling s i : Σ i = RGG T R T ; where G is the parameter obtained by diagonalizing the scale s i , and R is the rotation matrix obtained by converting from the quaternion rotation r i ; μ i , r i , s i respectively represent the i-th elements disassembled from the matrices M p , M r , M s .
[0073] As Figure 1 shown in the "3DGS parameter prediction network" module in, this module outputs the parameters r i , s i and the opacity α i . Specifically, this module consists of an encoder E depth and a decoder , and the method for calculating the parameters of the Gaussian radiation sphere is:
[0074]
[0075] Step 7: View rendering;
[0076] Corresponding to Figure 1 the rendering module in, using the obtained three-dimensional Gaussian point cloud, renders the new view T 1 , T 2 to obtain the rendered image C 1 , C 2 . The formula for obtaining the color C(i, j) of the pixel (i, j) by volume rendering is:
[0077]
[0078] where M is a subset of Gaussian points involved in pixel (i, j), and the elements of the subset are sorted by depth. During the rendering process, the final rendered image is obtained by weighted summation of the colors and transparencies of all Gaussian points.
[0079] Step 8: Loop optimization and training;
[0080] The above steps will be repeated. By traversing multiple image pairs, the weights of the convolutional neural network and the pixel-level Gaussian parameter prediction network are continuously updated until the optimization process converges. The overall optimization loss function L used in the training process consists of the following parts:
[0081] L = αL color + βL sdim + γL d
[0082] where α is the color L 1 loss weight, β is the color structure similarity loss weight, γ is the depth L 1 loss, L cslor represents the color L 1 loss function, L ssim represents the structure similarity loss function, L d is the depth L 1 loss.
[0083] After training is completed, the obtained model can perform Gaussian distribution regression on new scenes and targets, thereby realizing 3D reconstruction and new view synthesis of any new scene. The trained model has strong generalization ability, can quickly respond to dynamic scene changes, and provide real-time 3D reconstruction results.
[0084] The present invention proposes to use PatchMatch to mine the depths of reliably matchable multi-view feature points, and use the results as the "scale reference" for lightweight monocular depth estimation, thereby realizing the correction of monocular depth on a global scale. In pixel regions rich in texture or feature information, PatchMatch can obtain relatively accurate and stable multi-view matching depths. For regions where PatchMatch is difficult to match or there is texture degradation, monocular depth estimation can give reasonable depth initial values. By using the sparse depth obtained by PatchMatch as a constraint, the monocular depth output can be aligned on a global scale, avoiding large-scale scale drift generated by the monocular network. Compared with traditional iterative MVS, the local feature search of PatchMatch and monocular depth alignment can complete the global depth inference in a shorter time, which helps to achieve fast processing in dynamic scenes or real-time applications.
[0085] Finally, using the multi-view depth results obtained by the above fusion, the Gaussian particle parameters (position, color, opacity, variance matrix, etc.) learned in the view plane can be "lifted" into the three-dimensional coordinate system, and the three-dimensional Gaussian point clouds from multiple perspectives can be aggregated to complete the high-fidelity three-dimensional reconstruction and new view rendering of the person or scene. At the same time, this process is differentiable from multi-view depth to 3D Gaussian Splatting and then to multi-view RGB rendering, laying a foundation for end-to-end training of the prediction model of pixel-level three-dimensional Gaussian particle parameters. Since PatchMatch combined with the lightweight monocular depth estimation scheme takes into account both speed and accuracy in the process of depth information acquisition and has a strong matching ability for key feature points in complex self-occlusion scenarios, it can effectively replace the time-consuming MVS process in GPS-Gaussian, thus meeting the real-time or near-real-time application requirements. In the Nvidia-A6000 graphics card environment, it can reach 13 FPS under the condition of inputting 1K resolution and rendering 2K resolution, and 33 FPS under the condition of inputting 512×512 resolution and rendering 1K resolution.
[0086] The embodiments of the present invention are trained and verified on the THuman2.0 dataset, and the verification results are as Figure 2 shown, where the left view and the right view are the inputs of the model, and the rendering results of the new view are the outputs of the model.
[0087] In summary, on the premise of integrating PatchMatch multi-view matching and monocular depth estimation, the present invention effectively overcomes the scale ambiguity of monocular depth estimation, and while greatly reducing the computational amount, it maintains relatively reliable depth accuracy. Combining the fast rendering characteristics of 3DGS, the present invention can achieve real-time three-dimensional reconstruction and new view synthesis of the target person or scene in a sparse multi-view scenario, providing a new and efficient technical solution for practical applications such as sports broadcasting, virtual conferences, and interactive entertainment.
[0088] The present invention successfully solves the problems existing in multi-view depth estimation, such as large computational amount, slow speed, and inconsistent depth scale, by combining the PatchMatch algorithm with the lightweight monocular depth estimation technology. At the same time, by using the 3DGS method to achieve efficient three-dimensional reconstruction, the balance between real-time performance and high accuracy is achieved. This technology provides a practical solution for real-time three-dimensional reconstruction in the fields of virtual reality, sports broadcasting, interactive entertainment, etc.
[0089] The present invention provides a real-time dynamic three-dimensional reconstruction method and system based on the fusion of patch matching and lightweight monocular depth estimation. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A real-time dynamic 3D reconstruction method based on the fusion of patch matching and lightweight monocular depth estimation, characterized in that: The following steps are involved: Step 1: Prepare a monocular depth feature extraction model Initialize the convolutional neural network for patch matching PatchMatch and monocular deep feature fusion Prepare the encoder E required for the pixel-by-pixel Gaussian parameter prediction model depth and decoders for predicting rotation, scale, and opacity of Gaussian points, respectively Step 2: Input the training image pair (I1, I2) and the corresponding pose matrix (T1, T2), as well as the camera's intrinsic parameter matrix K; where I1, I2 represent the first training image and the second training image input, respectively, and T1, ...2 represent the pose matrix corresponding to the first training image I1 and the pose matrix corresponding to the second training image I2, respectively; Step 3: Perform patch matching PatchMatch in the original pixel space of the image pair (I1, I2) to obtain the sparse depth map D of the matching area of I1 and I2 respectively. r1 ,D r2 , and matching confidence S r1 ,S r2 ; Step 4: Input the image pair (I1, I2) into the monocular depth prediction model In the above example, we get the monocular depth feature map of I1. And the monocular depth feature map of I2 Where S is the dimension of the monocular depth feature map; Step 5: transform the sparse depth map I r1 ,D r2 , matching confidence S r1 ,S r2 , Monocular depth feature map Spliced into multi-channel input, forming a fusion feature of (S+2)×H×W dimensions, and then passed through a convolutional neural network The fused features are fused based on confidence, and the absolute depth maps D1 and D2 of the images I1 and I2 after global scale alignment are output; H and W are the height and width of the feature map, respectively, and the sparse depth map D r1 ,D r2 and matching confidence S r1 ,S r2 It is necessary to use bilinear interpolation to the height and width of the monocular depth feature map for merging; Step 6: Use the absolute depth maps D1 and D2 to lift the two-dimensional parameters to three-dimensional space, construct a preliminary 3D Gaussian distribution, and pass it through the encoder E depth and decoder Get the pixel-by-pixel position matrix M p , color matrix M c , rotation matrix M r , scaling matrix M s and the opacity matrix M a ; The color matrix of Gaussian points is obtained from the training image pair (I1, I2); Step 7, obtain images C1 and C2 under the perspectives of camera pose matrices T1 and T2 respectively through volume rendering; Step 8: Repeat steps 2 to 7, traverse the image pairs at more than two moments, and update the convolutional neural network. and Gaussian parameters predict the parameters of the model until the optimization converges.
2. The method according to claim 1, characterized in that In step 4, the monocular depth feature map and Through the monocular depth prediction model The calculation formula is:
3. The method according to claim 2, characterized in that In step 6, the two-dimensional parameters are lifted to three-dimensional space using the following formula: M c (i,j)=I(i,j) Among them, M p (i,j) represents the position matrix M p The i-th row and j-th column of M c (i,j) represents the i-th row and j-th column of the color matrix; D(i,j) represents the i-th row and j-th column of the absolute depth map D; I(i,j) represents the i-th row and j-th column of the image I; I represents the image I1 or I2; D represents the absolute depth map D1 or D2.
4. The method according to claim 3, characterized in that In step 6, the expression of the Gaussian radiation sphere is: Among them G i (X) is the distribution of the i-th Gaussian radiation ball at the spatial position X, μ i is the center position of the i-th Gaussian radiation sphere, e is a natural constant, Σ i By the rotation quaternion r i and zooms i Decision: Σ i =RGG T R T ; Where G is the scale s i The parameters obtained by diagonalization, R is the quaternion rotation r i The rotation matrix obtained by transformation; μ i ,r i ,s i Respectively represent the matrix M p ,M r ,M s The i-th element disassembled from above; The parameter r of the Gaussian radiation sphere is i 、s i and opacity α i By encoder E depth and decoder get: Where I is the input image, are features of different scales extracted from monocular depth, D is the calculated monocular depth; Normalize is the normalization operation, and ClampMax is the maximum value truncation operation.
5. The method according to claim 4, characterized in that Step 7 includes: Obtain color C(i,j) through volume rendering: Where M is the subset of Gaussian points involved in pixel (i, j), and the elements of the subset are sorted by depth; c k Represents the matrix M c The kth element disassembled; a j ,a k Respectively represent the matrix M a The j-th element and the k-th element are disassembled.
6. The method according to claim 5, characterized in that In step 8, the following optimization loss function L is established: L=αL color +βL ssim +γL d Among them, α is the color L1 loss weight, β is the color structure similarity loss weight, γ is the depth L1 loss, L color represents the color L1 loss function, L ssim represents the structural similarity loss function, L d is the deep L1 loss.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A three-dimensional scene real-time reconstruction system implemented by the method according to any one of claims 1 to 6, characterized in that: include: An input module is used to receive training image pairs (I1, I2), pose matrices (T1, T2) and camera intrinsic parameter matrix K; The PatchMatch module is used to perform patch matching PatchMatch in the original pixel space of the image pair (I1, I2) to obtain a sparse depth map D r1 ,D r2 ; The monocular depth prediction module is used to input the image pair (I1, I2) into the monocular depth prediction model to obtain the monocular depth feature map Feature fusion module, used to splice the sparse depth map with the monocular depth feature map and pass it through the convolutional neural network Perform feature fusion and output absolute depth maps D1 and D2 after global scale alignment; The three-dimensional Gaussian distribution construction module is used to use the absolute depth maps D1 and D2 to lift the two-dimensional parameters to the three-dimensional space and construct a preliminary 3D Gaussian distribution; The rendering module is used to render images C1, T2 from the perspective of the camera pose matrix T1, T2; The optimization module is used to traverse more than two image pairs and update the weights of the convolutional neural network and the Gaussian parameter prediction network until the optimization converges.
10. The system according to claim 9, characterized in that The system also includes a training module, which is used to enable the model to have the generalization ability to perform new perspective rendering on any new scene dynamic multi-perspective image pair after the training is completed.
Citation Information
Patent Citations
Deep learning-based image laser data fusion method for building reconstruction
CN115423978A
Neural rendering method, system and equipment based on depth unbiased estimation
CN117745924A
Method for synthesizing new view angle of lunar surface based on Gaussian radiation field
CN118469836A
Gaussian radiation field three-dimensional reconstruction method based on pre-training optical flow model geometric distillation
CN118887346A
Method for 3D scene dense reconstruction based on monocular visual slam
US20200273190A1
Cited By
Large-scene three-dimensional reconstruction method based on three-dimensional Gaussian sputtering
CN120472121A
Monocular dynamic scene reconstruction method and system based on self-supervised flow matching
CN120833442A
Image-driven three-dimensional scene reconstruction method and device without SfM initialization
CN121304906A
Image-driven three-dimensional scene reconstruction method and apparatus without sfm initialization
CN121304906B
Vector space video stream and three-dimensional reconstruction collaborative scene construction method and system
CN122176208A