Dense visual field reconstruction method and system based on MViT
The Single-view ViMGS-SLAM framework addresses scale drift and ambiguity in 3DGS-SLAM by using MViT for dense visual scene reconstruction, ensuring geometric and scale consistency for real-time high-precision 3D reconstruction.
Patent Information
- Application Number
- CN202510804154.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The monocular 3DGS-SLAM system faces the problems of sparse observation constraints, finite triangulation baselines and weak pose estimation coupling during the initial mapping stage, which leads to error-prone Gaussian initialization and the challenge of scale ambiguity and real-time high-precision reconstruction in dynamic scenarios.
The dense visual field reconstruction method based on MViT is adopted, and the depth estimation is carried out through the monocular ViMGS-SLAM framework combined with the MViT module, the 3D Gaussian representation module is reparameterized, the camera tracking module optimizes the position, the keyframe management module selects the keyframe, and the dense visual field reconstruction is carried out through the graph building module, solving the problem of scale drift and real-time high-precision reconstruction.
It realizes high-precision reconstruction of dense visual fields in dynamic scenarios, solves the problems of scale drift accumulation and real-time high-precision reconstruction of monocular systems, and improves the reliability and reconstruction accuracy of the system.
Smart Images

Figure CN120318391A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method and system for dense visual field reconstruction based on MViT (Multiscale Vision Transformers). Background Art
[0002] Visual Simultaneous Localization and Mapping (VSLAM) is an advanced technology in visual perception technology and has attracted much attention in the field of robot research. VSLAM can not only provide environmental perception and autonomous navigation capabilities for autonomous driving devices, but also achieve three-dimensional reconstruction of large areas of the earth's surface and remote sensing mapping. VSLAM estimates the position of a robot relying on visual information in an unknown environment while reconstructing the surrounding environment, which has important theoretical significance and application value.
[0003] In recent years, 3D Gaussian Splatting (3DGS) has emerged as a novel explicit scene representation and differentiable rendering paradigm, attracting extensive attention in the fields of computer vision and 3D graphics. Different from traditional point cloud or mesh representations, 3DGS parameterizes spatial points through anisotropic 3D Gaussian distributions with covariance matrices, improving the geometric fidelity of scene modeling through adaptive density control. This representation method not only facilitates realistic novel view synthesis but also supports real-time rendering capabilities, effectively meeting the key requirements of VSLAM. Pioneering studies such as GS-SLAM and SplaTAM have successfully integrated 3DGS into the RGB-D SLAM (Red Green Blue-depth Simultaneous Localization And Mapping) pipeline, showing superior performance metrics compared to traditional neural implicit methods. Additionally, the proposal of a monocular SLAM (Simultaneous Localization And Mapping) framework based entirely on 3DGS has also become an important milestone, which achieves real-time incremental reconstruction through optimizing Gaussian parameter updates. However, monocular 3DGS-SLAM systems face inherent challenges during the initial mapping stage, including sparse observation constraints, limited triangulation baselines, and weak pose estimation coupling, resulting in error-prone Gaussian initialization. To address these issues, YAN et al. proposed a hybrid architecture that combines direct sparse photometric measurements with 3DGS to achieve denser and more structurally coherent point cloud generation through geometric consistency constraints. HU et al. further developed a dynamic Gaussian management strategy to achieve real-time error pruning, observation fusion, and parameter refinement guided by photometric confidence masks. Scale ambiguity in monocular depth estimation is another key challenge. ZHANG et al. introduced a scale-aware alignment module that combines depth regularization based on multi-resolution grids and derived a closed-form solution for ray-Gaussian intersection depth calculation. LEE et al. proposed view-specific scale parameterization with anchors to achieve online scale calibration by minimizing scale-consistent depth loss.
[0004] Despite significant progress in 3D reconstruction quality, monocular 3DGS-SLAM is still in its infancy and faces numerous challenges, such as scale drift accumulation during long-term system operation and system reliability issues caused by the inherent scale ambiguity problem of the system, as well as the problem of real-time and high-precision 3D reconstruction of monocular RGB inputs in dynamic scenes. Therefore, developing an efficient monocular 3DGS-SLAM framework with metric scale consistency, reliable feature tracking, and dynamic scene adaptability is an important issue to be addressed in further research in this field. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a method and system for dense visual field reconstruction based on MViT. The technical problems to be solved by the present invention are realized through the following technical solutions: In a first aspect, the present invention proposes a method for dense visual field reconstruction based on MViT. This method is implemented based on the monocular ViMGS-SLAM (A Real-Time Monocular 3DGS-based SLAM via Multiscale Vision Transformers) framework. The above monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module. The method includes: Input the image to be reconstructed into the monocular ViMGS-SLAM framework. When the image to be reconstructed only includes a monocular RGB image, in the MViT module, depth estimation is performed on the monocular RGB image through a cross-scale attention mechanism to obtain a depth image. When the image to be reconstructed also includes an RGB-D image, the RGB-D image is used as the depth image; In the 3D Gaussian representation module, reparameterization processing is performed on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume; Based on the monocular RGB image and the depth image, camera pose optimization is performed in the camera tracking module to obtain camera pose information; Based on the camera pose information and the 3D Gaussian volume, key frame selection is performed in the key frame management module to obtain key frame information; Based on the key frame information and the 3D Gaussian volume, new view rendering is generated through isotropic regularization in the mapping module, thereby realizing dense visual field reconstruction.
[0006] In a second aspect, the present invention proposes a system for dense visual field reconstruction based on MViT, which is used to implement the method proposed in the first aspect of the present invention. This system is equipped with a monocular ViMGS-SLAM framework. The above monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module. Among them, The MViT module is used to perform depth estimation on the monocular RGB image in the image to be reconstructed through a cross-scale attention mechanism to obtain a depth image; The 3D Gaussian representation module is used to perform reparameterization processing on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume; The camera tracking module is used to perform camera pose optimization based on the monocular RGB image and the depth image to obtain camera pose information; A key - frame management module, which is used to select key - frames based on camera pose information and 3D Gaussian volumes to obtain key - frame information; A mapping module, which is used to generate new - view rendering through isotropic regularization based on key - frame information and 3D Gaussian volumes, thereby realizing dense visual - field reconstruction.
[0007] Advantages of the present invention: 1. A dense visual - field reconstruction method based on MViT provided by the present invention constructs a monocular ViMGS - SLAM framework, which includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key - frame management module, and a mapping module; for the input monocular RGB image, first, in the MViT module, a cross - scale attention mechanism is used to extract photometric features in the cross - spatial - frequency domain, thereby enhancing the geometric consistency of depth prediction and realizing accurate depth estimation; then, by combining the geometric modeling ability of the 3D Gaussian representation module, the camera pose optimization of the camera tracking module, the key - frame selection of the key - frame management module, and the parameter update and new - view rendering of the mapping module, dense visual - field reconstruction in a dynamic scene is realized, solving the problems of scale - drift accumulation in the long - term operation of the monocular system and real - time high - precision 3D reconstruction of monocular RGB input; meanwhile, this framework can also achieve real - time high - precision reconstruction on RGB and RGB - D data streams.
[0008] 2. In the dense visual - field reconstruction method based on MViT provided by the present invention, aiming at the continuous scale - ambiguity problem of the monocular system in the MViT module, a multi - scale ViT based on geometric - awareness enhancement is designed, and combined with a multi - scale feature aggregation mechanism, a scale - invariance constraint is enforced through a hierarchical feature - fusion process, enhancing the geometric consistency of depth prediction, and further improving the system reliability and reconstruction accuracy.
[0009] The present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0010] Figure 1 It is a schematic flowchart of a dense visual - field reconstruction method based on MViT provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of a monocular ViMGS - SLAM framework provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of the framework of the MViT module provided by an embodiment of the present invention; Figure 4 It is the scene - reconstruction result of the ViMGS - SLAM framework proposed by the present invention and the existing MonoGS framework on the Office3 and Room2 sequences of the Replica dataset. Detailed Embodiments
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0012] In VSLAM research, reconstructing a metrically consistent three-dimensional environment from a monocular image sequence remains a fundamental challenge. A large number of studies have shown that due to the stronger spatial constraints of elliptical primitives, their geometric positioning accuracy is superior to that of traditional point or line features. Zwicker et al. made pioneering progress by introducing the Elliptically Weighted Average (EWA) stitching technique - a complex rendering framework using an anisotropic Gaussian kernel. Their research demonstrated that the EWA stitching technique can achieve alias-free reconstruction while preserving sharp discontinuities in volume data and effectively handling non-spherical scattering effects. This method demonstrated excellent versatility on regular, curved, and unstructured volume datasets, establishing a unified framework for multi-modal data processing. Subsequent developments in ellipse-based pose estimation and stereo reconstruction expanded the application scope of dynamic scene understanding. However, practical applications face some limitations: these limitations hinder SLAM systems from adopting elliptical structures. The emergence of the 3DGS technology has completely changed this pattern, achieving unified optimization of geometry and appearance through variable volume rendering, significantly improving the SLAM reconstruction efficiency and visual quality. Despite these advancements, the implementation of monocular 3DGS-SLAM has not been fully explored, especially in terms of scale consistency in depth estimation and computational complexity management.
[0013] Based on this, the first aspect of the present invention constructs a monocular ViMGS-SLAM framework and proposes a dense visual field reconstruction method based on MViT. This method addresses these limitations through three key innovations: 1) a hybrid depth prediction network that combines monocular geometric priors with multi-scale Gaussian uncertainty modeling; 2) a neural-Gaussian fusion framework for realistic novel view synthesis; 3) an adaptive sparsification algorithm that enables real-time performance on RGB and RGB-D data streams.
[0014] Please refer jointly to Figure 1 and Figure 2 , Figure 1 which is a schematic flow diagram of a dense visual field reconstruction method based on MViT provided by the embodiments of the present invention, Figure 2 and is a schematic diagram of the monocular ViMGS-SLAM framework provided by the embodiments of the present invention. A dense visual field reconstruction method based on MViT proposed by the present invention is based on Figure 2The implementation of the monocular ViMGS-SLAM framework is shown. Among them, the monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module (i.e., the 3DGS representation module in Figure 2 , a camera tracking module (i.e., the tracking module in Figure 2 ), a key-frame management module (i.e., the key-frame module in Figure 2 ), and a mapping module. Then, the method for dense visual field reconstruction based on MViT specifically includes the following steps: Step 1: Input the image to be reconstructed into the monocular ViMGS-SLAM framework. When the image to be reconstructed only includes a monocular RGB image, in the MViT module, depth estimation is performed on the monocular RGB image through a cross-scale attention mechanism to obtain a depth image. When the image to be reconstructed also includes an RGB-D image, the RGB-D image is used as the depth image.
[0015] Specifically, when the image to be reconstructed only includes a monocular RGB image, that is, when there is no available depth information, the MViT module can be used to perform depth estimation on the monocular RGB image. By using a multi-scale attention mechanism, a dense depth prior is generated to ensure geometric consistency in depth estimation without affecting temporal consistency, and a depth image is obtained. When the image to be reconstructed includes both an RGB image and an RGB-D image at the same time, the RGB-D image contains directly available depth information. Therefore, the RGB-D image can be directly used as the depth image.
[0016] In this embodiment, the MViT module is a multi-scale vision transformer architecture proposed to meet the key requirements for geometrically accurate 3D reconstruction in monocular VSLAM. Please refer to Figure 3 , Figure 3 is a schematic diagram of the framework of the MViT module provided by an embodiment of the present invention. This architecture mainly includes a multi-scale feature encoding and position embedding module, an encoder module, a decoder module, and a depth information generation module. This architecture systematically preserves the metric consistency in object topology, terrain geometry, and depth estimation, integrates a pre-trained ViT encoder with a multi-scale feature aggregation and edge-aware geometry refinement module, and is specifically optimized for monocular depth prediction. Each module can specifically perform depth estimation according to the following steps 11)-14).
[0017] 11) The multi-scale feature encoding and position embedding module divides the monocular RGB image into multi-scale image blocks with different resolutions after symmetric padding, and performs linear projection on the image blocks of each scale to generate multi-scale feature vectors.
[0018] Please refer to Figure 3Understanding of the upper left module. First, the input RGB image with an original resolution of 640×480 pixels is symmetrically padded along the vertical dimension to obtain an RGB image after symmetric padding processing, and its resolution is adjusted to 640×640 pixels. The padding operation is to ensure an integer multiple for subsequent segmentation operations.
[0019] Then, using the multi-scale processing framework, the RGB image after symmetric padding processing is divided into three levels, which are described in steps A1, A2, and A3 respectively. These three steps can be processed in parallel. The purpose of dividing the input image into multiple scales is to capture multi-level features. For example, features at high resolution can retain local details (such as edges, textures), and features at low resolution can capture a wider context (such as object shape, overall layout).
[0020] Step A1: The RGB image after symmetric padding processing is divided into multiple image patches in an overlapping division manner as the first-scale image patches; Please refer to Figure 3 the branch where s = 1 in
[0021] In this branch, the RGB image after symmetric padding processing with a resolution of 640×640 pixels is divided into 16 image patches of the same size in an overlapping division manner of image patches as the first-scale image patches. Each image patch is a patch.
[0022] Step A2: The RGB image after symmetric padding processing is downsampled by a factor of 2 and then divided into multiple image patches in a non-overlapping division manner as the second-scale image patches; Please refer to Figure 3 the branch where s = 2 in
[0023] In this branch, the RGB image after symmetric padding processing with a resolution of 640×640 pixels is first downsampled by a factor of 2, making the resolution become 320×320 pixels; then it is divided into 16 image patches of the same size in a non-overlapping division manner of image patches as the second-scale image patches.
[0024] Step A3: The RGB image after symmetric padding processing is downsampled by a factor of 4, and the resulting image is used as the third-scale image patch; Please refer to Figure 3 the branch where s = 3 in
[0025] In this branch, the RGB image after symmetric padding processing with a resolution of 640×640 pixels is downsampled by a factor of 4, so that the resolution becomes 160×160 pixels, serving as the third-scale image patch.
[0026] Finally, the image patches at each scale are linearly projected into feature vectors with trainable positional encoding to obtain the feature vectors at each scale.
[0027] Linear projection means that each patch is mapped into a vector of a fixed dimension through a fully connected layer, called an embedding vector.
[0028] Positional encoding means adding positional information (such as coordinates, scale) to each embedding vector to endow it with spatial semantics.
[0029] After linear projection and positional encoding, the feature vectors at each scale are expressed as: ; In the formula, the input RGB image is expressed as , , represents that the mathematical property of the data type is the set of real numbers, and respectively represent the height and width of the input RGB image; represents the feature vector at the th scale, ; , representing a learnable embedding matrix; represents the image patch at the th scale; , and both represent the size of the image patch at the th scale; is the spatial dimension at the th scale, represents the dimension of the embedded feature vector, that is, through linear projection, each image patch is mapped into a -dimensional latent space; is a flattening function for multi-dimensional arrays, used to convert the image patch into a one-dimensional feature vector. This step integrates spatial locality and global position perception, which not only conforms to the paradigm of Vision Transformer (ViT), but also can handle geometric changes at a specific scale.
[0030] This part can be understood in combination with the relevant content of the ViT architecture and will not be elaborated here.
[0031] 12) The encoder module adopts a hierarchical Transformer and a cross-scale attention mechanism to obtain the feature vector after fusing all multi-scale feature vectors, and uses a local branch for feature extraction to obtain local multi-scale features. At the same time, the global branch is used to extract features from the feature vector of the lowest resolution scale in the multi-scale feature vectors to obtain global features.
[0032] Specifically, please refer to Figure 3 the module understanding in the upper right corner. The encoder module is implemented by a hierarchical Transformer and a cross-scale attention mechanism.
[0033] After merging the feature vectors of the three scales, 33 patches are obtained, and then they enter the local branch for feature extraction to obtain local branch feature information. At the same time, the feature vector of the third scale with the lowest resolution enters the global branch for feature extraction to obtain global branch feature information. Then, the local branch feature information and the global branch feature information are merged to obtain a set of feature information containing different resolutions. The set of feature information is shown in the box on the right side of the global branch, which contains feature information with different resolutions. The red square on the far left is the global branch feature information output by the global branch, and the remaining squares are the merged local branch feature information.
[0034] The encoder module follows the standard ViT principle, and local-global feature interaction is modeled by a two-branch Transformer. Both the local branch and the global branch are implemented using the architecture of the multi-head attention mechanism (Multi-Head Self-Attention, MSA) combined with a multi-layer perceptron (Multi-Layer Perceptron, MLP). Layer normalization (LN) is used to stabilize the activation, and the multi-head attention mechanism is used to simulate context dependencies; there are L layers of transformer layers stacked in sequence. The local branch uses 6 attention heads, and the global branch uses 12 attention heads; the multi-layer perceptrons of the local branch and the global branch include two fully connected layers, and the Gaussian Error Linear Unit (GELU) is used as the intermediate activation function.
[0035] In the L layers of transformer layers stacked in sequence, the processing of the layer is expressed as: ; ; In the formula, represents the multi-head attention mechanism; represents layer normalization, which is used to standardize the features, accelerate the model convergence and stabilize the training process; represents the multi-layer perceptron; Represents the basic features of the input of the layer, and information is transmitted through residual connections; Represents the intermediate features of the layer after processing;
[0036] 13) The decoder module performs local-global feature interaction and adaptive feature fusion on local multi-scale features and global features based on a hierarchical transformer with cross-scale attention, and introduces a multi-layer perceptron for progressive upsampling processing during the fusion process to restore the corresponding image information, and then uses the DPT decoder to reconstruct the inverse depth image.
[0037] Please refer to Figure 3 the module understanding in the lower right corner. The decoder module fuses the feature information set through cross-scale feature interaction based on the attention mechanism to obtain the fused feature information, which mainly includes feature alignment, dynamic weight adjustment, and non-linear interaction. The specific processes are as follows: Through cross-scale feature alignment based on the attention mechanism, the dynamic weights generated by the interaction between the global query vector and local key-value pairs are applied to features at each scale to achieve context-aware multi-scale feature fusion and obtain the fused feature information. This part corresponds to Figure 3 the adaptive feature fusion in By the cross-attention mechanism between the global query vector and local key-value pairs , the framework proposed by the present invention can dynamically adjust the weight coefficients of multi-scale features in a content-dependent manner to achieve context-aware adaptive fusion, thereby optimizing the perceptual granularity of hierarchical feature representations at different spatial resolutions. Among them, represents the spatial dimension of the global query vector . The local branch feature information in the feature information set serves as multi-resolution features , and the global branch feature information is the global context feature , and they are fused through an alignment method based on the attention mechanism.
[0038] Among them, the calculation formula for the fused feature information is: ; In the formula, represents the fused feature information, corresponding to Figure 3 in below the upsampling, representing the restored image information; represents the total number of multi-scales; represents the Local key of scale; Indicates transpose; Indicates the dimension of the feature vector after embedding; Indicates the Feature information of scale. The framework of the present invention mimics the biological visual system, and features from coarse to fine are dynamically weighted to achieve context coherence.
[0039] The decoder module uses the introduced multi-layer perceptron MLP to perform progressive upsampling on the fused feature information to restore the corresponding image information, and then uses the DPT decoder to reconstruct the inverse depth map. The process includes: Using the MLP to non-linearly refine the skip connection features, enhancing the edge alignment ability by combining the edge mask, and gradually upsampling through transposed convolution and feature alignment operations, and finally using the DPT decoder to reconstruct the inverse depth map; the formula used is: ; In the feature fusion and upsampling process of the decoder, the present invention introduces an MLP to non-linearly refine the features of the skip connection to enhance the edge alignment ability.
[0040] In the above formula, the DPT decoder is an existing open-source module. Transposed convolution and skip connection are the core operations in the upsampling process; Indicates the Image information restored at scale; Indicates the multi-layer perceptron; , Indicates the Skip connection features corresponding to scale. Those skilled in the art can understand that the skip connection refers to the connection part that directly connects the feature map of the encoder (feature extraction) part to the decoder (image reconstruction) part; Is the Spatial dimension of scale, Indicates the dimension of the feature vector after embedding; , Indicates the Edge mask corresponding to scale; Indicates element-wise multiplication to embed the edge information into the feature map. In the decoder module of the embodiment of the present invention, the non-linear interaction between the edge mask and the multi-scale features can improve the edge accuracy.
[0041] The decoder module uses transposed convolution and skip connection to reconstruct the inverse depth map. The calculation formula of the inverse depth map is expressed as: ; ; In the formula, denotes a transposed convolution; denotes a feature alignment operation; The image features after performing transposed convolution and feature alignment operations on the image information restored at the denotes element-wise addition; denotes an inverse depth map; is the Sigmoid function; is a 1×1 convolution kernel used to perform channel feature transformation; denotes the image features obtained at all scales; denotes a bias term. The skip connection can reduce information loss and retain high-frequency geometric details. The specific integration method of the skip connection is channel concatenation. The kernel size of the transposed convolution is 3×3, and the stride is 2×2.
[0042] 14) The depth information generation module generates a metric depth map representing absolute depth information and a refined edge depth map with the same resolution as the RGB image based on the inverse depth map as the depth image.
[0043] Please refer to Figure 3 the module understanding in the lower left corner, which obtains absolute depth based on geometric constraints.
[0044] To obtain a dense metric depth map , the present invention performs scaling through the horizontal field of view, and the horizontal field of view is represented by the focal length and the width .
[0045] Specifically, the process of the depth information generation module generating a metric depth map representing absolute depth information based on the inverse depth map includes: Generating a metric depth map representing absolute depth information based on the inverse depth map and geometric constraints; expressed by the formula: ; In the formula, denotes the metric depth map; denotes the camera focal length; denotes the projected width of the object on the image plane.
[0046] In the depth information generation module, an initial refined edge depth map ( Figure 3 with a medium resolution of 640×640 pixels) is generated based on the inverse depth map, and then cropped to obtain a refined edge depth map with the same resolution as the RGB image ( Figure 3 with a medium resolution of 640×480 pixels).
[0047] In response to the persistent scale ambiguity problem in monocular systems in the MViT module, this invention designs a multi-scale ViT enhanced by geometric perception, combines a multi-scale feature aggregation mechanism, enforces scale invariance constraints through a hierarchical feature fusion process, enhances the geometric consistency of depth prediction, and thereby improves system reliability and reconstruction accuracy. The proposed MViT framework includes three key innovations: 1) a hierarchical cross-attention mechanism that can uniformly model local-global context dependencies, 2) a multi-scale adaptive feature alignment module that solves inter-scale feature differences through learnable fusion weights, and 3) geometric constraint boundary refinement with sub-pixel accuracy to enhance spatial fidelity.
[0048] Step 2: In the 3D Gaussian representation module, reparameterize the monocular RGB image based on the depth image to obtain a 3D Gaussian volume.
[0049] To strictly standardize the proposed 3DGS-based SLAM framework, this embodiment integrates its core into several key equations, emphasizing compactness, computational efficiency, and geometric fidelity. Specifically, the 3D Gaussian representation module can obtain a 3D Gaussian volume according to the operations in the following steps 21)-24).
[0050] 21) Initialize the three-dimensional scene based on the monocular RGB image to represent the three-dimensional scene as a coupled representation of a set of anisotropic 3D Gaussian volumes, and decompose the three-dimensional covariance matrix of each 3D Gaussian volume into rotation and scaling operations.
[0051] Specifically, to achieve efficient and high-precision three-dimensional scene modeling and real-time rendering, represent the three-dimensional scene as a coupled representation of a set of anisotropic 3D Gaussian volumes G , which is specifically represented as: ; Among them, a single Gaussian volume contains three attributes, the position coordinates in the world coordinate system , the three-dimensional covariance matrix , and the opacity , that is, each anisotropic 3D Gaussian volume is defined by its position , covariance , and opacity , represents the total number of 3D Gaussian volumes. Compared with the standard Gaussian representation function, 3DGS omits the scale coefficient of the exponential term (which does not affect the ellipsoidal geometric characteristics), uses the origin of the model coordinate system as the default center for easy rotation and scaling operations, and only applies a translation transformation when mapping to the world space.
[0052] To ensure differentiability, the covariance is decomposed into rotation and scale , namely: .
[0053] 22) Project each 3D Gaussian volume onto the imaging plane. After sorting based on depth information in the depth image, use end-to-end optimization for scene rendering to obtain the rendered Gaussian volume.
[0054] First, project each 3D Gaussian volume onto the imaging plane through the camera intrinsic parameters and camera extrinsic parameters. This can relate the 3D scene geometry to the 2D image observation, which is crucial for differentiable rendering. The mean and covariance of the resulting 2D Gaussian volume are expressed as: ; ; In the formula, represents the mean of the th 2D Gaussian volume, represents the perspective projection operation, represents the camera intrinsic parameters, represents the camera extrinsic parameters, represents the rotation matrix, represents the mean of the th 3D Gaussian volume, represents the th covariance of the 2D Gaussian volume, represents the th covariance of the 3D Gaussian volume, represents the Jacobian matrix in the projection operation, and the superscript represents the matrix transpose.
[0055] Then, sort the 2D Gaussian volumes based on depth information in the depth image to obtain N ordered Gaussians.
[0056] Next, calculate the color of each pixel by blending multiple ordered Gaussians along the ray from the front end to the back end to achieve scene rendering and obtain the rendered Gaussian volume; among them, the calculation formula for the pixel color is: ; In the formula, represents the color of pixel , represents the number of Gaussian volumes, represents the opacity of the th 2D Gaussian volume, represents calculating the 2D Gaussian parameters at pixel , represents the th mean of the 2D Gaussian volume, represents the Covariance of a 2D Gaussian body, indicating the color of the th 2D Gaussian body, and opacity of the
[0057] 23) Generate a binary mask based on learnable RGB mask parameters to adaptively prune the rendered Gaussian bodies, obtaining the pruned Gaussian bodies.
[0058] Furthermore, since the original 3DGS method exhibits three key limitations in SLAM applications. First, redundant Gaussians that have little impact on the scene rendering quality proliferate within the system. Second, due to subpixel space effects, Gaussian distributions that are too small have a minimal impact on visual fidelity. Third, low-opacity Gaussians incur unnecessary computational overhead while providing marginal improvement to the final rendered output.
[0059] To address these inefficiencies, this embodiment proposes a two-layer densification framework with adaptive Gaussian pruning, introducing learnable RGB mask parameters , and generating a binary mask through a straight-through estimator .
[0060] Among them, the generation formula of the binary mask is: ; In the formula, represents the binary mask, represents the gradient truncation operator, represents the indicator function, which is 1 if the condition holds and 0 otherwise; represents the sigmoid function, represents the RGB mask parameter, represents the threshold of the mask.
[0061] 24) For the pruned Gaussian bodies, introduce an opacity filtering mechanism to compare the opacity value of each Gaussian body with a preset threshold to filter out Gaussian bodies whose opacity values do not exceed the preset threshold, obtaining the final 3D Gaussian bodies.
[0062] Specifically, to address the adverse effects of low-opacity Gaussians on the scene construction accuracy, an opacity-based filtering mechanism is introduced here. The filtering mechanism compares the opacity value of each Gaussian with a dynamically adjustable threshold parameter. This layer serves as a preprocessing stage to ensure that only Gaussians that exceed the predefined opacity threshold are merged into the rendering pipeline, and the formula is expressed as: ; In the formula, represents the set where the opacity of the Gaussian volume exceeds a predefined threshold, represents the opacity of the -th Gaussian volume, represents the
[0063] In this embodiment, this threshold-driven method effectively eliminates the Gaussian basis elements that contribute insufficiently to the scene geometry, while maintaining real-time performance through parallelizable calculations. It is worth noting that the separation of opacity evaluation and mask derivation is consistent with the best practices in resource-constrained rendering systems.
[0064] Through the above processing, the monocular RGB image is represented as different 3D Gaussian volumes.
[0065] Step 3: Based on the monocular RGB image and the depth image, optimize the camera pose in the camera tracking module to obtain the camera pose information.
[0066] First, initialize the camera pose.
[0067] Specifically, in the camera tracking module, this embodiment uses the Jacobian matrix of the camera pose to more precisely analyze and control the relationship between the camera pose and the rasterization process. The camera pose is represented as a rigid transformation matrix , and is parameterized using Lie algebra elements. The exponential map connects the two: ; where represents the skew-symmetric matrix representation of , and
[0068] The initial pose is projected onto the orthogonal space using quaternion translation decomposition: ; where represents the unified scaling factor for orthogonal consistency, represents initialization through quaternion normalization, represents the external camera parameters.
[0069] Then, based on the monocular RGB image and the depth image, iteratively optimize the camera pose using photometric residuals and depth residuals to obtain the camera pose information.
[0070] Specifically, when only monocular data is available, the photometric residual between the rendered image and the original image is defined as: ; ; In the formula, represents the photometric residual, represents the opacity of the th Gaussian body, represents the pixel domain, represents the hyperparameter, represents the value of the photometric change of the image filtered by the mask, represents the L1 loss function, represents the SSIM loss function, represents element-wise multiplication, represents the pixel region of the RGB mask, represents the rendered image, represents the pixel region, represents the pixel region corresponding pose matrix, represents the original image.
[0071] Regarding the analytical Jacobian is derived by the chain rule of "SE(3)", that is: ; where SE(3) is the group composed of all possible rigid body transformation matrices, utilizes the pose matrix to calculate the derivative of the Lie algebra parameter equal to the adjoint matrix of its inverse transformation, realizing efficient calculation of the gradient in Lie group optimization, that is: ; In the formula, is the adjoint matrix, which maintains the Lie group structure during the differentiation process.
[0072] To avoid the overhead of automatic differentiation, the derivative is directly calculated based on CUDA rasterization: ; where, represents the downward gradient, represents the rasterization function, represents the gradient of the rendered image, and the gradient of the rendered image is calculated in parallel. The pose is updated as follows: ; where, and respectively represent thek The Lie algebra parameters in the +1-th and the k -th iterations are empirically set to converge within 50 iterations, and is the learning rate.
[0073] The final pose update follows the manifold geometry of "SE(3)": ; .
[0074] where, and represent the pose matrices in the +1-th and the k -th iterations, and k represents the skew-symmetric matrix representation of . Furthermore, in this embodiment, the depth information of a single RGB image is obtained through the MViT module, that is, there is available depth information, then the camera position can be further optimized by considering photometric and depth residuals. In this embodiment, the depth residual is defined as:
[0075] ; ; ; wherein, represents the depth residual, represents the indicator function, represents the depth of the rendered image, represents the depth threshold, represents the pixel domain, represents the depth change value of the image filtered by the mask, represents the pixel region of the depth mask, and represents the original image depth.
[0076] Then the total tracking loss is defined as follows: ; wherein, represents the total tracking loss, when the depth information is not available, is 1, otherwise, .
[0077] Step 4: Based on the camera pose information and the 3D Gaussian volume, key frame selection is performed in the key frame management module to obtain key frame information.
[0078] In the key frame management module, the strategic selection and management of key frames are crucial for improving computational efficiency while maintaining map integrity. The optimization of Gaussian representation usually requires processing large-scale datasets, and the computational complexity exhibits a linear scaling characteristic with respect to the number of frames.
[0079] In this embodiment, when inserting a key frame, the current frame and the last key frame must have a visibility overlap rate lower than a threshold. For key frame culling, each candidate key frame is compared with the current key frame. Then the decision function for key frame management is formalized as: ; ; where, represents the key frame to be inserted, represents the indicator function, which is 1 if the condition holds and 0 otherwise, represents the current key frame, represents the last key frame, represents the selected key frame, and both represent the threshold function; represents the key frame to be deleted.
[0080] Furthermore, since this embodiment obtains the depth information of a single RGB image through the MViT module, that is, there is available depth information, then combining the depth information, the following parallel insertion criterion is implemented according to the relative displacement: ; where, represents the relative displacement, represents the median depth information of the candidate key frame, represents the corresponding threshold function.
[0081] It should be noted that when the depth information is missing, the threshold determination method relies on a fixed threshold criterion to judge the importance of frames without considering the complexity and dynamics of the scene. This will lead to the misidentification of some key frames and the determination of linking nearby visible Gaussians, resulting in a decrease in accuracy.
[0082] To address this challenge, this embodiment introduces adaptive key frame interval adjustment after key frame management, dynamically adjusts the key frame interval according to the scene complexity, and quantifies it through the tracking loss , that is: ; where, represents the adjusted key frame interval, Represents the default key-frame interval, Represents the total tracking loss, Represents the loss threshold.
[0083] Step 5: Based on the key-frame information and the 3D Gaussian volume, generate a new view rendering through isotropic regularization in the mapping module, thereby realizing the reconstruction of a dense visual field.
[0084] 51) Optimize the primitive position of the 3D Gaussian volume based on the key-frame information, and add an isotropic regularization term during the optimization process to obtain an optimized 3D Gaussian volume.
[0085] Specifically, in the mapping module, optimize the placement of the Gaussian primitives according to the estimated online camera pose set for the current frame. Areas with insufficient field-of-view coverage are vulnerable to unconstrained geometric deformations and artifacts because the previous Gaussian distributions lack observational constraints. To mitigate the excessive elongation of the Gaussian primitives, an isotropic regularization term is incorporated into the optimization framework, thereby strengthening geometric consistency and stabilizing the reconstruction process, that is: ; In the formula, Represents the loss function of isotropic regularization, Represents the Gaussian scale, Represents its average value.
[0086] It should be noted that isotropic regularization effectively restricts the Gaussian expansion to maintain a geometrically stable scale while minimizing reconstruction artifacts. In contrast, spherical Gaussian distributions tend to exhibit uncontrolled spatial propagation in unobserved regions, which may introduce representational inaccuracies during the scene reconstruction process.
[0087] 52) Based on the constructed total mapping loss, perform image construction on the optimized 3D Gaussian volume representation, and use the gradient magnitude and visibility statistics for densification or pruning during the image construction process to generate a new view rendering, thereby realizing the reconstruction of a dense visual field.
[0088] Specifically, the total mapping loss is defined as follows: ; In the formula, and are respectively and weights. If the depth information is not available, then is 1, otherwise, .
[0089] Furthermore, in this embodiment, a densification process is performed during the construction of the graph. Gaussians are densified or pruned based on gradient magnitude and visibility statistics. When the maximum gradient magnitude of all Gaussians exceeds a threshold, the densification condition is triggered: ; wherein, is the average position relative to the th Gaussian, represents the total loss gradient, is a predefined gradient threshold.
[0090] At this time, the Gaussian volume needs to be densified.
[0091] When the number of key frames in which the Gaussian volume is observed is less than the threshold number of key frames, the Gaussian volume is pruned, that is: ; wherein, represents the number of times the i th Gaussian volume is observed by other Gaussian volumes, represents the number of visible key frames of the product at this time.
[0092] At this time, the Gaussian volume needs to be pruned.
[0093] Through the above processing of the mapping module, the current reconstructed image is obtained.
[0094] It can be understood that under the monocular ViMGS-SLAM framework proposed by the present invention, based on the current reconstructed image, camera pose optimization can be assisted, and by updating the Gaussian parameters of the three-dimensional scene and key frame selection, real-time continuous scene reconstruction can be achieved.
[0095] A dense visual field reconstruction method based on MViT provided by the present invention constructs a monocular ViMGS-SLAM framework, which includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module; for the input monocular RGB image, first, cross-scale attention mechanism is used in the MViT module to extract photometric features in the cross-spatial frequency domain, so as to enhance the geometric consistency of depth prediction and achieve accurate depth estimation; then, combined with the geometric modeling ability of the 3D Gaussian representation module, the camera pose optimization of the camera tracking module, the key frame selection of the key frame management module, and the parameter update and new view rendering of the mapping module, dense visual field reconstruction in a dynamic scene is achieved, and the problems of scale drift accumulation in the long-term operation of the monocular system and real-time high-precision 3D reconstruction of monocular RGB input are solved; at the same time, this framework can also achieve real-time high-precision reconstruction on RGB and RGB-D data streams.
[0096] Based on the same inventive concept, the second aspect of the present invention further provides a dense visual field reconstruction system based on MViT. This system is equipped with a monocular ViMGS-SLAM framework, and the monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module. Among them, The MViT module is used to estimate the depth of the monocular RGB image in the image to be reconstructed through a cross-scale attention mechanism, and obtain a depth image. The 3D Gaussian representation module is used to perform reparameterization processing on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume. The camera tracking module is used to optimize the camera pose based on the monocular RGB image and the depth image, and obtain camera pose information. The key frame management module is used to select key frames based on the camera pose information and the 3D Gaussian volume, and obtain key frame information. The mapping module is used to generate new view rendering through isotropic regularization based on the key frame information and the 3D Gaussian volume, so as to realize dense visual field reconstruction.
[0097] It should be noted that for the system embodiment, it can implement the above-mentioned dense visual field reconstruction method based on MViT. For the relevant parts, refer to the partial description of the method embodiment. Therefore, this system can also achieve the same or similar beneficial effects as the method embodiment.
[0098] Next, the effectiveness of the method proposed by the present invention is verified and illustrated through experiments.
[0099] 1. Experimental Setup In this experiment, the algorithm proposed by the present invention is evaluated on the Replica dataset and the TUM dataset widely recognized in the SLAM community, including monocular sequences and depth-enhanced RGB-D sequences. The TUM RGB dataset faces inherent challenges in sensor acquisition. Its characteristics are obvious motion blur in color images, and due to the limitations of the Kinect v1 sensor, depth measurements are very sparse. These characteristics severely test the ability of the algorithm to preserve details under real-world noise conditions. In contrast, the synthetic environment provided by the Replica dataset has ground truth poses and dense depth maps with millimeter-level accuracy, making it an ideal benchmark for evaluating the performance upper limit in a controlled scenario. This experimental scheme follows the established evaluation metrics, using eight Replica sequences (office 0 - 4 and room 0 - 2) and three representative TUM RGB-D sequences (fr1 desk, fr2 xyz, fr3 office) to ensure comprehensive verification under synthetic and real-world conditions.
[0100] In terms of RGB rendering quality assessment, three widely adopted perceptual metrics were used this time: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). In camera trajectory evaluation, the Root Mean Square Error (RMSE) of the Trajectory Absolute Error (ATE) metric was used to quantify the positioning accuracy. To evaluate the performance of RGB only, ORB-SLAM2, MonoGS, and MotionGS were benchmarked on the TUM RGB dataset this time. To comprehensively analyze RGB-D scenes, this experiment was compared with NeRF-based SLAM implementations (Vox-Fusion, NICE-SLAM, Point-SLAM) and 3DGS-based systems (SplaTAM, GS-SLAM, MonoGS). All comparisons used the benchmark results published in the original papers and the reproduction experiments conducted using open-source implementations under the same hardware configuration.
[0101] 2. Experimental Results and Discussion 2.1 Novel View Synthesis Tables 1 and 2 respectively show the comparison results of the rendering performance of different SLAM methods on the Replica dataset.
[0102] Table 1 Rendering Performance of NeRF-based SLAM Methods on the Replica Dataset
[0103] Table 2 Rendering Performance of DGS-based SLAM Methods on the Replica Dataset
[0104] Among them, the symbol "↑" after the metric indicates that the higher the metric, the better the effect, and "↓" indicates that the lower the metric, the worse the effect.
[0105] As can be seen from Table 1 and Table 2, the proposed ViMGS-SLAM framework in the present invention has created new state-of-the-art performance on the Replica dataset, achieving breakthroughs in both reconstruction fidelity and computational efficiency. The results show that ViMGS-SLAM has comprehensive advantages in all metrics: 1) Radiometric accuracy. The proposed method reached 39.60 dB (Δ +1.69 dB compared to the second-best method). This result is close to the lossless reconstruction threshold (≥40 dB), indicating a very small deviation from the ground truth measurement. 2) Structural integrity. The structural similarity index is 0.976. Compared with the implementations of Point-SLAM and GS-SLAM, the proposed solution in the present invention shows a certain improvement (Δ +0.001). Although ViMGS-SLAM does not dominate in all scenarios, it always ranks among the top in terms of the overall SSIM performance. 3) Perceptual fidelity. The learned perceptual image patch similarity metric reached 0.042, with a 66.7% reduction in perceptual difference compared to the MonoGS method. At the same time, ViMGS-SLAM achieved real-time rendering at 1064.17 FPS, two orders of magnitude higher than NeRF-based SLAM and 38.38% higher than 3DGS-based SLAM. The proposed architecture achieved real-time differentiable rasterization by supporting the geometric processing of 3DGS, adopting a depth-aware sorting algorithm and an adaptive multi-resolution culling mechanism.
[0106] Figure 4 Scene reconstruction performance of the proposed ViMGS-SLAM framework in the present invention and the existing MonoGS framework on the Office3 and Room2 sequences of the Replica dataset.
[0107] It can be seen that using RGB-D sensor data, the ViMGS-SLAM framework of the present invention achieved better geometric reconstruction accuracy than MonoGS, especially in retaining high-frequency geometric details such as edge contours and surface irregularities. In terms of hierarchical detail retention, ViMGS-SLAM retained sub-centimeter-scale structural features (such as the clock device and the sofa gap in Office3) during the local reconstruction process, while improving the structural integrity of occluded areas through the MViT module. In terms of scene completion, the framework ensured semantically consistent hole filling (e.g., reconstructing the unobserved area under the cabinet in Room2) through structure-perception consistency and reduced texture artifacts at material boundaries (e.g., reconstructing the display panel in Office3) through gradient-aware fusion, significantly outperforming the MonoGS method.
[0108] 2.2 Ablation experiment analysis Using three TUM RGB sequences (fr1, fr2, fr3), with RMSE ATE in centimeters as the main metric, the contributions of each loss component were systematically evaluated through controlled ablation experiments. The ViMGS-SLAM framework uses photometric loss and depth loss for camera pose estimation during tracking, and combines photometric loss, depth loss, and anisotropic regularization for scene reconstruction during mapping. To study their respective roles, a systematic ablation study was conducted on these three parts, analyzing the effects of color error, geometric residuals, and isotropic regularization on the SLAM system, and the results are shown in Table 3.
[0109] Table 3 Ablation experiment results of ViMGS-SLAM in the TUM RGB scenario
[0110] The following conclusions can be drawn from the results in Table 3: 1) Subtract photometric loss. After eliminating the photometric loss, the system mainly relies on depth loss and anisotropic regularization for tracking and mapping. The experimental results show that compared with the complete system, the RMSE ATE increased by 2.41 times, highlighting the indispensable role of photometric constraints in maintaining positioning accuracy and system robustness. This degradation occurs because photometric loss provides crucial visual feedback through pixel intensity alignment between the current frame and the reference frame, which is particularly important for handling lighting changes and textureless regions.
[0111] 2) Eliminate depth loss. After eliminating the depth loss, the RMSE ATE increased by 1.8 times, highlighting the importance of depth loss for geometric accuracy. Depth loss provides a metric scale constraint by leveraging scene geometry, thereby alleviating the scale ambiguity inherent in monocular SLAM. However, the system still retains functional tracking and mapping capabilities, indicating that in static environments, photometric and regularization conditions can partially compensate for depth loss.
[0112] 3) Subtract anisotropic regularization. Eliminating anisotropic regularization results in a performance degradation comparable to that of the depth loss ablation, highlighting its importance in noise suppression and error propagation control. Anisotropic regularization selectively smooths the scene geometry while retaining edge features through direction diffusion constraints, which is particularly important for maintaining consistency in dynamic environments. Figure 1 consistency is particularly important.
[0113] 2.3 Runtime analysis Table 4 shows the runtime of VIMGS-SLAM on the TUM RGB-D_fr2 / desk dataset. The total execution time of the system, the FPS calculated by dividing the total number of processed frames by the total time, and the average number of mapping iterations for each newly added keyframe are reported.
[0114] Table 4 Experimental results of the running time of VIMGS-SLAM on TUM RGB-D_fr2 / desk
[0115] It can be seen that in the monocular RGB mode, the system achieves an end-to-end frame processing speed of 2.7 FPS, demonstrating its ability to effectively utilize photometric features (such as color gradients and texture patterns) for pose estimation and mapping. When integrating depth data from an RGB-D sensor, the processing speed drops to 2.3 FPS, indicating an increase in the computational overhead associated with depth map registration and 3D point cloud optimization. First, when depth information is unavailable, the system maintains operational robustness in dynamic environments through an adaptive feature selection mechanism. Second, the observed trade-off between frame rate and reconstruction accuracy quantifies the inherent computational cost-benefit relationship of depth-enhanced SLAM architectures. Runtime analysis confirms the balanced design philosophy of VIMGS-SLAM, which prioritizes geometric consistency and real-time operation constraints.
[0116] In summary, the monocular ViMGS-SLAM framework proposed in the present invention can achieve explicit geometric scene reconstruction and realistic rendering through 3DGS. This architecture introduces the MViT structure, aiming to collaboratively capture fine-grained local details and global context information. This hierarchical feature extraction mechanism combines dynamic feature fusion and boundary-aware optimization, effectively addressing the limitations of monocular depth estimation through adaptive scale perception. The ViMGS-SLAM framework consists of five tightly integrated components: the MViT module, the 3D Gaussian representation module, the camera tracking module, the keyframe management module, and the mapping module. ViMGS-SLAM supports taking only monocular RGB data as input, as well as taking RGB and RGB-D data as input simultaneously, and has made significant progress in camera pose estimation, scene reconstruction, and rendering, providing valuable insights and a solid foundation for the future development of more complex and efficient 3DGS-based SLAM systems.
[0117] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A dense visual field reconstruction method based on MViT, characterized in that The method is implemented based on a monocular ViMGS-SLAM framework, and the monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module; the method includes: Input the image to be reconstructed into the monocular ViMGS-SLAM framework. When the image to be reconstructed only includes a monocular RGB image, in the MViT module, perform depth estimation on the monocular RGB image through a cross-scale attention mechanism to obtain a depth image; when the image to be reconstructed also includes an RGB-D image, use the RGB-D image as the depth image; In the 3D Gaussian representation module, perform reparameterization processing on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume; Based on the monocular RGB image and the depth image, optimize the camera pose in the camera tracking module to obtain camera pose information; Based on the camera pose information and the 3D Gaussian volume, perform key frame selection in the key frame management module to obtain key frame information; Based on the key frame information and the 3D Gaussian volume, generate a new view rendering through isotropic regularization in the mapping module, thereby realizing the reconstruction of a dense visual field.
2. The dense visual field reconstruction method based on MViT according to claim 1, wherein, In the MViT module, perform depth estimation on the monocular RGB image through a cross-scale attention mechanism to obtain a depth image, specifically including: After symmetric padding of the monocular RGB image, divide it into multi-scale image patches of different resolutions, and perform linear projection on the image patches of each scale to generate multi-scale feature vectors; Adopt a hierarchical Transformer and a cross-scale attention mechanism to obtain the feature vectors after fusing all multi-scale feature vectors, and use a local branch to extract features to obtain local multi-scale features; at the same time, use a global branch to extract features from the feature vectors of the lowest resolution scale in the multi-scale feature vectors to obtain global features; Based on a hierarchical transformer with cross-scale attention, perform local-global feature interaction and adaptive feature fusion on the local multi-scale features and the global features, and introduce a multi-layer perceptron for progressive upsampling processing during the fusion process to restore the corresponding image information, and then use a DPT decoder to reconstruct the inverse depth map; Generate a metric depth map representing absolute depth information and a refined edge depth map with the same resolution as the RGB image according to the inverse depth map as the depth image.
3. A dense visual field reconstruction method based on MViT according to claim 1, characterized in that In the 3D Gaussian representation module, perform reparameterization processing on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume, specifically including: Initialize the three-dimensional scene based on the monocular RGB image to represent the three-dimensional scene as a coupled representation of a group of anisotropic 3D Gaussian volumes, and decompose the three-dimensional covariance matrix of each 3D Gaussian volume into rotation and scaling operations; Project each 3D Gaussian volume onto the imaging plane, sort them based on the depth information according to the depth image, and use end-to-end optimization for scene rendering to obtain the rendered Gaussian volumes; Generate a binary mask based on learnable RGB mask parameters to adaptively prune the rendered Gaussian volume, obtaining a pruned Gaussian volume; For the pruned Gaussian volume, introduce an opacity filtering mechanism to compare the opacity value of each Gaussian volume with a preset threshold to filter out Gaussian volumes whose opacity values do not exceed the preset threshold, obtaining the final 3D Gaussian volume.
4. A dense visual field reconstruction method based on MViT according to claim 3, characterized in that, Project each 3D Gaussian volume onto the imaging plane. After sorting based on the depth information in the depth image, use end-to-end optimization for scene rendering to obtain the rendered Gaussian volume, specifically including: Project each 3D Gaussian volume onto the imaging plane through the camera intrinsic parameters and camera extrinsic parameters to obtain the projected 2D Gaussian volume; where the mean and covariance of the 2D Gaussian volume are expressed as: ; ; In the formula, represents the mean of the -th 2D Gaussian volume, represents the perspective projection operation, represents the camera intrinsic parameters, represents the rotation matrix, represents the camera extrinsic parameters, represents the mean of the -th 3D Gaussian volume, represents the -th covariance of the 2D Gaussian volume, represents the -th covariance of the 3D Gaussian volume, represents the Jacobian matrix in the projection operation, and the superscript represents the matrix transpose; Sort the 2D Gaussian volumes based on the depth information in the depth image; Calculate the color of each pixel by blending multiple ordered Gaussians along the ray from the front end to the back end to achieve scene rendering, obtaining the rendered Gaussian volume; where the calculation formula for the pixel color is: ; In the formula, represents the color of a pixel ; represents the number of Gaussian bodies, the -th opacity of the 2D Gaussian body, represents the 2D Gaussian parameters calculated at pixel ; the -th mean value of the 2D Gaussian body, the -th covariance of the 2D Gaussian body, the -th color of the 2D Gaussian body, the -th opacity of the 2D Gaussian body.
5. A dense visual field reconstruction method based on MViT according to claim 1, characterized in that, Based on the monocular RGB image and the depth image, perform camera pose optimization in the camera tracking module to obtain camera pose information, including: Initialize the camera pose; Based on the monocular RGB image and the depth image, iteratively optimize the camera pose using photometric residuals and depth residuals to obtain camera pose information; Among them, the photometric residual is expressed as: ; ; In the formula, represents the photometric residual, represents the opacity of the th Gaussian body, represents the pixel domain, represents the hyperparameter, represents the photometric change value of the image filtered by the mask, represents the L1 loss function, represents the SSIM loss function, represents element-wise multiplication, represents the pixel area of the RGB mask, represents the rendered image, represents the pixel area, represents the pixel area corresponding pose matrix, represents the original image; The depth residual is expressed as: ; ; In the formula, represents the depth residual, represents the indicator function, represents the depth of the rendered image, represents the depth threshold, represents the depth change value of the image filtered by the mask, represents the pixel area of the depth mask, represents the original image depth.
6. A dense visual field reconstruction method based on MViT according to claim 1, characterized in that When performing key frame selection in the key frame management module based on the camera pose information and the 3D Gaussian volume, insert key frames according to the following decision function: ; Wherein, represents the key frame to be inserted, represents the indicator function, represents the current key frame, represents the last key frame, represents the threshold function; At the same time, combined with the depth information, implement the following parallel insertion criterion according to the relative displacement: ; In the formula, represents the relative displacement, represents the median depth information of the candidate key frame, represents the corresponding threshold function.
7. A method for reconstructing a dense visual field based on MViT according to claim 6, characterized in that, During the process of performing key frame selection in the key frame management module, it also includes: Adaptive adjustment of the key frame interval; where the adjustment formula for the key frame interval is: ; In the formula, represents the adjusted key frame interval, represents the default key frame interval, represents the tracking loss, represents the loss threshold.
8. A dense visual field reconstruction method based on MViT according to claim 1, characterized in that, Based on the key frame information and the 3D Gaussian volume, generate a new view rendering through isotropic regularization in the mapping module, thereby realizing real-time dense visual field reconstruction, including: Optimize the primitive position of the 3D Gaussian volume based on the key frame information, and add an isotropic regularization term during the optimization process to obtain the optimized 3D Gaussian volume; According to the constructed total mapping loss, perform image construction on the optimized 3D Gaussian volume representation, and use gradient magnitude and visibility statistics for densification or pruning during the image construction process to generate a new view rendering, thereby realizing dense visual field reconstruction.
9. A method for reconstructing a dense visual field based on MViT according to claim 8, characterized in that, Using gradient magnitude and visibility statistics for densification or pruning during the image construction process includes: When the maximum gradient magnitude of all Gaussian volumes exceeds the preset gradient threshold, perform densification processing on the Gaussian volumes; When the number of key frames in which a Gaussian volume is observed is less than the threshold number of key frames, perform pruning processing on the Gaussian volume.
10. A dense visual field reconstruction system based on MViT for implementing the method according to any one of claims 1-9, characterized in that, The system is equipped with a monocular ViMGS-SLAM framework, and the monocular ViMGS-SLAM framework includes an MViT module, a 3D Gaussian representation module, a camera tracking module, a key frame management module, and a mapping module; where, The MViT module is used to perform depth estimation on the monocular RGB image in the image to be reconstructed through a cross-scale attention mechanism to obtain a depth image; The 3D Gaussian representation module is used to perform reparameterization processing on the monocular RGB image based on the depth image to obtain a 3D Gaussian volume; The camera tracking module is used to optimize the camera pose based on the monocular RGB image and the depth image to obtain camera pose information; The key frame management module is used to select key frames based on the camera pose information and the 3D Gaussian volume to obtain key frame information; The mapping module is used to generate new view rendering through isotropic regularization based on the key frame information and the 3D Gaussian volume, thereby realizing dense visual field reconstruction.
Citation Information
Patent Citations
Visual and inertial integrated positioning system based on point-line feature rapid fusion
CN117804438A
Dense RGB-D SLAM method based on multistage 3D Gaussian
CN118781189A
Laser enhanced vision three-dimensional reconstruction method and system based on Gaussian splashing
CN119180908A
Dense visual scene reconstruction method and system
CN119478277A
High-precision three-dimensional reconstruction method based on sparse depth
CN119784947A
Cited By
Three-dimensional scene reconstruction method and system based on monocular depth estimation
CN121482285A
Three-dimensional scene reconstruction method and system based on monocular depth estimation
CN121482285B
Dynamic scene-oriented gradient sensing self-supervising monocular depth estimation method and device
CN121810754A
Gradient-aware self-supervised monocular depth estimation method and device for dynamic scene
CN121810754B