Semantic dense vision SLAM method and system based on 3D Gaussian representation
By adopting a semantic intensive visual SLAM method based on 3D Gaussian representation in the SLAM system, combining multiple technologies to optimize Gaussian point parameters and camera poses, the problem of insufficient positioning accuracy and robustness of traditional SLAM systems in complex environments is solved, and more efficient semantic information utilization and dynamic object processing is achieved, which significantly improves the system's map construction accuracy and robustness.
Patent Information
- Application Number
- CN202510372816.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional SLAM systems have insufficient positioning accuracy and robustness in complex environments, difficult to effectively utilize semantic information, and difficult to deal with dynamic object interference, and there are also problems with map representation efficiency and accuracy.
The semantic intensive visual SLAM method based on 3D Gaussian representation is adopted, combining Gaussian voting point cloud technology, dynamic object processing, semantic feature fusion, feature-level loss optimization, semantic perception bundled adjustment, virtual camera view pruning, multi-channel optimization, loopback detection and differentiable rendering technology to optimize Gaussian point parameters and camera poses to improve the system's graph construction accuracy and robustness.
It significantly improves the accuracy and robustness of the system's map construction, allowing it to show better performance in complex environments, enable a more comprehensive understanding of the scene, reduce dynamic object interference, and improve the accuracy and consistency of the map.
Smart Images

Figure CN120219462A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of embodied intelligence, and relates to the technical fields of image detection, robot mapping and positioning. Specifically, it relates to a semantic dense visual SLAM method and system based on 3D Gaussian representation. Background Art
[0002] In the field of embodied intelligence, agents face unprecedented challenges. They need to achieve autonomous perception, decision-making and interaction in complex, ever-changing and uncertain real-world environments. The realization of this goal places extremely high requirements on the environmental cognition ability of agents. Agents not only need to be able to construct a model of the surrounding environment in real time and accurately to capture the spatial structure and layout of the environment, but also need to deeply understand the semantic information in the environment, such as the category, attributes of objects and the relationships between them, so as to be able to make reasonable action decisions based on this comprehensive information.
[0003] However, traditional visual SLAM (Simultaneous Localization and Mapping) systems have many defects when constructing environmental maps, which seriously limit the application and development of agents in the field of embodied intelligence.
[0004] First of all, the positioning accuracy and robustness of traditional SLAM systems in complex environments are significantly insufficient. The core idea of the feature point-based SLAM method is to describe the environment by extracting sparse feature points in the environment and perform positioning and map construction. However, this method only relies on limited feature point information, resulting in the loss of a large amount of environmental detail information and making it difficult to meet the requirements of embodied intelligence for refined environmental perception. In complex and ever-changing environments, the stability and reliability of feature points will also be severely affected, thereby affecting the positioning accuracy and map accuracy. While the direct method SLAM can construct a dense map using image pixel information, providing richer environmental details, its computational complexity is huge, it has high requirements for hardware resources, and it is easily affected by environmental factors such as light changes, and its stability is poor.
[0005] Secondly, the low utilization efficiency of semantic information is also a major problem faced by traditional SLAM systems. At the semantic understanding level, most existing semantic SLAM systems adopt the method of simply combining the semantic segmentation results with traditional SLAM. This simple fusion method does not fully explore the internal connection between semantic information and geometric information, resulting in the role of semantic information not being fully exerted in the process of map construction and positioning. The inefficient utilization of semantic information not only wastes valuable semantic resources, but also limits the agent's in-depth understanding of the environment and intelligent interaction ability.
[0006] In addition, the interference of dynamic objects is also a problem that traditional SLAM systems are difficult to overcome. In the real world, there are often a large number of dynamic objects in the environment, such as walking pedestrians, moving vehicles, etc. The presence of these dynamic objects will cause serious interference to the positioning and map construction of the SLAM system, resulting in increased positioning errors, map distortion and other problems.
[0007] Finally, the efficiency and accuracy of map representation are also urgent problems to be solved by traditional SLAM systems. For the representation of 3D scenes, traditional methods such as voxel representation and point cloud representation have their own drawbacks. Although voxel representation can accurately describe the three-dimensional structure and details of the scene, it consumes huge memory and has low computational efficiency, making it difficult to meet real-time requirements. Point cloud representation lacks effective expression of the smoothness and continuity of the scene surface, making it difficult to use for subsequent rendering and analysis tasks, limiting the intelligent agent's comprehensive understanding and utilization of the environment. Therefore, how to efficiently and accurately represent 3D scenes has become a key issue that needs to be solved in the field of embodied intelligence. Summary of the invention
[0008] In view of this, the purpose of the present invention is to provide a semantic dense visual SLAM method and system based on 3D Gaussian representation, to overcome the problems of insufficient positioning accuracy and robustness of traditional SLAM systems in complex environments, and to construct a semantic dense visual SLAM system based on 3D Gaussian representation that integrates Gaussian voting point cloud technology, dynamic object processing, semantic feature fusion, feature-level loss optimization, semantic-aware bundle adjustment, virtual camera view pruning, multi-channel optimization, loop detection and differentiable rendering technology, which significantly improves the mapping accuracy and robustness of the system and enables it to exhibit better performance in complex environments.
[0009] In order to achieve the above object, the present invention provides the following technical solutions:
[0010] Solution 1: A semantic dense visual SLAM method based on 3D Gaussian representation, which obtains the image data of the scene through sensors, constructs a scene model represented by 3D Gaussian, fuses semantic features with appearance features to generate high-dimensional semantic features, optimizes the parameters of Gaussian points based on appearance, geometry and semantic constraints, updates the scene model and optimizes the camera pose according to the optimized parameters, generates rendered images and tracks the location through differentiable rendering technology, and uses image features to detect loops and optimize global map accuracy.
[0011] Furthermore, image data of the scene is obtained through the sensor, specifically including: collecting RGB images and depth images in real time through the RGB-D sensor, extracting initial semantic features and appearance features from the RGB images, generating a preliminary geometric structure of the scene according to the depth image, aligning the preliminary geometric structure with the initial semantic features and appearance features, and generating input data for constructing a 3D Gaussian representation.
[0012] Furthermore, a scene model of 3D Gaussian representation is constructed, specifically including: extracting position and depth information from the initial frame image data to generate multiple Gaussian points, where each Gaussian point is defined by position (mean), covariance matrix, color, opacity, and semantic features; performing semantic segmentation on the initial frame according to a pre-trained network to obtain the semantic segmentation result, assigning semantic labels to each Gaussian point through the Gaussian voting point cloud technology, and integrating the semantic information of the Gaussian points according to the semantic labels to generate an initial 3D Gaussian representation.
[0013] Furthermore, semantic features and appearance features are fused to generate high-dimensional semantic features, specifically including: obtaining RGB images and depth images through an RGB-D sensor, extracting semantic features from the RGB images using a pre-trained semantic segmentation network, and at the same time, extracting appearance features from the RGB images using a feature extraction method or a deep learning-based method; then, fusing the semantic features and appearance features to generate high-dimensional semantic features consistent with the scene space.
[0014] Furthermore, the parameters of Gaussian points are optimized by combining appearance, geometric, and semantic constraints, specifically including: optimizing appearance parameters such as the color of Gaussian points by minimizing the difference between the rendered image and the real image; optimizing the position and shape of Gaussian points through geometric constraints; optimizing the semantic features of Gaussian points using semantic constraints; integrating the optimization results of color parameters, position, shape, and semantic features using multi-channel optimization technology and feature-level loss optimization technology to generate optimized Gaussian point parameters.
[0015] Furthermore, the scene model is updated according to the optimized parameters and the camera pose is optimized, specifically including: guiding key frame selection and optimization using semantic-aware bundle adjustment technology, specifically selecting key frames semantically related to the current frame according to the semantic labels of Gaussian points, and jointly optimizing Gaussian point parameters and camera pose through multi-view semantic constraints; at the same time, detecting and removing feature points of dynamic objects using semantic information and geometric constraints, verifying the multi-view consistency of Gaussian points from different angles using virtual views, and updating the scene model and optimizing the camera pose according to the verification results.
[0016] Furthermore, a rendered image is generated and camera positioning is tracked through differentiable rendering technology, specifically including: generating a rendered image according to the current scene model using differentiable rendering technology, predicting the semantic mask and depth image of the real image using a pre-trained segmentation network, adjusting the Gaussian point parameters according to the differences between the rendered image and the semantic mask and depth image of the real image, generating a rendered image highly consistent with the real image, and tracking camera positioning according to the adjusted rendered image.
[0017] Furthermore, loop closure is detected using image features and the global map accuracy is optimized, specifically including: quantifying image features using the BoW method to generate a visual vocabulary, quantifying feature points in the image into visual words, thereby forming a visual word histogram; calculating the similarity between images based on the visual vocabulary word histogram, screening loop closure candidate frames, using the ICP algorithm to align and verify the 3D Gaussian point sets of the loop closure candidate frames, and then adjusting the pose constraints of the loop closure frames through the pose graph optimization method to update the accuracy of the global map.
[0018] Solution 2: A semantic dense visual SLAM system based on 3D Gaussian representation, specifically including:
[0019] Initialization module: Construct a scene model in 3D Gaussian representation, specifically including: extracting position and depth information from the initial frame image data to generate multiple Gaussian points, each Gaussian point is defined by position (mean), covariance matrix, color, opacity, and semantic features; performing semantic segmentation on the initial frame according to the pre-trained network to obtain the semantic segmentation result, assigning semantic labels to each Gaussian point through the Gaussian voting point cloud technology, and integrating the semantic information of the Gaussian points according to the semantic labels to generate the initial 3D Gaussian representation;
[0020] Data acquisition and preprocessing module: Obtain the image data of the scene through sensors, specifically including: real-time collecting RGB images and depth images through an RGB-D sensor, extracting initial semantic features and appearance features for the RGB images, generating a preliminary geometric structure of the scene according to the depth images, and aligning the preliminary geometric structure with the initial semantic features and appearance features to generate the input data for constructing the 3D Gaussian representation;
[0021] Core optimization and mapping module: Optimize the parameters of Gaussian points by combining appearance, geometric, and semantic constraints, specifically including: optimizing appearance parameters such as the color of Gaussian points by minimizing the difference between the rendered image and the real image; optimizing the position and shape of Gaussian points through geometric constraints; adopting semantic constraints to optimize the semantic features of Gaussian points; using multi-channel optimization technology and feature-level loss optimization technology to integrate the optimization results of color parameters, position, shape, and semantic features to generate optimized Gaussian point parameters;
[0022] Then update the scene model and determine the camera pose according to the optimized parameters, specifically including: guiding key frame selection and optimization using semantic-aware bundle adjustment technology, specifically selecting key frames semantically related to the current frame according to the semantic labels of Gaussian points, and jointly optimizing Gaussian point parameters and camera pose through multi-view semantic constraints; at the same time, detecting and removing the feature points of dynamic objects using semantic information and geometric constraints, verifying the multi-view consistency of Gaussian points from different angles using virtual views, and updating the scene model and optimizing the camera pose according to the verification results;
[0023] Rendering and Tracking Module: Generates rendered images and tracks the position through differentiable rendering technology, specifically including: Generating rendered images according to the current scene model using differentiable rendering technology, predicting the semantic mask and depth image of the real image using a pre-trained segmentation network, adjusting the Gaussian point parameters based on the differences between the rendered image and the semantic mask and depth image of the real image, generating a rendered image highly consistent with the real image, and tracking the camera position based on the adjusted rendered image;
[0024] Loop Detection Module: Detects loops using image features and optimizes the global map accuracy, specifically including: Quantifying image features using the BoW method to generate a visual dictionary, quantifying the feature points in the image into visual words, thus forming a visual word histogram; Calculating the similarity between images based on the visual dictionary word histogram, screening loop candidate frames, using the ICP algorithm to align and verify the 3D Gaussian point sets of the loop candidate frames, and then adjusting the pose constraints of the loop frames through the pose graph optimization method to update the accuracy of the global map.
[0025] The beneficial effects of the present invention are as follows: The present invention constructs a semantic dense visual SLAM system based on 3D Gaussian representation that integrates Gaussian voting point cloud technology, dynamic object processing, semantic feature fusion, feature-level loss optimization, semantic-aware bundle adjustment, virtual camera view pruning, multi-channel optimization, loop detection, and differentiable rendering technology, significantly improving the mapping accuracy and robustness of the system and enabling it to exhibit more excellent performance in complex environments. The specific advantages are as follows:
[0026] (1) 3D Gaussian Representation Modeling: Using 3D Gaussian points to represent the scene, each Gaussian point has parameters such as position, covariance matrix, color, opacity, and semantic features, which can finely depict the geometric shape and semantic information of objects in the scene. Compared with traditional voxel representation and point cloud representation, it overcomes problems such as large memory consumption, low computational efficiency, and lack of surface smoothness expression;
[0027] (2) Multi-channel Optimization: Combining appearance, geometry, and semantic constraints during the optimization process, comprehensively considering different types of feature information, improving the quality and accuracy of map reconstruction, and ensuring that the optimized Gaussian point parameters are more in line with the semantic and geometric structure of the actual scene;
[0028] (3) Semantic Feature Fusion: Combining semantic features with appearance features, enabling Gaussian points to not only have the appearance information of objects but also contain their category information, thus more comprehensively understanding the scene and improving the system's understanding and modeling accuracy of the scene;
[0029] (4) Feature-level Loss Optimization: Introducing a feature-level loss function to provide higher-level guidance for the optimization process, more precisely adjusting the Gaussian point parameters to make them more in line with the semantic and geometric structure of the actual scene.
[0030] (5) Semantic-aware Bundle Adjustment: Select key frames relevant to the semantics of the current frame based on the semantic tags of Gaussian points, and jointly optimize the parameters of Gaussian points and camera poses through multi-view semantic constraints to improve the geometric accuracy and semantic consistency of mapping.
[0031] (6) Dynamic Object Handling: Detect and remove the feature points of dynamic objects using semantic information and geometric constraints, reduce the interference of dynamic objects on localization and mapping, and improve the geometric accuracy and semantic consistency of mapping.
[0032] (7) Virtual Camera View Pruning: Observe the scene from different angles through virtual cameras, perform multi-view consistency checks on Gaussian points, and remove Gaussian points that do not meet multi-view consistency to reduce the impact of incorrect points on scene representation and improve the quality and accuracy of mapping. Figure 1 Reduce the impact of incorrect points on scene representation and improve the quality and accuracy of mapping.
[0033] (8) Differentiable Rendering Technique: Use the real RGB image, depth image, and semantic mask predicted by the pre-trained segmentation head for multi-channel supervision during the rendering process to achieve joint optimization of GS parameters, improve the consistency between the rendered image and the real image, and thus improve the localization accuracy and map reconstruction quality of the system.
[0034] (9) Loop Closure Detection Optimization: Use the BoW method to quantify image features, generate a visual dictionary, screen out loop closure candidate frames, use the ICP algorithm to align and verify the 3D Gaussian point set of loop closure candidate frames, and further optimize the loop closure constraint through pose graph optimization (PGO) to improve the global map accuracy.
[0035] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0036] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0037] Figure 1 It is a structural diagram of a semantic dense visual SLAM system based on 3D Gaussian representation incorporating multiple technologies. Detailed Embodiments
[0038] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0039] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0040] Embodiment 1:
[0041] Please refer to Figure 1 , in order to achieve accurate 3D semantic segmentation and high-fidelity reconstruction, the embodiments of the present invention add various technologies on the basis of SGS-SLAM (Semantic Gaussian Splatting For Neural Dense SLAM) to construct a semantic dense visual SLAM system based on 3D Gaussian representation that integrates Gaussian voting point cloud technology, dynamic object processing, semantic feature fusion, feature-level loss optimization, semantic-aware bundle adjustment, virtual camera view pruning, multi-channel optimization, loop detection, and differentiable rendering technology, so as to improve the mapping accuracy and robustness of the system and make it perform better in complex environments.
[0042] (1) Gaussian voting point cloud technology: In the initialization stage, vote on Gaussian points according to the semantic segmentation results to determine their semantic labels, ensuring that Gaussian points of the same object have consistent semantic labels. Integrate semantic information through the voting mechanism to reduce errors and inconsistencies in semantic annotation, making the semantic information in mapping more accurate.
[0043] (2) Dynamic object processing: During the tracking and mapping process, use semantic information and geometric constraints to detect and eliminate feature points of dynamic objects. By analyzing the semantic labels and motion trajectories of Gaussian points, identify the Gaussian points corresponding to dynamic objects and exclude them during the optimization process to reduce the interference of dynamic objects on positioning and mapping, and improve the geometric accuracy and semantic consistency of mapping.
[0044] (3) Semantic Feature Fusion: In the feature extraction stage, semantic features are fused with appearance features. A pre-trained semantic segmentation network is used to extract semantic features of the RGB image, while traditional feature extraction methods are utilized to extract appearance features. Then, these two types of features are combined through a fusion module to generate a richer feature representation, which is embedded into 3D Gaussian points. The fused features enable the Gaussian points to more accurately describe scene objects, improving the system's understanding and modeling accuracy of the scene.
[0045] (4) Feature-Level Loss Optimization: A feature-level loss function is introduced during the optimization process. When calculating the loss of Gaussian point parameters, not only appearance and geometric losses are considered, but also semantic feature loss is added. By comparing the semantic features of Gaussian points with the true semantic labels, the optimization of Gaussian point parameters is guided. Feature-level loss optimization provides a higher-level guidance for the optimization process, more precisely adjusting Gaussian point parameters to better conform to the semantic and geometric structures of the actual scene.
[0046] (5) Semantic-Aware Bundle Adjustment: During the key frame selection and optimization stage, semantic associations are used for bundle adjustment. Based on the semantic labels of Gaussian points, key frames semantically related to the current frame are selected, and then the parameters of Gaussian points and camera poses are jointly optimized through multi-view semantic constraints. Using semantic information to guide key frame selection and optimization can more accurately estimate camera poses and Gaussian point parameters, improving the geometric accuracy and semantic consistency of mapping.
[0047] (6) Virtual Camera View Pruning: During the mapping process, the scene is observed from different angles through virtual cameras to perform multi-view consistency checks on Gaussian points. For each Gaussian point, it is projected from multiple perspectives through virtual cameras, and its visibility and consistency under different perspectives are calculated. Gaussian points that do not meet multi-view Figure 1 consistency are removed. Removing outlier Gaussian points reduces the impact of incorrect points on the scene representation, improving the quality and accuracy of mapping.
[0048] (7) Multi-Channel Optimization: During the mapping process, a multi-channel optimization strategy is adopted, combining appearance, geometric, and semantic constraints to optimize the parameters of Gaussian points. This includes optimizing parameters such as the color, position, and shape of Gaussian points to accurately fit the geometric shape and appearance color of the scene, while ensuring the consistency of semantic information. Through multi-channel optimization, different types of feature information are comprehensively considered to improve the quality and accuracy of map reconstruction.
[0049] (8) Differentiable Rendering Technology: During the rendering process, multi-channel supervision is performed using the true RGB image, depth image, and semantic mask predicted by the pre-trained segmentation head to jointly optimize the GS parameters. Through differentiable rendering technology, different types of supervision information are combined to improve the consistency between the rendered image and the true image, thereby enhancing the system's localization accuracy and map reconstruction quality.
[0050] (9) Loop Detection: During the process of map update and optimization, the loop detection module starts to play its role. The system uses the BoW method to quantify image features, generating a visual dictionary, quantifying the feature points in the image into visual words, and thus forming a visual word histogram h(w) to represent the visual word distribution of the image. The similarity between images is quantified by calculating the cosine similarity. By calculating the similarity between images, loop candidate frames are screened out. Next, the system uses the ICP algorithm to align the 3D Gaussian point sets of the loop candidate frames, thereby verifying the geometric consistency of the loop. In ICP, the system minimizes the geometric error between the current frame and the candidate frame by optimizing the objective function. Once the loop candidate frame passes the ICP verification, the system uses Pose Graph Optimization (PGO) to further optimize the loop constraint and improve the global map accuracy. Pose graph optimization updates the pose nodes in the pose graph by minimizing the error in the loop constraint.
[0051] As Figure 1 shown, the working process of this system is as follows:
[0052] In the initialization stage, the system first performs a preliminary 3D Gaussian representation modeling of the scene. Each Gaussian point is defined by parameters such as position (mean μ), covariance matrix Σ, color, opacity α, and semantic feature embedding. The system uses the semantic segmentation result of the initial frame and assigns semantic labels to each Gaussian point through the Gaussian voting point cloud technology. Specifically, according to the semantic segmentation result, votes are cast for the Gaussian points to determine their semantic labels, ensuring that the Gaussian points of the same object have consistent semantic labels. By integrating semantic information through the voting mechanism, errors and inconsistencies in semantic annotation are reduced, making the semantic information in mapping more accurate.
[0053] Entering the data acquisition and preprocessing stage, the system obtains image data through an RGB-D camera, including RGB images and depth images. A pre-trained semantic segmentation network is used to process the RGB images to extract semantic features. At the same time, traditional feature extraction methods (such as SIFT, ORB, etc.) or deep learning-based methods are used to extract the appearance features of the images. Then, the semantic features and appearance features are fused to generate spatially consistent high-dimensional semantic features. In this step, the semantic feature fusion technology is crucial. It combines semantic features with appearance features, enabling Gaussian points to not only know the appearance of objects but also their categories, thus understanding the scene more comprehensively.
[0054] In the core optimization and mapping stage, the system combines appearance, geometric, and semantic constraints to optimize the parameters of Gaussian points. By minimizing the difference between the rendered image and the real image, it optimizes appearance parameters such as the color of Gaussian points; through geometric constraints, it optimizes the position and shape of Gaussian points to accurately fit the geometric structure of the scene; through semantic constraints, it optimizes the semantic features of Gaussian points to more accurately represent the semantic information in the scene. In this process, the multi-channel optimization technology ensures comprehensive consideration of different types of feature information, improving the quality and accuracy of map reconstruction. At the same time, the feature-level loss optimization technology introduces a feature-level loss function to provide higher-level guidance for the optimization process, more precisely adjusting the Gaussian point parameters to make them more conform to the semantic and geometric structure of the actual scene.
[0055] In the key-frame selection and optimization stage, the system selects key frames semantically related to the current frame based on the semantic labels of Gaussian points, and then jointly optimizes the parameters of Gaussian points and the camera pose through multi-view semantic constraints. The semantic-aware bundle adjustment technology plays an important role here, using semantic information to guide key-frame selection and optimization, more accurately estimating the camera pose and Gaussian point parameters, and improving the geometric accuracy and semantic consistency of mapping. At the same time, the system uses semantic information and geometric constraints to detect and eliminate the feature points of dynamic objects, reducing the interference of dynamic objects on localization and mapping, and improving the geometric accuracy and semantic consistency of mapping.
[0056] As new frames are added, the system continuously updates and optimizes the map, adds new Gaussian points to complete incremental mapping, and further improves the accuracy and semantic consistency of the map through key-frame optimization. In this process, the virtual camera view pruning technology observes the scene from different angles through a virtual camera, performs multi-view consistency checks on Gaussian points, and eliminates Gaussian points that do not meet multi-view Figure 1 consistency, reducing the impact of incorrect points on scene representation and improving the quality and accuracy of mapping.
[0057] In the rendering and tracking stage, the system uses differentiable rendering technology to perform multi-channel supervision using the real RGB image, depth image, and semantic mask predicted by the pre-trained segmentation head, achieving joint optimization of GS parameters. The differentiable rendering technology ensures the combination of different types of supervision information here, improving the consistency between the rendered image and the real image, thereby enhancing the localization accuracy and map reconstruction quality of the system.
[0058] Finally, during the process of map updating and optimization, the loop detection module starts to play its role. The system uses the BoW method to quantify image features, generating a visual dictionary, quantifying the feature points in the image into visual words, and thus forming a visual word histogram h(w) to represent the visual word distribution of the image. The similarity between images is quantified by calculating the cosine similarity. By calculating the similarity between images, loop candidate frames are screened out. Next, the system uses the ICP algorithm to align the 3D Gaussian point sets of the loop candidate frames to verify the geometric consistency of the loop. In ICP, the system minimizes the geometric error between the current frame and the candidate frame by optimizing the objective function. Once the loop candidate frame passes the ICP verification, the system uses pose graph optimization (PGO) to further optimize the loop constraint and improve the global map accuracy. Pose graph optimization updates the pose nodes in the pose graph by minimizing the error in the loop constraint.
[0059] Comparison and verification experiment:
[0060] The loop detection module can effectively detect the situation where the agent returns to the previous scene, correct the cumulative error through loop information, and optimize the map and positioning information. Compared with the SGS-SLAM system without loop detection, in long-distance and complex environments, the accuracy of camera pose estimation can be improved by about 20% - 30%, the accuracy and consistency of map construction are better, effectively avoiding the map drift problem, and providing a more reliable basis for the navigation and decision-making of embodied agents in complex environments.
[0061] By integrating technologies such as virtual camera view pruning and loop detection into SGS-SLAM, the constructed semantic dense visual SLAM system is optimized in terms of positioning accuracy, semantic understanding, and stability, achieving more efficient real-time rendering and scene reconstruction. During the execution of embodied intelligent tasks, the overall operation efficiency is increased by about 20% - 30%, which can better meet the application requirements of agents in complex real environments and provide more powerful technical support for the development of the embodied intelligence field.
[0062] In summary, this system can make the fusion of semantic and geometric information, and the computational efficiency, scene representation, and rendering effect are better than traditional methods.
[0063] Example 2: The semantic feature fusion technology in Example 1 can use the fusion method of multi-sensor data depth maps to fuse the depth information of the lidar and the image information of the camera to improve the accuracy of scene understanding.
[0064] Example 3: For the loop detection in item 2 of Example 1, a deep learning-based loop detection method can be used to identify the similarity between images by training a neural network model, thereby screening out loop candidate frames.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A semantic dense visual SLAM method based on 3D Gaussian representation, characterized in that: The image data of the scene is obtained through the sensor, and a scene model represented by 3D Gaussian is constructed. The semantic features and appearance features are fused to generate high-dimensional semantic features. The parameters of the Gaussian points are optimized based on the appearance, geometry and semantic constraints. The scene model is updated and the camera pose is optimized according to the optimized parameters. The rendered image is generated and tracked through differentiable rendering technology. The image features are used to detect loops and optimize the global map accuracy.
2. The semantic dense visual SLAM method according to claim 1, characterized in that The method of obtaining image data of a scene through a sensor specifically includes: collecting RGB images and depth images in real time through an RGB-D sensor, extracting initial semantic features and appearance features from the RGB images, generating a preliminary geometric structure of the scene according to the depth image, aligning the preliminary geometric structure with the initial semantic features and appearance features, and generating input data for constructing a 3D Gaussian representation.
3. The semantic dense visual SLAM method according to claim 1, characterized in that The construction of the scene model of 3D Gaussian representation specifically includes: extracting position and depth information from the initial frame image data to generate multiple Gaussian points, each Gaussian point is defined by position, covariance matrix, color, opacity and semantic features; performing semantic segmentation on the initial frame according to the pre-trained network to obtain the semantic segmentation result, assigning a semantic label to each Gaussian point through Gaussian voting point cloud technology, integrating the semantic information of the Gaussian points according to the semantic labels, and generating an initial 3D Gaussian representation.
4. The semantic dense visual SLAM method according to claim 1, characterized in that The method of fusing semantic features with appearance features to generate high-dimensional semantic features specifically includes: acquiring RGB images and depth images through an RGB-D sensor, extracting semantic features from the RGB images using a pre-trained semantic segmentation network, and at the same time, extracting appearance features from the RGB images using a feature extraction method or a deep learning-based method; then, fusing the semantic features with the appearance features to generate high-dimensional semantic features consistent with the scene space.
5. The semantic dense visual SLAM method according to claim 1, characterized in that The method of optimizing the parameters of Gaussian points by combining appearance, geometry and semantic constraints specifically includes: optimizing the color parameters of Gaussian points by minimizing the difference between the rendered image and the real image; optimizing the position and shape of Gaussian points by geometric constraints; optimizing the semantic features of Gaussian points by semantic constraints; and integrating the optimization results of color parameters, position, shape and semantic features by multi-channel optimization technology and feature-level loss optimization technology to generate optimized Gaussian point parameters.
6. The semantic dense visual SLAM method according to claim 1, characterized in that The method of updating the scene model and optimizing the camera pose according to the optimized parameters specifically includes: using semantically-aware bundle adjustment technology to guide key frame selection and optimization, specifically selecting key frames related to the semantics of the current frame according to the semantic labels of the Gaussian points, and jointly optimizing the Gaussian point parameters and the camera pose through multi-view semantic constraints; at the same time, using semantic information and geometric constraints to detect and eliminate feature points of dynamic objects, using virtual views to verify the multi-view consistency of the Gaussian points from different angles, and updating the scene model and optimizing the camera pose according to the verification results.
7. The semantic dense visual SLAM method according to claim 1, characterized in that The method of generating a rendered image and tracking the positioning by using differentiable rendering technology specifically includes: using differentiable rendering technology to generate a rendered image according to a current scene model, using a pre-trained segmentation network to predict a semantic mask and a depth image of a real image, adjusting Gaussian point parameters according to the difference between the rendered image and the semantic mask and the depth image of the real image, generating a rendered image that is highly consistent with the real image, and tracking the camera positioning according to the adjusted rendered image.
8. The semantic dense visual SLAM method according to claim 1, characterized in that: The method of detecting loops using image features and optimizing the accuracy of the global map specifically includes: quantifying image features using the BoW method, generating a visual dictionary, quantifying feature points in the image into visual words, and thus forming a visual word histogram; calculating the similarity between images according to the visual dictionary word histogram, screening loop candidate frames, and aligning and verifying the 3D Gaussian point sets of the loop candidate frames using the ICP algorithm, and then adjusting the pose constraints of the loop frames through a pose graph optimization method to update the accuracy of the global map.
9. A semantic dense visual SLAM system based on 3D Gaussian representation, characterized in that: The system specifically includes: Initialization module: Construct a scene model represented by 3D Gaussian, including: extracting position and depth information from the initial frame image data, generating multiple Gaussian points, each of which is defined by position, covariance matrix, color, opacity and semantic features; performing semantic segmentation on the initial frame according to the pre-trained network, obtaining the semantic segmentation results, assigning semantic labels to each Gaussian point through Gaussian voting point cloud technology, integrating the semantic information of Gaussian points according to the semantic labels, and generating the initial 3D Gaussian representation; Data acquisition and preprocessing module: Acquire image data of the scene through sensors, including: real-time acquisition of RGB images and depth images through RGB-D sensors, extraction of initial semantic features and appearance features from RGB images, generation of preliminary geometric structures of the scene based on depth images, alignment of preliminary geometric structures with initial semantic features and appearance features, and generation of input data for constructing 3D Gaussian representation; Core optimization and mapping module: optimizes the parameters of Gaussian points by combining appearance, geometry and semantic constraints, including: optimizing the color parameters of Gaussian points by minimizing the difference between the rendered image and the real image; optimizing the position and shape of Gaussian points by geometric constraints; optimizing the semantic features of Gaussian points by semantic constraints; integrating the optimization results of color parameters, position, shape and semantic features by multi-channel optimization technology and feature-level loss optimization technology to generate optimized Gaussian point parameters; Then, the scene model is updated and the camera pose is determined based on the optimized parameters, which includes: using semantic-aware bundle adjustment technology to guide key frame selection and optimization, specifically selecting key frames related to the current frame semantics based on the semantic labels of Gaussian points, and jointly optimizing Gaussian point parameters and camera pose through multi-view semantic constraints; at the same time, semantic information and geometric constraints are used to detect and remove feature points of dynamic objects, and virtual views are used to verify the multi-view consistency of Gaussian points from different angles. Based on the verification results, the scene model is updated and the camera pose is optimized; Rendering and tracking module: Generates rendered images and tracks positioning through differentiable rendering technology, including: using differentiable rendering technology to generate rendered images according to the current scene model, using a pre-trained segmentation network to predict the semantic mask and depth image of the real image, adjusting the Gaussian point parameters according to the differences between the semantic mask and depth image of the rendered image and the real image, generating a rendered image that is highly consistent with the real image, and tracking the camera positioning according to the adjusted rendered image; Loop detection module: Use image features to detect loops and optimize the accuracy of the global map, including: using the BoW method to quantify image features, generate a visual dictionary, quantify the feature points in the image into visual words, and thus form a visual word histogram; calculate the similarity between images based on the visual dictionary word histogram, screen the loop candidate frames, and use the ICP algorithm to align and verify the 3D Gaussian point set of the loop candidate frames, and then adjust the pose constraints of the loop frames through the pose graph optimization method to update the accuracy of the global map.
Citation Information
Cited By
Semantic simultaneous localization and mapping method and system based on Gaussian splashing
CN121236253A
Real-time high-fidelity visual synchronous localization and mapping method and related device
CN121861222A