Method and device for real-time semantic segmentation of visual SLAM
Through the improved semantic segmentation network and background restoration technology, dynamic areas are identified and eliminated to generate a complete static background image, which solves the problem of incomplete map construction of the visual SLAM system in dynamic scenes and improves positioning accuracy and robustness.
Patent Information
- Application Number
- CN202510712500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-30
AI Technical Summary
Existing visual SLAM systems cannot effectively process multi-scale features in dynamic scenes and segment dynamic objects in complex scenes, resulting in incomplete map construction and affecting positioning accuracy and robustness.
An improved semantic segmentation network is used for image segmentation to identify static and dynamic areas, remove feature points in dynamic areas, and generate a complete static background image through background restoration technology. The real-time update mechanism is combined for pose estimation and map construction.
The positioning accuracy and robustness of the visual SLAM system in dynamic scenes are improved, the integrity of map construction and the relocation effect are ensured, and real-time semantic segmentation and positioning mapping of dynamic scenes are realized.
Smart Images

Figure CN120726467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of SLAM technology, and in particular to a method and device for real-time semantic segmentation visual SLAM. Background Art
[0002] With the continuous development of computer vision technology, visual SLAM (Simultaneous Localization and Mapping) technology has been widely used in fields such as robotic navigation, autonomous driving, and augmented reality. Visual SLAM systems use cameras to acquire environmental image information, achieving real-time self-localization and environmental mapping. However, in dynamic scenes, due to the presence of moving objects or people, traditional visual SLAM methods face numerous challenges, such as feature point drift caused by dynamic objects, inaccurate map construction, and reduced positioning accuracy, which seriously affect the robustness and practicality of the system.
[0003] While some existing approaches attempt to improve the performance of visual SLAM in dynamic scenes by using semantic segmentation to identify and remove dynamic objects, some shortcomings remain. For example, some methods rely solely on simple semantic segmentation networks, which cannot effectively handle dynamic object segmentation in complex scenes with multi-scale features. Other methods fail to effectively repair the background after removing dynamic objects, resulting in incomplete map construction and affecting subsequent relocalization. Summary of the Invention
[0004] The present invention proposes a real-time semantic segmentation visual SLAM method and device to solve the problems of the existing inability to effectively process multi-scale features and dynamic object segmentation in complex scenes, and the failure to effectively repair the background, resulting in incomplete map construction and affecting the subsequent relocalization effect.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] The method for real-time semantic segmentation visual SLAM comprises: extracting semantic information of an input image, performing semantic segmentation on the input image through an improved semantic segmentation network, and obtaining pixel segmentation results of different semantic categories in the image; identifying static areas and dynamic areas in the image based on the pixel segmentation results; in a visual SLAM system, using feature points of the static area to estimate camera pose, and eliminating feature points of the dynamic area to improve the positioning accuracy and robustness of the system; using static information of a previous view to synthesize an image that does not contain dynamic objects, so as to perform background repair on a key frame after eliminating the dynamic area; and updating the pose estimation and map construction of the visual SLAM system in real time to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
[0007] Furthermore, the improved semantic segmentation network includes: a high-resolution feature extraction module for extracting high-resolution image features; a multi-scale feature extraction enhancement module for enhancing the ability to extract multi-scale features; and an attention mechanism module for optimizing feature representation by adaptively adjusting feature weights.
[0008] Furthermore, the high-resolution feature extraction module includes: multiple parallel feature extraction branches, each branch is used to extract image features of different resolutions; and a feature fusion mechanism is used to fuse feature maps of different resolutions to maintain high-resolution information and integrate multi-scale features.
[0009] Furthermore, the multi-scale feature extraction enhancement module includes: multiple dilated convolutional layers with different dilation rates, which are used to expand the receptive field and enhance the ability to capture multi-scale features; and a feature splicing and adjustment mechanism is used to splice and channel-adjust feature maps with different dilation rates to generate a comprehensive multi-scale feature representation.
[0010] Furthermore, the attention mechanism module includes: adopting a channel attention mechanism to enhance important features in the image by weighting each channel; adopting a spatial attention mechanism to weight features at different spatial positions so that the network can better focus on important spatial areas; and adopting a fusion mechanism to combine channel attention and spatial attention to generate more accurate feature representation.
[0011] Furthermore, the dynamic area is identified based on instance segmentation technology: dynamic objects in the dynamic scene are extracted through an instance segmentation network, and the corresponding feature points are identified as feature points of the dynamic area.
[0012] Furthermore, the background restoration includes: projecting the RGB and depth channels of the previous key frames to the dynamic part of the current frame; leaving the part that cannot be projected blank, and filling the area blocked by the dynamic object by synthesizing static background information.
[0013] Furthermore, the real-time update of the pose estimation and map construction of the visual SLAM system is carried out to realize real-time semantic segmentation and positioning mapping of dynamic scenes, including: by adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry part of the visual SLAM system, real-time dynamic object detection and removal, and background repair, thereby realizing real-time semantic segmentation and positioning mapping of dynamic scenes.
[0014] Furthermore, the feature fusion mechanism adopts a weighted fusion strategy to dynamically adjust the fusion weight according to the importance of feature maps of different resolutions; the feature splicing and adjustment mechanism optimizes the channel weight distribution of multi-scale features by introducing an adaptive channel attention mechanism; the fusion mechanism adopts a nonlinear fusion strategy to adjust the fusion ratio of channel attention and spatial attention by introducing an activation function; the instance segmentation network adopts the Mask R-CNN architecture and introduces an adaptive feature pyramid module to improve the segmentation accuracy of dynamic objects; the background restoration also includes multi-scale background reconstruction of the occluded area of dynamic objects based on the fusion of depth information and semantic information to generate a more realistic static background image; the multi-view geometry thread adopts a multi-perspective geometric constraint algorithm and introduces a depth information correction mechanism to improve the accuracy of dynamic object detection.
[0015] A device for real-time semantic segmentation visual SLAM includes: a first unit, which extracts semantic information of an input image, performs semantic segmentation on the input image through an improved semantic segmentation network, and obtains pixel segmentation results of different semantic categories in the image; a second unit, which identifies static areas and dynamic areas in the image based on the pixel segmentation results; a third unit, which uses feature points of the static area to estimate the camera pose in the visual SLAM system, and eliminates feature points of the dynamic area to improve the positioning accuracy and robustness of the system; a fourth unit, which uses static information of a previous view to synthesize an image that does not contain dynamic objects, so as to perform background repair on the key frame after the dynamic area is eliminated; and a fifth unit, which updates the pose estimation and map construction of the visual SLAM system in real time to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
[0016] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0017] 1. The present invention extracts semantic information from an input image, performs semantic segmentation on the input image through an improved semantic segmentation network, obtains pixel segmentation results of different semantic categories in the image, thereby improving the accuracy and efficiency of semantic segmentation, forming obvious technical differences from the semantic segmentation methods in existing patents, and providing a new processing method; according to the pixel segmentation results, static areas and dynamic areas in the image are identified; in a visual SLAM system, feature points of the static area are used to estimate the camera pose, and feature points of the dynamic area are eliminated to improve the positioning accuracy and robustness of the system; static information of the previous view is used to synthesize an image that does not contain dynamic objects, so as to perform background repair on the key frame after the dynamic area is eliminated, thereby improving the mapping effect and enhancing the map repositioning effect; the pose estimation and map construction of the visual SLAM system are updated in real time to realize real-time semantic segmentation and positioning mapping of dynamic scenes, thereby improving the real-time and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flowchart of the real-time semantic segmentation visual SLAM method proposed in the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] like Figure 1 The method for real-time semantic segmentation visual SLAM shown in the figure includes: S1, extracting semantic information of an input image, performing semantic segmentation on the input image through an improved semantic segmentation network, and obtaining pixel segmentation results of different semantic categories in the image; S2, identifying static areas and dynamic areas in the image based on the pixel segmentation results; S3, in a visual SLAM system, using feature points of the static area to estimate the camera pose, and eliminating feature points of the dynamic area to improve the positioning accuracy and robustness of the system; S4, using static information of a previous view to synthesize an image that does not contain dynamic objects, so as to perform background repair on the key frame after eliminating the dynamic area; S5, updating the pose estimation and map construction of the visual SLAM system in real time to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
[0021] The previous view is the image perspective corresponding to several keyframes before the current frame. The static information contained in these previous views is used in the background inpainting step to synthesize realistic images without dynamic objects, thereby improving the mapping effect and the relocalization effect after the map is created.
[0022] Static information refers to image content in an image sequence or video stream that is unrelated to dynamic or moving objects. This information typically includes background scenes, fixed objects, or static elements in the environment. The key characteristic of static information is that it is relatively stable over time and does not change significantly over time. This information is collected, summarized, and classified into static regions.
[0023] Relatively stable content can serve as a reference for subsequent changes, preventing distortion. Using static information from previous views to synthesize realistic images without dynamic objects is a form of restoration. Improving mapping quality and post-map relocalization is actually the effect or benefit of restoration.
[0024] The specific process of repair includes: one is to extract the static information of the previous view and extract the information of the static area from the previous key frame; the other is to synthesize the static background and use this static information to synthesize the area occluded by the dynamic object in the current frame to generate a complete static background image that does not contain dynamic objects. The repair effect also includes: one is to improve the mapping effect. By repairing the area occluded by dynamic objects, the generated static background image is more complete and accurate, thereby improving the quality of map construction; the other is to improve the repositioning effect. After the map is created, when the visual SLAM system needs to reposition, the repaired static background image can provide more accurate reference information, thereby improving the accuracy and reliability of repositioning.
[0025] The improved semantic segmentation network includes: a high-resolution feature extraction module for extracting high-resolution image features; a multi-scale feature extraction enhancement module for enhancing the ability to extract multi-scale features; and an attention mechanism module for optimizing feature representation by adaptively adjusting feature weights.
[0026] The present invention aims to improve existing visual SLAM systems, or to utilize visual SLAM systems for improvement. That is, within the framework of existing visual SLAM systems, by introducing improved semantic segmentation technology and background restoration technology, the challenges faced by traditional visual SLAM systems in dynamic scenes are addressed. Specifically, the existing visual SLAM systems are improved in the following ways:
[0027] Dynamic region identification and removal: Through an improved semantic segmentation network and instance segmentation technology, dynamic regions in the image are accurately identified and feature points in these regions are removed to prevent dynamic objects from interfering with positioning and mapping.
[0028] Background inpainting: After removing dynamic areas, the static information of the previous view is used to synthesize an image without dynamic objects, filling the holes left by the movement of dynamic objects to generate a complete static background image, improving the integrity of the map and the relocalization effect.
[0029] Real-time update mechanism: By adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry part of the visual SLAM system, dynamic object detection and removal, as well as background restoration, are performed in real time, achieving real-time semantic segmentation and positioning mapping of dynamic scenes.
[0030] Modular Integration: The improved solution of this invention can be seamlessly integrated into existing visual SLAM systems, such as open-source systems like OpenVSLAM. By adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry portion of the visual SLAM system, real-time processing of dynamic scenes is achieved.
[0031] Real-time performance and efficiency: This invention improves the real-time performance and efficiency of the system through parallel processing and optimization algorithms, ensuring rapid response and accurate processing in dynamic scenarios.
[0032] Improved robustness: Through the identification and elimination of dynamic areas, as well as background restoration technology, the system's robustness and positioning accuracy are significantly improved, enabling it to operate stably even in complex dynamic scenes.
[0033] A "keyframe" is a representative image frame selected from an image sequence in a visual SLAM system to construct a map and estimate the camera pose. The selection of keyframes is based on the following criteria:
[0034] 1. Salient features: Keyframes should contain rich, stable feature points that can provide sufficient information for camera pose estimation and map construction. For example, door frames and wall edges in indoor environments, building edges and road signs in outdoor environments, etc.
[0035] 2. Motion change: The motion change between keyframes should be large enough to ensure the integrity and accuracy of map construction. Usually, a new keyframe is selected when the camera moves to a new position or perspective.
[0036] Dynamic region culling: Keyframes should avoid containing dynamic or moving objects. If a keyframe contains dynamic regions, a static background image without the dynamic objects can be generated using background inpainting techniques. These inpainted frames can then be used as new keyframes.
[0037] Real-time update in indoor environments: Suppose in an indoor environment, the camera moves along a corridor with multiple rooms on both sides. In a visual SLAM system, the process of real-time update of pose estimation and map construction is as follows:
[0038] 1. Keyframe Selection: As the camera moves from one room to another, the system selects frames containing significant features, such as door frames and wall edges, as keyframes based on the changes in motion between images and the distribution of feature points. These keyframes provide stable structural information, which facilitates camera pose estimation and map construction.
[0039] 2. Dynamic Region Identification and Removal: Using an improved semantic segmentation network and instance segmentation technology, dynamic regions in the image (such as moving pedestrians) are identified. The system removes the feature points in these dynamic regions to prevent dynamic objects from interfering with positioning and mapping.
[0040] 3. Background inpainting: If there are dynamic objects (such as pedestrians) in the keyframe, the system uses background inpainting technology to fill in the gaps left by the pedestrians' movement, generating a complete static background image. These inpainted image frames can be used as new keyframes for map updates and camera pose estimation.
[0041] 4. Real-time Update: When the camera moves to the next room, the real-time semantic segmentation visual SLAM device updates the keyframes based on the new image frames. For example, if the camera enters a new room, the real-time semantic segmentation visual SLAM device selects the image frame containing the door frame and wall edges of the new room as the new keyframe and updates the map and camera pose estimation. The real-time semantic segmentation visual SLAM device performs real-time dynamic object detection and removal, as well as background repair, by adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry portion of the visual SLAM system, achieving real-time semantic segmentation and positioning mapping of dynamic scenes.
[0042] Real-time update in outdoor environments: Suppose in an outdoor environment, the camera moves along a street with pedestrians and vehicles. Correspondingly, in a visual SLAM system, the process of real-time update of pose estimation and map construction is as follows:
[0043] 1. Keyframe Selection: As the camera moves along the street, the real-time semantic segmentation visual SLAM device selects image frames containing significant features such as building edges and road signs as keyframes based on the changes in motion between images and the distribution of feature points. These keyframes provide stable reference information, which helps estimate the camera pose and construct the map.
[0044] 2. Dynamic Region Identification and Removal: Dynamic regions in the image (such as moving vehicles and pedestrians) are identified using an improved semantic segmentation network and instance segmentation techniques. The real-time semantic segmentation visual SLAM device removes the feature points in these dynamic regions, preventing dynamic objects from interfering with positioning and mapping.
[0045] 3. Background inpainting: If a dynamic object (such as a vehicle) is present in a keyframe, the real-time semantic segmentation visual SLAM device uses background inpainting technology to fill in the gaps left by the vehicle's movement, generating a complete static background image. These inpainted image frames can be used as new keyframes for map updates and camera pose estimation.
[0046] 4. Real-time Updates: When the camera passes through an intersection, the real-time semantic segmentation visual SLAM device updates the keyframes based on the new image frames. For example, when the camera enters a new block, the real-time semantic segmentation visual SLAM device selects the image frames containing the building edges and road signs in the new block as new keyframes and updates the map and camera pose estimation. The real-time semantic segmentation visual SLAM device performs real-time dynamic object detection and removal, as well as background restoration, by adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry portion of the visual SLAM system, achieving real-time semantic segmentation and positioning mapping of dynamic scenes.
[0047] Real-time updates in dynamic scenes: Imagine a dynamic scene where the camera is moving in a shopping mall with pedestrians and moving display shelves. Correspondingly, in a visual SLAM system, the process of real-time updates of pose estimation and map construction is as follows:
[0048] 1. Keyframe Selection: As the camera moves through the mall, the real-time semantic segmentation visual SLAM device selects image frames containing mall walls, pillars, and fixed display shelves as keyframes based on the motion changes and feature point distribution between images. These keyframes provide stable reference information, facilitating camera pose estimation and map construction.
[0049] 2. Dynamic Region Identification and Removal: Dynamic regions in the image (such as moving pedestrians and display shelves) are identified using an improved semantic segmentation network and instance segmentation technology. The real-time semantic segmentation visual SLAM device removes the feature points in these dynamic regions, preventing dynamic objects from interfering with positioning and mapping.
[0050] 3. Background inpainting: If dynamic objects (such as pedestrians and display shelves) are present in the keyframes, the real-time semantic segmentation visual SLAM device uses background inpainting technology to fill in the gaps left by these moving objects, generating a complete static background image. These inpainted image frames can be used as new keyframes for map updates and camera pose estimation.
[0051] 4. Real-time update: When the camera moves to a new store area, the real-time semantic segmentation visual SLAM device will update the keyframes based on the new image frames. For example, if the camera enters a new store, the real-time semantic segmentation visual SLAM device will select the image frame containing the new store door frame and wall edges as the new keyframe and update the map and camera pose estimation. The real-time semantic segmentation visual SLAM device adds instance segmentation threads and multi-view geometry threads in parallel to the visual odometry part of the visual SLAM system to perform real-time dynamic object detection and removal, as well as background repair, to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
[0052] The high-resolution feature extraction module includes: multiple parallel feature extraction branches, each branch is used to extract image features at different resolutions; a feature fusion mechanism is used to fuse feature maps of different resolutions to maintain high-resolution information and integrate multi-scale features.
[0053] The multi-scale feature extraction enhancement module includes: multiple hollow convolution layers with different hollowing rates, which are used to expand the receptive field and enhance the ability to capture multi-scale features; and a feature splicing and adjustment mechanism is used to splice and adjust the channels of feature maps with different hollowing rates to generate a comprehensive multi-scale feature representation.
[0054] The attention mechanism module includes: adopting a channel attention mechanism to enhance important features in the image by weighting each channel; adopting a spatial attention mechanism to weight features at different spatial positions so that the network can better focus on important spatial areas; and adopting a fusion mechanism to combine channel attention and spatial attention to generate more accurate feature representation.
[0055] The dynamic area is identified based on instance segmentation technology: dynamic objects in the dynamic scene are extracted through an instance segmentation network, and the corresponding feature points are identified as feature points of the dynamic area.
[0056] The background restoration includes: projecting the RGB and depth channels of the previous key frames to the dynamic part of the current frame; leaving the parts that cannot be projected blank, and filling the areas blocked by dynamic objects by synthesizing static background information.
[0057] The real-time update of the pose estimation and map construction of the visual SLAM system to achieve real-time semantic segmentation and positioning mapping of dynamic scenes includes: adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry part of the visual SLAM system, performing real-time dynamic object detection and removal, and background repair, thereby achieving real-time semantic segmentation and positioning mapping of dynamic scenes.
[0058] The feature fusion mechanism adopts a weighted fusion strategy to dynamically adjust the fusion weight according to the importance of feature maps of different resolutions; the feature splicing and adjustment mechanism optimizes the channel weight distribution of multi-scale features by introducing an adaptive channel attention mechanism; the fusion mechanism adopts a nonlinear fusion strategy to adjust the fusion ratio of channel attention and spatial attention by introducing an activation function; the instance segmentation network adopts the Mask R-CNN architecture and introduces an adaptive feature pyramid module to improve the segmentation accuracy of dynamic objects; the background restoration also includes multi-scale background reconstruction of the occluded area of dynamic objects based on the fusion of depth information and semantic information to generate more realistic static background images; the multi-view geometry thread adopts a multi-perspective geometric constraint algorithm and introduces a depth information correction mechanism to improve the accuracy of dynamic object detection.
[0059] Defects of the existing technology:
[0060] Hole problem: The movement of dynamic objects will cause holes in the map construction process, affecting the integrity and practicality of the map.
[0061] Error accumulation: The presence of dynamic objects can lead to error accumulation during map construction, further affecting the accuracy and reliability of the map.
[0062] Improvements of the present invention:
[0063] Multi-scale background reconstruction: This technology uses a fusion of depth and semantic information to reconstruct the multi-scale background of areas occluded by dynamic objects. This technology generates more realistic static background images and fills the gaps left by moving objects, thereby improving the integrity and accuracy of the map.
[0064] Real-time Update Mechanism: This invention adds an instance segmentation thread and a multi-view geometry thread to the visual odometry portion of the visual SLAM system, performing real-time dynamic object detection and removal, as well as background inpainting. This real-time update mechanism effectively handles complex dynamic scenes, avoids error accumulation, and improves the accuracy and reliability of map construction.
[0065] In terms of poor relocation effect, the defects of the existing technology are:
[0066] Dynamic interference: The presence of dynamic objects will affect the repositioning effect of the system, resulting in a decrease in positioning accuracy.
[0067] Incomplete map: Incomplete map construction will further affect the accuracy of relocalization, especially in scenes with frequent movement of dynamic objects.
[0068] Improvements of the present invention:
[0069] Static background restoration: Through background restoration technology, the present invention can generate a complete static background image, providing more accurate reference information for relocalization. This technology significantly improves the accuracy and reliability of relocalization.
[0070] Real-time update: The real-time update mechanism ensures that the system can handle changes in dynamic scenes in a timely manner, avoids interference with relocalization caused by dynamic objects, and further improves the robustness of the system and the relocalization effect.
[0071] In terms of lack of real-time performance, the defects of existing technologies are:
[0072] Complex calculations: Existing technologies often require complex calculations when processing dynamic scenes, making it difficult to meet real-time requirements.
[0073] Inefficiency: The detection and elimination of dynamic objects usually require more computing resources, resulting in longer system response times and difficulty in achieving real-time performance.
[0074] Improvements of the present invention:
[0075] Parallel Processing: This invention implements real-time dynamic object detection and removal, as well as background inpainting, by adding instance segmentation threads and multi-view geometry threads to the visual odometry portion of the visual SLAM system. This parallel processing mechanism significantly improves the system's real-time performance and efficiency.
[0076] Optimization algorithm: The improved semantic segmentation network and instance segmentation technology improve the accuracy and efficiency of semantic segmentation by introducing a high-resolution feature extraction module, a multi-scale feature extraction enhancement module, and an attention mechanism module, further optimizing the real-time performance of the system.
[0077] Specific examples:
[0078] Regarding moving pedestrians: In scenes such as shopping malls or streets, moving pedestrians are common dynamic objects. When dealing with such scenes, existing technologies often cause feature point drift due to pedestrians, affecting positioning accuracy and map construction. The present invention, through an improved semantic segmentation network and instance segmentation technology, can accurately identify and remove feature points corresponding to pedestrians, avoiding the interference of pedestrians on positioning and mapping. At the same time, background restoration technology can fill the gaps left by pedestrians and generate a complete static background image, significantly improving the mapping and relocation effects.
[0079] Moving vehicles or autonomous driving: In autonomous driving or traffic monitoring scenarios, moving vehicles are the main dynamic objects. When dealing with such scenarios, existing technologies often cause incomplete map construction due to vehicles, affecting the robustness of the system and the repositioning effect. The present invention uses an improved semantic segmentation network and instance segmentation technology to accurately identify and eliminate feature points corresponding to vehicles, thereby improving positioning accuracy and map construction accuracy. Background repair technology can fill the gaps left by vehicle movement and generate a complete static background image, further improving the integrity of the map and the repositioning effect.
[0080] A device for real-time semantic segmentation visual SLAM includes: a first unit, which extracts semantic information of an input image, performs semantic segmentation on the input image through an improved semantic segmentation network, and obtains pixel segmentation results of different semantic categories in the image; a second unit, which identifies static areas and dynamic areas in the image based on the pixel segmentation results; a third unit, which uses feature points of the static area to estimate the camera pose in the visual SLAM system, and eliminates feature points of the dynamic area to improve the positioning accuracy and robustness of the system; a fourth unit, which uses static information of a previous view to synthesize an image that does not contain dynamic objects, so as to perform background repair on the key frame after the dynamic area is eliminated; and a fifth unit, which updates the pose estimation and map construction of the visual SLAM system in real time to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
[0081] A real-time semantic segmentation visual SLAM method and device, through an improved semantic segmentation network and instance segmentation technology, combined with background restoration and real-time update mechanism, effectively solves the positioning accuracy and mapping effect problems of the visual SLAM system in dynamic scenes. Its innovation lies in four aspects
[0082] 1. Improved semantic segmentation network. This network uses a high-resolution feature extraction module, a multi-scale feature extraction enhancement module, and an attention mechanism module to improve the accuracy and efficiency of semantic segmentation, enabling more accurate identification of static and dynamic regions in images.
[0083] 2. Dynamic region identification and removal. Based on instance segmentation technology, the Mask R-CNN architecture, which incorporates an adaptive feature pyramid module, improves the segmentation accuracy of dynamic objects, effectively removes feature points in dynamic areas, and prevents interference from dynamic objects on positioning and mapping.
[0084] 3. Background inpainting technology. After removing dynamic areas, it uses static information from previous views to synthesize a realistic image without dynamic objects. By fusing depth and semantic information, it reconstructs the multi-scale background of areas occluded by dynamic objects, improving mapping quality and post-map relocalization.
[0085] 4. Real-time update mechanism. By adding an instance segmentation thread and a multi-view geometry thread in parallel to the visual odometry portion of the visual SLAM system, dynamic object detection and removal, as well as background inpainting, are performed in real time. This enables real-time semantic segmentation and positioning mapping of dynamic scenes, improving the system's real-time performance and adaptability.
[0086] The above description is a detailed description of the preferred embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications completed under the technical spirit suggested by the present invention should fall within the patent scope covered by the present invention.
Claims
1. A method for real-time semantic segmentation visual SLAM, characterized in that: include: Extract the semantic information of the input image, perform semantic segmentation on the input image through the improved semantic segmentation network, and obtain pixel segmentation results of different semantic categories in the image; Identifying static areas and dynamic areas in the image based on the pixel segmentation result; In a visual SLAM system, the feature points of the static area are used to estimate the camera pose, and the feature points of the dynamic area are eliminated to improve the positioning accuracy and robustness of the system; Using static information from previous views to synthesize images without dynamic objects, the background of the keyframe after dynamic areas are removed is repaired. Update the pose estimation and map construction of the visual SLAM system in real time to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.
2. The method for real-time semantic segmentation visual SLAM according to claim 1, wherein The improved semantic segmentation network includes: High-resolution feature extraction module, used to extract high-resolution image features; Multi-scale feature extraction enhancement module, used to enhance the ability to extract multi-scale features; Attention mechanism module, used to optimize feature representation by adaptively adjusting feature weights.
3. The method for real-time semantic segmentation visual SLAM according to claim 2, wherein The high-resolution feature extraction module includes: Multiple parallel feature extraction branches, each branch is used to extract image features at different resolutions; A feature fusion mechanism is used to fuse feature maps of different resolutions to maintain high-resolution information and integrate multi-scale features.
4. The method for real-time semantic segmentation visual SLAM according to claim 3, wherein The multi-scale feature extraction enhancement module includes: Multiple dilated convolutional layers with different dilation rates are used to expand the receptive field and enhance the ability to capture multi-scale features; A feature splicing and adjustment mechanism is used to splice and channel-adjust feature maps with different void ratios to generate a comprehensive multi-scale feature representation.
5. The method for real-time semantic segmentation visual SLAM according to claim 4, wherein The attention mechanism module includes: Adopt channel attention mechanism to enhance important features in the image by weighting each channel; The spatial attention mechanism is used to weight features at different spatial locations, so that the network can better focus on important spatial regions. A fusion mechanism is adopted to combine channel attention and spatial attention to generate more accurate feature representation.
6. The method for real-time semantic segmentation visual SLAM according to claim 5, wherein Identify the dynamic region based on instance segmentation technology: Dynamic objects in dynamic scenes are extracted through instance segmentation network, and their corresponding feature points are identified as feature points of dynamic areas.
7. The method for real-time semantic segmentation visual SLAM according to claim 6, wherein The background repair includes: Project the RGB and depth channels of the previous keyframes to the dynamic part of the current frame; The parts that cannot be mapped are left blank, and the areas blocked by dynamic objects are filled by synthesizing static background information.
8. The method for real-time semantic segmentation visual SLAM according to claim 7, wherein The real-time update of the pose estimation and map construction of the visual SLAM system to achieve real-time semantic segmentation and positioning mapping of dynamic scenes includes: By adding instance segmentation threads and multi-view geometry threads in parallel to the visual odometry part of the visual SLAM system, dynamic object detection and removal, as well as background repair, are performed in real time, thereby achieving real-time semantic segmentation and positioning mapping of dynamic scenes.
9. The method for real-time semantic segmentation visual SLAM according to claim 8, wherein The feature fusion mechanism adopts a weighted fusion strategy to dynamically adjust the fusion weight according to the importance of feature maps of different resolutions; The feature splicing and adjustment mechanism optimizes the channel weight distribution of multi-scale features by introducing an adaptive channel attention mechanism; The fusion mechanism adopts a nonlinear fusion strategy and adjusts the fusion ratio of channel attention and spatial attention by introducing an activation function; The instance segmentation network adopts the Mask R-CNN architecture and introduces an adaptive feature pyramid module to improve the segmentation accuracy of dynamic objects; The background restoration also includes multi-scale background reconstruction of the dynamic object occluded area based on the fusion of depth information and semantic information to generate a more realistic static background image; The multi-view geometry thread adopts a multi-view geometry constraint algorithm and introduces a depth information correction mechanism to improve the accuracy of dynamic object detection.
10. A real-time semantic segmentation visual SLAM device, characterized in that: include: The first unit extracts the semantic information of the input image and performs semantic segmentation on the input image through the improved semantic segmentation network to obtain pixel segmentation results of different semantic categories in the image; The second unit identifies the static area and the dynamic area in the image according to the pixel segmentation result; Unit 3, in a visual SLAM system, using the feature points of the static area to estimate the camera pose, and eliminating the feature points of the dynamic area to improve the positioning accuracy and robustness of the system; The fourth unit uses the static information of the previous view to synthesize an image without dynamic objects to perform background restoration on the key frame after removing the dynamic area; Unit 5: Real-time update of the pose estimation and map construction of the visual SLAM system to achieve real-time semantic segmentation and positioning mapping of dynamic scenes.