Semantic mapping method and system based on visual SLAM and YOLOv5

By introducing YOLOv5 semantic segmentation and surface element superset fusion technology into the SLAM algorithm, the problem that traditional SLAM cannot build semantic maps is solved, and high-precision, globally consistent semantic map construction is realized, which improves the semantic understanding ability of robot environment interaction.

CN119992145AActive Publication Date: 2025-05-13YANTAI ZHISHI ELECTRONIC TECHNOLOGY CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411887649.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Traditional SLAM algorithms cannot build environment maps with semantic information, limiting the semantic understanding ability of robots when performing advanced environmental interaction tasks.

Method used

The semantic graphing method based on visual SLAM and YOLOv5 is adopted, and the semantic information is obtained by optimizing line feature matching and pose optimization, combined with the YOLOv5 semantic segmentation network, and a globally consistent semantic map is constructed using the method of face elements and supersets.

Benefits of technology

It realizes the construction of high-precision, globally consistent semantic maps, improves the robot's semantic understanding ability in environmental interaction, and surpasses the performance of ManhattanSLAM and other advanced algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992145A_ABST
    Figure CN119992145A_ABST
Patent Text Reader

Abstract

According to the semantic mapping method and system based on the visual SLAM and the YOLOv5, positioning optimization and global consistent semantic mapping are achieved. According to the design, firstly, a more accurate camera pose is obtained through a line feature optimization and weighted optimization strategy so as to improve the mapping quality; then semantic information is obtained in combination with a YOLOv5 semantic segmentation network, a semantic map construction method combining surface elements and supersets is used, the problem of semantic inconsistency in the map construction process is solved, and finally a global consistent semantic map is constructed. According to the method, real-time tracking, mapping and back-end optimization can be realized, so that high-precision pose estimation and global consistent semantic map construction are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction and semantic map construction, and in particular to a semantic mapping method and system based on visual SLAM and YOLOv5. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) uses sensors on a robot to achieve positioning, tracking, and map building in an unknown environment. Among them, camera-based SLAM is called visual SLAM. Compared with other sensors, cameras have the characteristics of simple installation, low price, and rich data.

[0003] With the vigorous development of science and technology, mobile robots have begun to be widely used in various industries. Among them, simultaneous localization and mapping (SLAM) technology provides solutions for robot navigation and obstacle avoidance in unknown environments, and has become a key technology in related fields. However, traditional SLAM algorithms cannot build environmental maps with semantic information, which makes it impossible for robots to have a high-level semantic understanding of the environment when performing advanced environmental interaction tasks, thus limiting the development of robots. Therefore, realizing a SLAM algorithm that can build an environmental semantic map has become a key step in improving the intelligence of robots.

[0004] The commonly used visual SLAM algorithms are usually located by tracking features on the image. For example, the most advanced ManhattanSLAM achieves accurate positioning by tracking point, line and surface features and Manhattan constraints, and uses the surface element model to build a dense map of the environment. In ManhattanSLAM, point features provide the most stable constraints for pose estimation with their stability and matching accuracy, while line features and surface features not only provide additional constraints for pose optimization in complex environments, but also provide sufficient constraints for positioning in weak texture environments where the number of point features is insufficient. However, the stability of line features is affected by feature extraction algorithms and depth map holes, which may introduce erroneous constraints and reduce system accuracy. Moreover, ManhattanSLAM cannot construct semantic maps containing semantic information, and cannot provide information for the system's advanced environmental interaction tasks such as object recognition and positioning, and environmental identification. Existing visual SLAM algorithms that can construct environmental semantic information often do not consider the semantic inconsistency problem in the mapping process, and the quality of the constructed semantic maps is poor. Summary of the invention

[0005] The purpose of the present invention is to provide a semantic mapping method and system based on visual SLAM and YOLOv5 to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A semantic mapping method based on visual SLAM and YOLOv5, including:

[0008] During the positioning process, line features are merged to improve the stability of line features. The depth of line feature endpoints is optimized through clustering and straight line fitting to solve the problem of missing line feature endpoint depth caused by depth map holes. The line feature matching method is improved to make line feature matching more efficient and faster;

[0009] After completing a stable match, the algorithm will combine point, line and surface errors to jointly optimize the camera pose, adding weights to the point, line and surface error model to obtain a more accurate pose.

[0010] The key frame is selected by judging the feature points of the current frame and the inter-frame pose. After confirming the key frame, the frame is passed to the backend optimization and mapping thread;

[0011] In the map construction thread, semantic segmentation is first performed on the new keyframe to obtain semantic information and construct a semantic map to prepare for the subsequent semantic map construction;

[0012] After obtaining the semantic map, superpixel segmentation is performed based on the color map, and then superpixels with semantic information are obtained in combination with the semantic map. Then, data fusion semantic superpixels are extracted from the facet database and the superset database. Semantic facets are created based on the remaining superpixels. After creation, the newly created semantic facets are superset fused, and the semantic facets that have not been fused are used to create supersets. Finally, the superset is updated based on the facet data under the superset, and a globally consistent semantic map is constructed based on the superset information.

[0013] Preferably, after detecting the line features, unstable short line features are merged into more stable long line features according to the following constraints: 1) The angle between the line features l1 and l2 is within the threshold. 2) The distance between the line features l1 and l2 is within the threshold, and the line feature distance is determined by the distance between the midpoints of the line segments. 3) To prevent the line features from being included, the projections of the two line features on the XY axis are set to be non-repetitive.

[0014] Preferably, the depth loss problem in the three-dimensionalization process of line features is optimized to obtain more accurate depth values. Specifically, for the detected line features, all pixel points with depth are first obtained in combination with the depth map and projected to the camera coordinate system to obtain three-dimensional points. After obtaining the three-dimensional points, distance clustering is performed on them, and then straight line fitting is performed based on the clustering results. Finally, the depth value of the line feature endpoint is obtained based on the straight line fitting result.

[0015] Preferably, a matching method based on point-to-line distance constraint. The specific steps are as follows: 1) Line feature sampling, that is, uniformly sampling nine points on the line feature to be matched. These sampling points can be selected at equal intervals on the line feature to ensure that different parts of the line feature are covered. 2) Back-projection to the image, that is, back-projecting the sampling points onto the image plane to obtain the corresponding pixel coordinates. This can be calculated using the camera's intrinsic parameter matrix and depth information. 3) Calculating the distance, that is, for each sampling point, calculating its average distance to the matching line segment. This can be obtained by accumulating the distance between the sampling point and the matching line segment and dividing it by the number of sampling points to obtain the average distance. 4) Matching judgment, that is, comparing the average distance with a pre-set threshold. If the average distance is less than the threshold, the line feature is considered to be matched. Through these four steps, the algorithm can achieve fast and stable line feature matching while ensuring matching accuracy.

[0016] Preferably, since point feature matching is more accurate and stable, a higher weight is given to its error term. To ensure the real-time performance of the system, avoid complex weight setting methods. Point, line and surface weight w p 、w l and w φ Number of matches based on point features N p The allocation is set up as shown in the formula:

[0017]

[0018] The threshold N is set to 40. The Levenberg-Marquardt algorithm is used to minimize the following energy function to optimize the camera pose:

[0019]

[0020] Where {R* cw, t* cw} is the current camera pose and is the optimal solution; N p 、N l and N φ are the number of point feature matches, line feature matches, and surface feature matches of the current frame respectively; w p 、w l and w φ is the weight of the error term; H p , H l and H φ is the Huber robust loss function; e pi 、e lj and e φk are the errors of points, lines and surfaces respectively; Σ p ,Σ l and Σ φ Represents the covariance matrix of points, lines, and areas.

[0021] Preferably, the keyframe is passed into the trained YOLOv5s-seg model, and the model outputs a set of semantic information. Each semantic information contains the category to which the object belongs, the confidence, the prediction box and the segmentation mask. Based on the output information, a semantic graph is constructed to save the semantics. The semantic graph is set as a two-channel image, the first channel stores the semantic category C, and the second channel stores the semantic confidence P.

[0022] Preferably, a dense semantic map is constructed by combining semantic information with the facet-based dense map construction method. The mapping method reconstructs a globally consistent semantic map by inputting a color map, a depth map, and a frame pose of the system.

[0023] Preferably, instead of creating a bin for each pixel, the color image is first segmented into superpixels, and then bins are constructed for superpixels. This can not only greatly reduce the amount of system computation, but also reduce the impact of outliers and depth holes in low-quality depth images on map construction. The improved SLIC method is used to obtain superpixels, and semantic information is added to each superpixel. For each superpixel, its category C is initially defined. sp is the background object, and the confidence P sp Set to 0.2. In the superpixel segmentation process, the semantic category of the superpixel is updated to the semantic category with the largest number of pixels within the superpixel range, and the confidence is set to the confidence of the category. The confidence update is shown in the formula:

[0024]

[0025] Among them C TH Represents the semantic category as an object; C BG represents the semantic category as background; P c (*) is the confidence of the corresponding category; MaxNum(*) is the semantic category with the largest number of pixels in the range.

[0026] Preferably, semantic information is also added to the face element, and its initial category C is set sl is the background object, and the confidence P sl Set to 0.2. In addition, the serial number of the superset to which each facet belongs is added to establish the connection between the facet and the superset. After completing the semantic superpixel segmentation, the facets in the system are used to fuse these superpixels, and the facets and corresponding supersets are updated at the same time.

[0027] Furthermore, the physical distance and normal vector angle between the facet and superpixel determine whether to fuse the two. If the conditions are met, the superpixel is combined to update the attributes of the facet. Among them, the semantically related attributes are updated by comparing the semantic confidence P of the semantic facet. sl and the semantic confidence P of the superpixel sp To determine. sp >Psl When C sl =C sp , P sp =P sl When the semantic information of a face element changes, the superset that controls the face element is updated. First, determine whether the face element is merged into the superset. If not, create a new superset based on the face element data. The category of the newly created superset is defined as the semantic category of the face element, and the corresponding confidence PC st in the confidence list is updated as: PC st = PC st + P sl , radius R st Set to the surface radius R sl The superset position is set to the position of the face element. If the face element has been fused, only the semantic confidence attribute of the corresponding superset is updated.

[0028] Optionally, for the unfused superpixels, new facets are created based on them. The relevant attributes of the new facets are set to the corresponding attributes of the superpixels. In order to control the semantic information of these facets, the new facets are fused using the superset in the system. At the same time, the superset also needs to be updated to reflect the existence of these new facets.

[0029] Furthermore, after creating a new facet, the facets in the system are used for fusion. When the following conditions are met at the same time, the facets are fused with the corresponding superset:

[0030] Preferably, the facet class C sl Superset category C st or background C BG , that is: C sl ∈{C st ,C BG};

[0031] Preferably, the distance D between the facet and the superset is smaller than the radius R of the superset. st Add the radius R of the panel sl , that is: D <R st +R sl ;

[0032] Optionally, if the new facet and all supersets do not meet the above conditions at the same time, and the category of the new facet is not background, a new superset is created based on the new facet. The creation steps here are the same as those for creating a new superset when facets are fused, so they are not repeated here. When the above conditions are met, the semantic confidence attribute of the superset is updated based on the facet.

[0033] Furthermore, after the face element update is completed, all supersets are updated. Specifically, the algorithm updates the center position, radius R, and st and category C stAttributes. For a superset, traverse all the face elements under it and record the maximum radius, the maximum value and the minimum value in the XYZ direction respectively. Use the collected data to update the superset. Specifically, update the center position of the superset as:

[0034]

[0035] Update the superset radius to:

[0036]

[0037] Update superset semantic categories to:

[0038]

[0039] PC st is the confidence of a category among all categories.

[0040] In the case of inconsistent semantic segmentation results for the same object during the mapping process, the semantics of the object facets constructed by the present invention will change, but the semantic information of the superset to which it belongs will not change suddenly. Moreover, the semantic information of the facets in the map is provided by the superset, thus maintaining the global semantic consistency of the object.

[0041] On the other hand, to achieve the above object, the present invention also provides a semantic mapping system of visual SLAM and YOLOv5, comprising:

[0042] Line feature merging module: merges the detected short line features into long line features.

[0043] Line feature endpoint depth optimization module: perform straight line fitting based on the points with depth on the line feature, and obtain the line feature endpoint depth through the fitted straight line.

[0044] Line feature matching module: Determine the matching degree of line features by calculating the geometric constraints between line features.

[0045] Weighted optimization module: Add weights to the optimization items involved in posture optimization, and the weights change dynamically.

[0046] Globally consistent semantic mapping module: Constructs semantic maps through image data and pose data, and maintains semantic consistency during the mapping process.

[0047] Preferably, the semantic mapping module includes:

[0048] The first point is superpixel segmentation: the input image data is first segmented into superpixels, and then semantic superpixels are given semantics to construct semantic superpixels;

[0049] The second point is surface element fusion: surface element fusion is created through superpixels, and the original surface data and superset data are updated at the same time.

[0050] The third point is superset update: update the superset information through the facets under the superset.

[0051] Compared with the prior art, the present invention has the following advantages and technical effects:

[0052] (1) The present invention provides a semantic mapping method based on visual SLAM and YOLOv5. The design first obtains more accurate camera poses to improve the mapping quality through line feature optimization and weighted optimization strategies; then, the semantic information is obtained by combining the YOLOv5 semantic segmentation network, and a semantic map construction method combining facets and supersets is used to solve the semantic inconsistency problem in the mapping process, and finally a globally consistent semantic map is constructed. Experimental verification shows that the proposed algorithm not only exceeds the ManhattanSLAM algorithm in terms of accuracy, but also constructs a semantic map that is better than other advanced mapping algorithms.

[0053] (2) Based on ManhattanSLAM, the present invention optimizes line features to solve the problem of unstable matching of line feature extraction; adds weights for optimization to obtain more accurate poses; combines the YOLOv5s-seg semantic segmentation model to semantically segment key frames, obtains semantic information of the surrounding environment while maintaining algorithm efficiency; and proposes a new incremental semantic map construction method to solve the impact of semantic over-segmentation or under-segmentation on map construction, and uses the superset as the semantic controller of the face element after semantic regularization. Experimental verification shows that the present invention can not only obtain high-precision poses, but also build a globally consistent semantic map of the environment.

[0054] (3) The present invention improves line feature extraction and matching, adds weights for optimization, and obtains better positioning accuracy and robustness than other advanced algorithms. The present invention obtains higher positioning accuracy than other advanced algorithms; in terms of mapping, the present invention can not only reconstruct high-quality semantic maps, but also effectively handle semantic inconsistency problems during the mapping process. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 A flowchart of a semantic mapping method based on visual SLAM and YOLOv5 provided in an embodiment of the present invention;

[0057] Figure 2It is a framework diagram of an algorithm system of a semantic mapping method based on visual SLAM and YOLOv5 according to an embodiment of the present invention;

[0058] Figure 3 It is a flowchart of line feature endpoint depth optimization according to an embodiment of the present invention;

[0059] Figure 4 This is a schematic diagram of the architecture of a globally consistent semantic mapping method according to an embodiment of the present invention;

[0060] Figure 5 This is a flowchart of facet fusion update according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] This embodiment provides a semantic mapping method based on visual SLAM and YOLOv5. The overall process is as follows: Figure 1 As shown in the figure, the overall framework of the algorithm is as follows Figure 2 As shown, the algorithm flow includes:

[0062] During the positioning process, line features are merged to improve the stability of line features. The depth of line feature endpoints is optimized through clustering and straight line fitting to solve the problem of missing line feature endpoint depth caused by depth map holes. The line feature matching method is improved to make line feature matching more efficient and faster;

[0063] Specifically, after detecting the line features, this example merges the unstable short line features into more stable long line features according to the following constraints: 1) The angle between the two-dimensional line features l1 and l2 is within the threshold. 2) The distance between the two-dimensional line features l1 and l2 is within the threshold, and the line feature distance is determined by the distance between the midpoints of the line segments. 3) To prevent the line features from being included, the projections of the two line features on the XY axis are set to be non-repetitive.

[0064] Furthermore, the depth loss problem in the three-dimensionalization process of line features is optimized to obtain more accurate depth values. The process is as follows: Figure 2 Specifically, for the detected line features, firstly, all the pixels with depth are obtained by combining the depth map and projected to the camera coordinate system to obtain 3D points. After obtaining the 3D points, distance clustering is performed on them, and then straight line fitting is performed based on the clustering results. Finally, the depth value of the line feature endpoint is obtained based on the straight line fitting result.

[0065] Furthermore, a matching method based on point-to-line distance constraint is used to improve the line feature matching method, making line feature matching more efficient and faster. The specific steps are as follows: 1) Line feature sampling, that is, uniformly sampling nine points on the line feature to be matched. These sampling points can be selected at equal intervals on the line feature to ensure that different parts of the line feature are covered. 2) Back-projection to the image, that is, back-projecting the sampling points onto the image plane to obtain the corresponding pixel coordinates. This can be calculated using the camera's intrinsic parameter matrix and depth information. 3) Calculating the distance, that is, for each sampling point, calculating its average distance to the matching line segment. This can be obtained by accumulating the distance between the sampling point and the matching line segment and dividing it by the number of sampling points to obtain the average distance. 4) Matching judgment, that is, comparing the average distance with a pre-set threshold. If the average distance is less than the threshold, the line feature is considered to be matched. Through these four steps, the algorithm can achieve fast and stable line feature matching while ensuring matching accuracy.

[0066] After completing stable matching, the algorithm will jointly optimize the camera pose with point, line and surface errors, adding weights to the point, line and surface error model to obtain a more accurate pose.

[0067] Specifically, after completing the stable matching, this example gives the error term a higher weight because the point feature matching is more accurate and stable. To ensure the real-time performance of the system, avoid complex weight setting methods. p 、w l and w φ Number of matches based on point features N p The allocation is set up as shown in the formula:

[0068]

[0069] The threshold N is set to 40.

[0070] Furthermore, the Levenberg-Marquardt algorithm is used to minimize the following energy function to optimize the camera pose:

[0071]

[0072] where {R * cw ,t * cw} is the current camera pose, which is the optimal solution; N p 、N l and N φ are the number of point feature matches, line feature matches, and surface feature matches of the current frame respectively; w p 、w l and w φ is the weight of the error term; H p , Hl and H φ are the Huber robust loss functions for points, lines, and surfaces respectively; e pi 、e lj and e φk are the errors of points, lines and surfaces respectively; Σ p ,Σ l and Σ φ Represents the covariance matrix of points, lines, and areas.

[0073] The key frame is selected by judging the feature points of the current frame and the inter-frame pose. After confirming the key frame, the frame is passed to the backend optimization and mapping thread. The mapping method architecture is as follows Figure 4 As shown;

[0074] Specifically, in this example, the key frame is passed into the trained YOLOv5s-seg model, and the model outputs a set of semantic information. Each semantic information contains the category to which the object belongs, the confidence, the prediction box, and the segmentation mask. Based on the output information, a semantic graph is constructed to save the semantics. The semantic graph is set as a two-channel image, the first channel stores the semantic category C, and the second channel stores the semantic confidence P.

[0075] In the map construction thread, semantic segmentation is first performed on the new keyframe to obtain semantic information and construct a semantic map to prepare for the subsequent semantic map construction;

[0076] Specifically, based on the facet-based dense map construction method, a dense semantic map is constructed by combining semantic information. The mapping method reconstructs a globally consistent semantic map by inputting the system's color map, depth map, and frame pose. Unlike creating facets for each pixel, the color map is first segmented into superpixels, and then facets are constructed for the superpixels. This not only greatly reduces the amount of computation in the system, but also reduces the impact of outliers and depth holes in low-quality depth maps on mapping. Superpixels are obtained using the improved SLIC method, and semantic information is added to each superpixel. For each superpixel, its category Csp is initially defined as a background object, and the confidence Psp is set to 0.2. During the superpixel segmentation process, the semantic category of the superpixel is updated to the semantic category with the largest number of pixels within the superpixel range, and the confidence is set to the confidence of the category. The confidence update is shown in the formula:

[0077]

[0078] Where C is the semantic category, C TH Represents the semantic category as an object; C BG represents the semantic category as background; P c (*) is the confidence of the corresponding category; MaxNum(*) is the semantic category with the largest number of pixels in the range.

[0079] Furthermore, semantic information is added to the face element and its initial category C is set sl is the background object, and the confidence P sl Set to 0.2. In addition, the serial number of the superset to which each facet belongs is added to establish the connection between the facet and the superset. After completing the semantic superpixel segmentation, the facets in the system are used to fuse these superpixels, and the facets and corresponding supersets are updated at the same time.

[0080] The update process of the face element is as follows Figure 5 As shown in Figure 1, first, the physical distance and normal vector angle between the face element and the superpixel determine whether to fuse the two. If the conditions are met, the superpixel is combined to update the attributes of the face element. Among them, the semantically related attributes are updated by comparing the semantic confidence P of the semantic face element. sl and the semantic confidence P of the superpixel sp To determine. sp >P sl When C sl =C sp , P sp =P sl When the semantic information of a face element changes, the superset that controls the face element is updated. First, determine whether the face element is merged into the superset. If not, create a new superset based on the face element data. The category of the newly created superset is defined as the semantic category of the face element, and the corresponding confidence PC st in the confidence list is updated as: PC st = PC st + P sl , radius R st Set to the surface radius R sl The superset position is set to the position of the face element. If the face element has been fused, only the semantic confidence attribute of the corresponding superset is updated.

[0081] Furthermore, for the unfused superpixels, new facets are created based on them. The relevant attributes of the new facets are set to the corresponding attributes of the superpixels. In order to control the semantic information of these facets, the new facets are fused using the superset in the system. At the same time, the superset also needs to be updated to reflect the existence of these new facets.

[0082] After creating a new facet, use the facets in the system to fuse them. When the following conditions are met at the same time, use the corresponding superset to fuse the facets:

[0083] (1) Facet category C sl Superset category C st or background C BG , that is: C sl ∈{C st ,C BG};

[0084] (2) The distance D between the face and the superset is less than the radius R of the supersetst Add the radius R of the panel sl , that is: D <R st +R sl ;

[0085] If the new facet and all supersets do not meet the above conditions at the same time, and the category of the new facet is not background, a new superset is created based on the new facet. The creation steps here are the same as those for creating a new superset when facets are fused, so they will not be repeated. When the above conditions are met, the semantic confidence attribute of the superset is updated based on the facet.

[0086] Furthermore, after the face element update is completed, all supersets are updated. Specifically, the algorithm updates the center position, radius R, and st and category C st Attributes. For a superset, traverse all the face elements under it and record the maximum radius, the maximum value and the minimum value in the XYZ direction respectively. Use the collected data to update the superset. Specifically, update the center position of the superset as:

[0087]

[0088] Update the superset radius to:

[0089]

[0090] PC st is the confidence of a category among all categories; MaxX is the maximum value in the X direction, MinX is the minimum value in the X direction, and the same applies to MaxY, MinY, MaxZ and MinZ.

[0091] In the case of inconsistent semantic segmentation results for the same object during the mapping process, the semantics of the object facets constructed by the present invention will change, but the semantic information of the superset to which it belongs will not change suddenly. Moreover, the semantic information of the facets in the map is provided by the superset, thus maintaining the global semantic consistency of the object.

[0092] After obtaining the semantic map, superpixel segmentation is performed based on the color map, and then superpixels with semantic information are obtained by combining the semantic map. Then, data fusion semantic superpixels are extracted from the facet database and the superset database. Semantic facets are created based on the remaining superpixels. After creation, the newly created semantic facets are superset fused, and the unfused semantic facets are used to create supersets. Finally, the superset is updated based on the facet data under the superset, and a globally consistent semantic map is constructed based on the superset information. The complete algorithm process is as follows: Figure 1 shown.

[0093] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A semantic mapping method and system based on visual SLAM and YOLOv5, characterized by: The semantic mapping method comprises the following steps: Step 1: During the positioning process, line features are merged to improve the stability of line features. The depth of line feature endpoints is optimized through clustering and straight line fitting to solve the problem of missing line feature endpoint depth caused by depth map holes. The line feature matching method is improved to make line feature matching more efficient and faster; Step 2: After completing the stable matching, the algorithm will jointly optimize the camera pose with the point-line-plane error, adding weights to the point-line-plane error model to obtain a more accurate pose; Step 3: Select the key frame by judging the feature points of the current frame and the inter-frame pose. After confirming the key frame, pass the frame to the backend optimization and mapping thread; Step 4: In the map building thread, firstly, semantic segmentation is performed on the new keyframe to obtain semantic information and construct a semantic map to prepare for the subsequent semantic map building; Step 5: After obtaining the semantic map, superpixel segmentation is performed based on the color map, and then superpixels with semantic information are obtained in combination with the semantic map; Step 6: extract data from the facet database and the superset database to fuse semantic superpixels. Create semantic facets based on the remaining superpixels. After creation, perform superset fusion on the newly created semantic facets, and create supersets based on the semantic facets that have not been fused. Step 7: Update the superset according to the face metadata under the superset, and build a globally consistent semantic map based on the superset information.

2. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The process of merging line features described in step 1 is: The split short line features are merged and restored into long line features according to the direction constraint and the distance constraint, and after the line features are merged, the short line features with smaller length are filtered out.

3. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The process of solving the problem of missing depth at the endpoint of line features described in step 1 is: First, all the pixels with depth on the line feature are obtained by combining the depth map and projected to the camera coordinate system to obtain 3D points. Then, the 3D points are clustered to exclude noise points. Finally, Ransac line fitting is performed based on the clustering results and the depth of the line feature endpoint is obtained based on the fitted line.

4. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The improved line feature matching method described in step 1 includes: A matching method based on point-to-line distance constraints is used, and the LBD line feature descriptor that is easily affected by illumination is not used. The specific steps are as follows: nine points are uniformly sampled on the line feature to be matched, and the sampling points are back-projected onto the image plane to obtain the corresponding pixel coordinates. For each sampling point, the distance between the sampling point and the matching line segment is accumulated and divided by the number of sampling points to obtain the average distance, and the average distance is compared with a pre-set threshold. If the average distance is less than the threshold, the line feature is considered to be matched. This method achieves fast and stable line feature matching while ensuring matching accuracy.

5. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The method of combining point, line and surface errors described in step 2 includes: Combine the reprojection error constraints of feature points, the reprojection distance constraints of line feature endpoints, and the matching constraints of planes to jointly optimize the camera pose to obtain a more accurate pose. Consider assigning different weights to the constraints constructed by point, line, and surface features to maximize positioning accuracy while ensuring system stability. To ensure the real-time performance of the system, avoid complex weight settings.

6. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The construction of the semantic graph described in step 4 includes: The facet-based dense map construction method uses facets of different sizes to reconstruct the surface of real objects. In the constructed map, real objects can be represented as a set of single or multiple adjacent facets, and these facets should have consistent semantic labels. The present invention defines a superset based on this: a superset is a set composed of one or more facets; objects in the real world are represented by one or more adjacent supersets in the map; the facets in the superset have consistent semantic labels. Over-segmentation or under-segmentation during semantic segmentation will cause inconsistency in the semantic labels of individual facets in the map, but will not easily change the semantic label of the superset.

7. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The superpixel segmentation based on the color image described in step 5 includes: Different from creating a bin for each pixel, the present invention first performs superpixel segmentation on the color image and then constructs bins for the superpixels. This can not only greatly reduce the amount of system computation, but also reduce the impact of outliers and depth holes in low-quality depth images on image construction.

8. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The superset fusion described in step 6 includes: The semantic facet update process is an incremental map creation and update process, the purpose of which is to merge the local map created in this frame into the global map. The first step is to merge the facet and superpixel. If the conditions are met, the attributes of the facet are updated by combining the superpixel. If not, a new facet is created based on the superpixel. After the facet is created, super fusion is performed based on the relationship between the facet and the superset. Finally, the result of superset fusion is used to determine whether to update the superset with reverse mapping or to create a new superset. In the fusion of facets and superpixels, the physical distance and normal vector angle between the facet and the superpixel are used to determine whether to merge the two.

9. A semantic mapping method based on visual SLAM and YOLOv5 according to claim 1, characterized in that, The superset update described in step 7 includes: According to the attributes of the facets under the superset, the average geometric attributes are calculated and the superset attributes such as position, size and semantics are updated respectively.

10. A semantic mapping system based on visual SLAM and YOLOv5, applied to the method according to any one of claims 1 to 9, characterized in that: include: Line feature merging module: merges the detected short line features into long line features. Line feature endpoint depth optimization module: perform straight line fitting based on the points with depth on the line feature, and obtain the line feature endpoint depth through the fitted straight line. Line feature matching module: Determine the matching degree of line features by calculating the geometric constraints between line features. Weighted optimization module: Add weights to the optimization items involved in pose optimization, and the weights change dynamically. Globally consistent semantic mapping module: Constructs semantic maps through image data and pose data, and maintains semantic consistency during the mapping process.

11. A semantic mapping system based on visual SLAM and YOLOv5, applied to the method according to any one of claims 10, characterized in that: The globally consistent semantic mapping module includes: The first point is the superpixel segmentation unit: the input image data is first segmented into superpixels, and then semantic superpixels are given semantics to construct semantic superpixels; The second point is the surface element fusion unit: surface element fusion is created through superpixels, and the original surface data and superset data are updated at the same time. The third point is superset update unit: it updates superset information through the facets under the superset.

Citation Information

Patent Citations

  • Improved dynamic environment SLAM method based on YOLOv6 algorithm

    CN116678401A

  • Local refinement mapping system and method based on SLAM and semantic segmentation

    CN116772820A

  • Structured scene visual slam method based on point line surface features

    WO2023184968A1