Robot positioning recovery method and system based on multi-modal perception
By constructing a semantic map and multimodal perception fusion, using point cloud registration and dynamic environment index, the problem of robot positioning system automatically recovering positioning under dynamic obstacles is solved, achieving efficient and accurate positioning recovery.
Patent Information
- Application Number
- CN202510710739.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-29
AI Technical Summary
Existing robot positioning systems are prone to lose positioning in scenarios where there are many dynamic obstacles and cannot automatically restore positioning, and require manual or additional system intervention.
The semantic map is constructed, and RGB images and lidar data are fused using multimodal perception, and the optimal positioning is determined through point cloud registration and random sampling consistency algorithm, and combined with dynamic environmental index to detect positioning abnormalities to achieve automatic positioning recovery.
It realizes the robot's rapid and accurate positioning and recovery in dynamic obstacle environments, avoids manual intervention and misoperation, and improves the efficiency and accuracy of positioning and recovery.
Smart Images

Figure CN120385362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot positioning. More specifically, the present invention relates to a robot positioning recovery method and system based on multi-modal perception. Background Art
[0002] In robot positioning, existing methods generally implement it based on the simultaneous localization and mapping technology of 2D / 3D lidar. Through point cloud matching algorithms, such as ICP (Iterative Closest Point) and NDT (Normal Distributions Transform), pose estimation is achieved, such as systems like Google Cartographer and Hector SLAM. These systems rely on high-precision lidar and perform well in fixed-structure environments. However, in characteristic scenarios with many dynamic obstacles, long corridors, open areas, etc., cumulative errors are likely to occur. After positioning is lost, it cannot be automatically recovered.
[0003] Or a visual positioning method that uses a monocular / binocular camera for feature point tracking, such as systems like ORB-SLAM and VINS-Fusion. It calculates pose changes through feature point matching and SFM (Structure from Motion). However, this solution is sensitive to light changes and is prone to losing feature tracking in areas with a single texture (such as flat grassland), resulting in positioning failure.
[0004] To address the problems of low positioning accuracy, easy failure, or loss of a single sensor mentioned above, a multi-sensor fusion solution has been proposed in the prior art, such as the laser-vision tightly coupled method (LOAM series), or a filtering fusion solution that combines an IMU (Inertial Measurement Unit) and an odometer. However, these solutions mostly focus on improving positioning accuracy in conventional environments and lack an active detection and positioning recovery mechanism for abnormal positions. For example, when a robot is in a scenario with many dynamic obstacles, the existing system still continuously outputs incorrect positioning. After the dynamic obstacles (such as moving crowds) decrease, the positioning recovery mechanism cannot be automatically performed, and manual intervention is required to re-perform positioning.
[0005] Therefore, how to enable a robot to automatically perform positioning recovery without relying on manual or additional system intervention after positioning fails or is lost is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0006] To solve the technical problem that a robot cannot automatically perform positioning recovery after positioning fails, the present invention provides solutions in the following aspects.
[0007] In a first aspect, the present invention provides a robot positioning recovery method based on multi-modal perception, including: constructing a semantic map, which stores static obstacle information, and the static obstacle information includes a point cloud template and a semantic label; in response to detecting positioning failure, selecting at least a preset number of static obstacles with known semantic labels in the field of view; respectively performing point cloud registration on each static obstacle and the corresponding obstacle in the semantic map to obtain a plurality of point cloud registration results; determining an optimal pose from the poses generated by the plurality of point cloud registration results; and taking the optimal pose as the robot pose.
[0008] Further, after determining the robot pose, the method of the present invention further includes: calculating the matching degree between the local point cloud and the global point cloud according to the robot pose, and in response to the matching degree being greater than a first preset threshold, determining that the positioning recovery is successful; in response to the matching degree being less than or equal to the first preset threshold, moving in the direction of entropy reduction to obtain more static obstacles.
[0009] Further, determining an optimal pose from the poses generated by the plurality of point cloud registration results includes: generating candidate poses based on each point cloud registration result; and selecting, based on the random sample consensus algorithm, the candidate pose with the largest number of inliers among the candidate poses as the optimal pose.
[0010] Further, it further includes: determining the point cloud state consistency, the motion anomaly degree, and the feature tracking stability; wherein, the point cloud state consistency is a piecewise linear function of the point cloud registration score between the current frame point cloud and the point cloud in the semantic map; the motion anomaly degree is a piecewise linear function of the distance between the trajectories estimated by the IMU and the wheeled odometer and the SLAM trajectory; the feature tracking stability is positively correlated with the statistic of the tracking rate for a continuous first preset number of frames; performing a weighted sum on the point cloud state consistency, the motion anomaly degree, and the feature tracking stability to obtain a dynamic environment index; and determining whether there are dynamic environment indices less than a second preset threshold for a continuous second preset number of frames, and if so, determining that the positioning fails.
[0011] Further, the calculation expression of the point cloud state consistency is:
[0012]
[0013] In the formula, C is the point cloud state consistency, C is the point cloud registration score, C max is a third preset threshold, and min() is a minimum value function.
[0014] Further, the calculation expression of the motion anomaly degree is:
[0015]
[0016] Wherein, M is the abnormal motion degree, H is the distance, M max is the fourth preset threshold value, and min() is the minimum value function.
[0017] Further, the method for determining the feature tracking stability includes: statistically calculating the feature point tracking rate of each frame in a continuous first preset number of frames, where the feature point tracking rate is positively correlated with the number of feature points tracked in the current frame; determining the feature tracking stability based on the average value of the feature point tracking rates, and the feature tracking stability is positively correlated with the average value.
[0018] Further, the calculation expression of the feature tracking stability is:
[0019]
[0020] Wherein, T is the feature tracking stability, is the average value.
[0021] Further, performing point cloud registration on each static obstacle and the corresponding obstacle in the semantic map respectively includes: for any static obstacle, extracting the point cloud cluster of the static obstacle, and performing ICP registration on the point cloud cluster and the point cloud template of the corresponding obstacle in the semantic map.
[0022] In a second aspect, the present invention provides a robot positioning recovery system based on multi-modal perception, including a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a robot positioning recovery method according to any one of the first aspect is implemented.
[0023] The beneficial effects of the present invention are as follows: By fusing multi-modal data to establish a dynamic environment index to determine whether there is a positioning loss situation, the present invention can quickly and accurately identify whether the robot is in an abnormal positioning state, thereby improving the accuracy of robot abnormal state detection and reducing the false trigger rate; after detecting positioning failure or loss, based on the semantic map, pose estimation is performed according to the surrounding static obstacles, realizing automatic positioning recovery, while avoiding the interference of dynamic and temporary obstacles, as well as avoiding invalid positioning recovery operations and the intervention of additional systems or humans, thereby improving the accuracy and efficiency of positioning recovery; by generating multiple candidate poses and selecting the optimal pose from the candidate poses, it can tolerate a single obstacle being occluded or moved, and still stably output the pose in the case of partial feature loss, thereby avoiding the situation where positioning cannot be automatically recovered. Description of the Drawings
[0024] Figure 1It is a flowchart schematically showing a multi-modal perception-based robot positioning recovery method according to Embodiment 1 of the present invention;
[0025] Figure 2 It is a flowchart schematically showing a multi-modal perception-based robot positioning recovery method according to Embodiment 2 of the present invention;
[0026] Figure 3 It is a flowchart schematically showing abnormal positioning detection according to Embodiment 2 of the present invention;
[0027] Figure 4 It is a structural block diagram schematically showing a multi-modal perception-based robot positioning recovery system according to Embodiment 3 of the present invention. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0029] Next, the detailed implementation manners of the present invention will be described in detail with reference to the accompanying drawings.
[0030] Embodiment 1
[0031] Figure 1 It is a flowchart schematically showing a multi-modal perception-based robot positioning recovery method according to an embodiment of the present invention.
[0032] To solve the problem in the prior art that after detecting positioning failure, the positioning cannot be automatically and accurately restored, in Embodiment 1, the present invention provides a multi-modal perception-based robot positioning recovery method. As Figure 1 shown, the method of the present invention includes:
[0033] S101. Construct a semantic map.
[0034] Specifically, RGB images (a type of visible light image) are collected by a camera, laser point cloud data is collected by a lidar, and IMU data is collected by an IMU (Inertial Measurement Unit).
[0035] Using a large vision model (such as the SAM model, Segment Anything Model), perform instance segmentation on RGB images, identify static / fixed obstacle categories (such as fountains, statues, road signs, tree trunks, store signs, etc.), and extract the contour feature point set of the fixed obstacles. Use a point cloud clustering algorithm (such as DBSCAN, i.e., density-based spatial clustering algorithm) to fuse multiple frames of point clouds of the same fixed obstacle, and then calculate the three-dimensional centroid coordinates of the fixed obstacle.
[0036] Furthermore, construct a semantic feature layer in the map, and store static obstacle information in this semantic feature layer. In this embodiment, the static obstacle information includes obstacle type, geometric features (size, shape), semantic labels (such as "fountain-01"), and three-dimensional centroid coordinates, thereby obtaining a retrievable topological map (i.e., semantic map).
[0037] By constructing a semantic map, the robot can use only static semantic features (i.e., static obstacles) during localization recovery or anomaly detection, avoiding localization drift caused by dynamic interference, thereby improving the accuracy of anomaly detection and localization recovery.
[0038] S102. In response to detecting a positioning failure, perform positioning recovery based on the semantic map.
[0039] Specifically, when it is detected that a positioning failure occurs, stop moving and start a 360° environmental scan. Select at least a preset number (in this embodiment, set to at least 3) of static obstacles with known semantic labels in the field of view. For any one static obstacle, extract the point cloud cluster of this static obstacle, and perform point cloud registration with the point cloud template of the corresponding obstacle in the semantic map. Generate candidate poses based on the registration result.
[0040] In this embodiment, when performing point cloud registration, ICP registration can be used. In alternative other embodiments, those skilled in the art can select according to actual needs, for example, use the NDT algorithm for registration.
[0041] By selecting static obstacles for pose calculation, it is possible to avoid the interference of dynamic / temporary obstacles on pose estimation, thereby improving the accuracy of pose estimation. In addition, the traditional ICP global search is time-consuming, and the larger the map, the longer the time-consuming. The present invention reduces the search space to a local area (i.e., only select three static obstacles with known semantic labels in the field of view) through semantic screening, reduces the search time, and improves the efficiency of positioning recovery.
[0042] Furthermore, based on the Random Sample Consensus (RANSAC) algorithm, select the candidate pose with the largest number of inliers from multiple candidate poses as the optimal pose of the robot.
[0043] By generating multiple candidate poses and selecting the optimal pose from the candidate poses, it is possible to tolerate the occlusion or movement of a single obstacle and still stably output the pose even when some features are lost, thus avoiding the situation where positioning cannot be automatically restored.
[0044] By performing positioning based on the surrounding environmental features after detecting the failure or loss of positioning, the automatic positioning recovery of the robot is achieved, avoiding the problems of waiting for additional systems or manual intervention, as well as misoperations and ineffective positioning recovery operations, thereby improving the efficiency and accuracy of positioning recovery.
[0045] In one embodiment, the method of the present invention further includes: inputting the optimal pose into a SLAM (Simultaneous Localization and Mapping) system, calculating the matching degree between the local point cloud and the global point cloud. If the matching degree is greater than a first preset threshold, it is determined that the positioning recovery is successful; if the matching degree is less than or equal to the first preset threshold, indicating that the positioning recovery fails, then move along the entropy reduction direction to obtain more features, and then re-perform positioning recovery based on the obtained features (i.e., static obstacles). It should be noted that the entropy reduction direction is the direction in which more static obstacles can be obtained.
[0046] In this embodiment, the NDT score of the local point cloud and the global map can be used as the matching degree between the local point cloud and the global map.
[0047] By calculating the NDT score of the local point cloud and the global map to determine whether the positioning recovery is successful, the situation of mis-matched positioning is avoided, and the reliability of positioning recovery is further improved. After the positioning fails, by re-obtaining more static obstacles for positioning, it is ensured that the robot can accurately restore the positioning.
[0048] Embodiment Two
[0049] In order to solve the problems in the prior art that it is impossible to quickly and accurately detect the positioning failure and abnormality, and after detecting the positioning failure, it is impossible to automatically and accurately restore the positioning, that is, the lack of an active detection and recovery mechanism for abnormal positioning. In Embodiment Two, the present invention provides a robot positioning recovery method based on multi-modal perception. As Figure 2 shown, the method of the present invention includes:
[0050] S201. Construct a semantic map.
[0051] S202. Perform abnormal positioning detection based on the semantic map and real-time data.
[0052] Specifically, as Figure 3 shown, the abnormal positioning detection in this embodiment includes:
[0053] S2021. Determine the point cloud state consistency, motion anomaly degree, and feature tracking stability.
[0054] The point cloud state consistency is a piecewise linear function of the point cloud registration score between the point cloud of the current frame and the point cloud of the corresponding feature in the semantic map; the motion anomaly degree is a piecewise linear function of the distance between the trajectories estimated by the IMU and wheel odometer and the SLAM trajectory.
[0055] In one embodiment, the method for determining the point cloud state consistency includes: performing point cloud registration on the laser point cloud of the current frame and the visual semantic features (such as the point cloud of the statue contour) in the semantic map, then calculating the point cloud registration score, and determining the point cloud state consistency of the current frame based on this point cloud registration score. In this embodiment, the NTD score between the laser point cloud and the semantic map is used as the point cloud registration score between the laser point cloud and the semantic map, and the range of this NTD score is from 0 to 1. Specifically, the calculation expression for the point cloud state consistency is:
[0056]
[0057] In the formula, C norm is the point cloud state consistency, C NDT is the point cloud registration score, C max is the third preset threshold, and min() is the minimum value function.
[0058] Among them, C max is determined according to the actual application scenario. For example, C max is set to 0.7.
[0059] Through the point cloud state consistency, the difference degree between the visual data and the lidar data can be clarified. It can be understood that the lower the point cloud state consistency, the higher the possibility of positioning anomalies and failures.
[0060] In one embodiment, the method for determining the motion anomaly degree includes: obtaining the trajectories estimated by the IMU and wheel odometer and the SLAM trajectory of the most recent N frames (such as 10 frames), after performing time synchronization alignment, calculating the distance between the trajectories estimated by the IMU and wheel odometer and the SLAM trajectory, and calculating the motion anomaly degree based on this distance. Specifically, the calculation expression for the motion anomaly degree is:
[0061]
[0062] In the formula, M is the motion anomaly degree, H is the distance, M max is the fourth preset threshold, and min() is the minimum value function.
[0063] In this embodiment, the distance between the IMU and wheel odometer estimated trajectories and the SLAM trajectory can be calculated by the bidirectional Hausdorff distance, and M max Different values can be set according to the outdoor scene and the outdoor scene. For example, it can be set to 0.7 outdoors and 0.6 indoors.
[0064] From the calculation expression of the motion abnormality degree, it can be seen that restricting the point cloud state consistency within the range of 0 to 1 can avoid the situation where the distance between the IMU and wheel odometer estimated trajectories and the SLAM trajectory is too large, resulting in abnormal calculation of the dynamic environment index (if it is not restricted within 0 to 1, misjudgment will occur), ensuring the reliability of positioning anomaly detection.
[0065] In one embodiment, the method for determining the feature tracking stability includes: First, count the feature point tracking rate of each frame in a continuous first preset number of frames (in this embodiment, set to 5 frames); Second, determine the feature tracking stability based on the average value of the feature point tracking rates of the first preset number of feature points, where the feature tracking stability is positively correlated with the average value. Specifically, the calculation expression of the feature point tracking rate is:
[0066]
[0067] In the formula, L i is the feature point tracking rate of the current frame, N init is the total number of initial feature points, N taracked is the number of feature points tracked in the current frame.
[0068] The calculation expression of the feature tracking stability is:
[0069]
[0070] In the formula, T is the feature tracking stability, is the average value of the feature point tracking rates.
[0071] By calculating the feature tracking stability in segments, the reliability of obtaining the feature tracking stability in different dynamic environments can be improved, thereby improving the reliability of subsequent positioning anomaly detection. In other optional embodiments, those skilled in the art can set the segmentation points of the piecewise function according to actual needs.
[0072] S2022. Perform weighted summation on the point cloud state consistency, motion abnormality degree, and feature tracking stability to obtain a dynamic environment index.
[0073] In one embodiment, the calculation expression of the dynamic environment index is:
[0074] EDI = αC + βM + γT;
[0075] In the formula, EDI is the dynamic environment index, C is the consistency of the point cloud state, M is the degree of motion anomaly, T is the stability of feature tracking, α is the weight of the consistency of the point cloud state, β is the weight of the degree of motion anomaly, and γ is the weight of the stability of feature tracking.
[0076] In this embodiment, the values of α, β, and γ (α + β + γ = 1) can be set according to the actual scenario. Specifically, in a dynamic scenario, the value of α can be increased; in a low-texture scenario (such as grassland, etc.), the value of γ can be increased.
[0077] According to the calculation expression of the dynamic environment index, the higher the consistency of the point cloud state, the higher the dynamic environment index, indicating that the probability / possibility of positioning failure is lower; the higher the degree of motion anomaly, the higher the dynamic environment index, indicating that the probability of positioning failure is lower; the higher the stability of feature tracking, the higher the dynamic environment index, indicating that the probability of positioning failure is lower.
[0078] S2023. Determine whether there is a positioning failure according to the magnitude relationship between the dynamic environment index and the second preset threshold.
[0079] Specifically, it is determined whether there are consecutive second preset frames (in this embodiment, set to 3 frames) whose dynamic environment indices are all less than the second preset threshold (in this embodiment, set to 0.5). If so, it is determined that there is a positioning failure; if not, it indicates that there is no positioning anomaly. In alternative other embodiments, those skilled in the art can set the second preset threshold according to actual needs, for example, set to 0.4.
[0080] By quantifying the consistency of the point cloud state, the degree of motion anomaly, the stability of feature tracking, and the dynamic environment index, a reliable basis is provided for positioning anomaly detection, so that it is possible to quickly and accurately detect whether the robot has a positioning anomaly or failure; by determining that there is a positioning failure only when the dynamic environment indices of multiple consecutive frames are all less than the second preset threshold, the probability of misjudgment can be reduced, thereby improving the reliability of positioning failure detection.
[0081] S203. In response to detecting a positioning failure, perform positioning recovery based on the semantic map.
[0082] It should be noted that S101 in Embodiment 1 is the same as S201 in Embodiment 2, and S102 in Embodiment 1 is the same as S203 in Embodiment 2. Therefore, S201 and S203 are not described in detail in Embodiment 2.
[0083] Embodiment 3
[0084] Figure 4 It schematically shows the structural block diagram of the robot positioning recovery system based on multi-modal perception in this embodiment.
[0085] In a second aspect, the present invention further provides a robot positioning recovery system based on multi-modal perception. As Figure 4 shown, the robot positioning recovery system based on multi-modal perception includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a robot positioning recovery method based on multi-modal perception according to the first aspect of the present invention is implemented.
[0086] The robot positioning recovery system based on multi-modal perception further includes other components well-known to those skilled in the art, such as a communication interface. Its settings and functions are known in the art, and thus will not be described in detail herein.
[0087] In the present invention, the foregoing memory may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as, for example, a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium may be part of the device or accessible or connectable to the device. Any application or module described in the present invention may be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.
[0088] In the description of this specification, "a plurality of" means at least two, for example, two, three, or more, etc., unless otherwise specifically defined.
[0089] Although this specification has shown and described several embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications, and alternative forms will occur to those skilled in the art without departing from the spirit and scope of the present invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.
Claims
1. A robot positioning recovery method based on multi-modal perception, characterized in that, Including: Construct a semantic map, which stores static obstacle information, and the static obstacle information includes a point cloud template and a semantic label; In response to detecting a positioning failure, select at least a preset number of static obstacles with known semantic labels in the field of view; Perform point cloud registration on each static obstacle and the corresponding obstacle in the semantic map respectively to obtain multiple point cloud registration results; Determine the optimal pose from the poses generated by the multiple point cloud registration results; Use the optimal pose as the robot pose.
2. The method for robot positioning recovery based on multi-modal perception according to claim 1, characterized in that, After determining the robot pose, it further includes: calculating the matching degree between the local point cloud and the global point cloud according to the robot pose. In response to the matching degree being greater than a first preset threshold, it is determined that the positioning recovery is successful; in response to the matching degree being less than or equal to the first preset threshold, move along the entropy reduction direction to obtain more static obstacles.
3. The method for robot positioning recovery based on multi-modal perception according to claim 1, wherein Determining the optimal pose from the poses generated by the multiple point cloud registration results includes: Generating candidate poses based on each point cloud registration result; Selecting the candidate pose with the largest number of inliers from the candidate poses based on the random sample consensus algorithm as the optimal pose.
4. The method for robot positioning recovery based on multi-modal perception according to claim 1, wherein, It also includes: Determine the point cloud state consistency, motion anomaly degree, and feature tracking stability; wherein, the point cloud state consistency is a piecewise linear function of the point cloud registration score between the current frame point cloud and the point cloud in the semantic map; the motion anomaly degree is a piecewise linear function of the distance between the trajectories estimated by the IMU and the wheel odometer and the SLAM trajectory; the feature tracking stability is positively correlated with the statistic of the tracking rate for a continuous first preset number of frames; Perform a weighted sum on the point cloud state consistency, motion anomaly degree, and feature tracking stability to obtain a dynamic environment index; Judge whether there are dynamic environment indices less than a second preset threshold for a continuous second preset number of frames. If so, it is determined that the positioning fails.
5. The method for robot positioning recovery based on multi-modal perception according to claim 4, wherein The calculation expression of the point cloud state consistency is: Where C is the point cloud state consistency, C NDT is the point cloud registration score, C max is the third preset threshold, and min() is the minimum value function.
6. The method for robot positioning recovery based on multi-modal perception according to claim 4, wherein The calculation expression of the motion anomaly degree is: Wherein, M is the abnormal motion degree, H is the distance, M max is the fourth preset threshold, and min() is the minimum value function.
7. The method for robot positioning recovery based on multi-modal perception according to claim 4, characterized in that The method for determining the feature tracking stability includes: Statistically calculate the feature point tracking rate for each frame in a continuous first preset number of frames, and the feature point tracking rate is positively correlated with the number of feature points tracked in the current frame; Determine the feature tracking stability based on the average value of the feature point tracking rate, and the feature tracking stability is positively correlated with the average value.
8. The method for robot positioning recovery based on multi-modal perception according to claim 7, characterized in that, The calculation expression of the feature tracking stability is: Where T is the feature tracking stability, is the average value.
9. The method for robot positioning recovery based on multi-modal perception according to claim 1, characterized in that, Performing point cloud registration on each static obstacle and the corresponding obstacle in the semantic map respectively includes: for any one static obstacle, extracting the point cloud cluster of the static obstacle and performing ICP registration on the point cloud cluster and the point cloud template of the corresponding obstacle in the semantic map.
10. A robot positioning recovery system based on multi-modal perception, characterized in that, Including a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, it realizes a robot positioning recovery method according to any one of claims 1-9.