Loop Detection Method, Device, Readable Storage Medium, and Removable Device

By employing sliding window techniques to compare multiple frames of visual and point cloud data, the method enhances the robustness and accuracy of loop closure detection in SLAM systems, reducing error rates and ensuring consistent map generation.

CN114494422BActive Publication Date: 2025-07-15MIDEA ROBOZONE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210127114.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-07-15
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

The existing loopback detection methods mainly rely on the comparison of current frame data and historical data, and are relatively low in robustness, resulting in a high loopback detection error rate.

Method used

Sliding window technology is adopted to compare the results after the fusion of continuous multi-frame data with historical data, combine visual data and point cloud data, and select sliding windows using the unchanging characteristics of the categories and quantity of objects, and morphology and graphics technology are used to improve detection speed and robustness.

Benefits of technology

It significantly improves the robustness of loopback detection, reduces the chance of error detection, and ensures the accuracy and consistency of the map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494422B_ABST
    Figure CN114494422B_ABST
Patent Text Reader

Abstract

The present invention provides a loop detection method, apparatus, readable storage medium and mobile device. The method includes: periodically acquiring image data in the field of view; determining a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of image data; and determining that there is a loop in the first sliding window and the second sliding window when the first sliding window and the second sliding window meet a first condition. By running this detection method, the robustness can be greatly improved and the probability of loop detection errors can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of control technologies, and more particularly, to a loop detection method, apparatus, readable storage medium, and mobile device. Background Art

[0002] Existing current loop detections all compare the current frame of data with historical data, resulting in low robustness and a high probability of loop detection errors. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art or related technologies.

[0004] To this end, in the first aspect of the present invention, there is provided a loop detection method.

[0005] In the second aspect of the present invention, there is provided a loop detection apparatus.

[0006] In the third aspect of the present invention, there is provided a loop detection apparatus.

[0007] In the fourth aspect of the present invention, there is provided a readable storage medium.

[0008] In the fifth aspect of the present invention, there is provided a mobile device.

[0009] In view of this, according to the first aspect of the present invention, there is provided a loop detection method, including: periodically acquiring image data in the field of view; determining a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of image data; and determining that there is a loop between the first sliding window and the second sliding window when the first sliding window and the second sliding window meet a first condition.

[0010] The technical solution of the present application proposes a loop detection method. By running this detection method, the robustness can be greatly improved, and the probability of loop detection errors can be reduced.

[0011] The technical solution of the present application is implemented based on the following principle. Specifically:

[0012] SLAM (Simultaneous Localization and Mapping), also known as CML (Concurrent Mapping and Localization), is simultaneous localization and mapping, or concurrent mapping and localization. The problem can be described as follows: If a robot is placed at an unknown location in an unknown environment, is there a way for the robot to gradually map the entire environment while moving? A complete map (a consistent map) means that the robot can move through the room without obstacles and reach every accessible corner.

[0013] Loop detection helps reduce the cumulative error of robot pose estimation and generate a consistent global map. In the past decade, the latest loop detection algorithms for visual and lidar SLAM mainly fall into three categories: vision-based loop detection methods, lidar-based loop detection methods, and deep learning-based loop detection methods. These methods currently compare the current single-frame data with historical data, and their robustness is not strong.

[0014] In the technical solution of this application, sliding windows are compared with each other. Since there is more than one frame of image data in the sliding window, the result of fusing multiple consecutive frames of data can be used to compare with historical data, thereby greatly improving the robustness and reducing false loop detections.

[0015] In one possible design, the field of view involved in this application can be understood as the image captured by the visual sensor.

[0016] In the above technical solution, due to the cumulative error, the map will eventually drift. A loop can be understood as pulling the position on the map with cumulative error back to the correct position when the same position is detected, thereby eliminating the cumulative error.

[0017] In addition, the loop detection method proposed in this application also has the following additional technical features.

[0018] In the above technical solution, the image data includes visual data and point cloud data. Determining the first sliding window and the second sliding window specifically includes: determining the objects in the field of view according to the visual data; determining the bounding boxes and categories of the objects; annotating the point cloud data according to the bounding boxes and categories of the objects; mapping the annotated point cloud data to a two-dimensional grid; determining that when the category and quantity of the objects remain unchanged within the first number of frames during the mapping process, the image data of the first number of frames is recorded as a sliding window; or determining that when the cumulative number of frames of the image data is greater than or equal to the second number of frames during the mapping process, the image data of the second number of frames is recorded as a sliding window.

[0019] In this technical solution, visual data can be acquired using a visual sensor, and point cloud data can be obtained using a TOF (Time of Flight) point cloud module. Among them, the form of the point cloud data can be in the XYZ form.

[0020] By marking a bounding box for the object, when annotating the point cloud data, the accuracy of the annotation can be ensured. At the same time, when mapping the point cloud data to a two-dimensional grid, the contour of the object in the two-dimensional grid can be determined, so as to use the contour of the object in the two-dimensional grid to achieve loop detection. In the above solution, morphological and graphics technologies are combined, which improves the speed of loop detection and the robustness of loop detection.

[0021] In the above technical solution, the selection scheme of the sliding window is also specifically defined. In the above solution, the selection of the sliding window is realized by using the characteristics that the category and quantity of the object no longer change, so as to minimize the size of the sliding window to the greatest extent, thereby reducing the computational complexity of loop detection.

[0022] In the above technical solution, it can be understood that in the continuous first frame number of image data, the determined category no longer changes, and at the same time, the quantity of the object also does not change, and it is considered that the category and quantity of the object no longer change.

[0023] In one possible design, the quantity of the object can be the quantity of any category of object or the quantity of all categories of objects.

[0024] In the above technical solution, the quantity of the object can be obtained by statistics during the process of determining the category of the object. Specifically, the center point and direction radius of the point cloud data of each category in the two-dimensional plane are calculated. Among them, the center point can be determined by using the depth value in the point cloud data. Specifically, the depth values are sorted, and the median point is selected as the center point of the object. At the same time, assuming that the point cloud conforms to a Gaussian distribution, we calculate the standard deviation in two directions of the two-dimensional plane as the corresponding direction radius respectively. In this way, the estimated center point and radius of each category are obtained.

[0025] For example, suppose there are two categories, chairs and trash cans. When a chair is detected in the first frame, a 1 is placed in the second bit of the binary sequence, representing 10, and the number of different chairs is recorded. As mentioned above, we already have the center point and radius of the current chair. In the second frame, if it is the same chair, we consider that the rectangle formed by the estimated center point and radius has a large intersection with that of the first frame, and the number of chairs is still 1 at this time. At the same time, the radius and center point position of the chair are updated according to the principle of multiplying Gaussian distributions. If the second frame is another chair, then the rectangle formed by the estimated center point and radius has a small intersection with that of the first frame. We increment the number of chairs by 1 and record it as 2.

[0026] It should be noted that the two-dimensional grid is updated not based on the estimated center point and radius of the object, but based on the projection of the actual point cloud belonging to the object category on the two-dimensional plane, so as to reflect the actual contour of the object. If a trash can is detected, then a 1 is placed in the first position of the binary sequence and the quantity is recorded.

[0027] The method of determining the sliding window by whether the cumulative number of frames of the image data is greater than the second number of frames ensures the integrity of the sliding window to the greatest extent.

[0028] In the above technical solution, the first number of frames is less than or equal to the second number of frames.

[0029] In the above technical solution, after the first sliding window, the same solution is used to determine the second sliding window.

[0030] In any of the above technical solutions, the first condition includes at least one of the following: the Euclidean distance between the binary sequences of the first sliding window and the second sliding window is less than the first value, the ratio of the first quantity of the objects of the corresponding category in the second sliding window to the second quantity of the objects of the corresponding category in the first sliding window is greater than or equal to the second value, and the distances between the corresponding same or different objects are the same.

[0031] In this technical solution, the possible contents included in the first condition are specifically given. In this solution, when determining the category of an object, statistics are performed on the first sliding window and the second sliding window, and the above statistical results are reflected in the form of a binary sequence. For example, if the object corresponds to the 12th bit in the binary sequence, if the object is detected, the 12th bit in the binary sequence is set to 1, otherwise, it is set to 0.

[0032] By calculating the Euclidean distance between the first sliding window and the second sliding window, the similarity degree between the first sliding window and the second sliding window can be calculated. When the Euclidean distance is less than the first value, it is considered that the similarity degree between the first sliding window and the second sliding window is relatively high.

[0033] Similarly, two other schemes can also be adopted for detection. Among them, the use of the above three schemes can be selected according to the actual usage scenario.

[0034] In the above technical solution, the distances between corresponding identical or different objects are the same. It can be understood that the distance between identical objects in the first sliding window is the same as the distance between the corresponding identical objects in the second sliding window, and the distance between different objects in the first sliding window is the same as the distance between the corresponding different objects in the second sliding window.

[0035] For example, if the distance between the centers of the chair and the trash can in sliding window A is 1 meter, then the distance between the centers of the chair and the trash can in sliding window B should also be around 1 meter.

[0036] In the above technical solution, when the first condition is met, it is considered that there is a loop between the first sliding window and the second sliding window.

[0037] In any of the above technical solutions, it further includes: obtaining first image data in the second sliding window, where the first image data is the image data closest to the current moment in the second sliding window; determining the target object included in the first image data; determining the pose candidate set including the target object; determining the target image data in the pose candidate set that matches the first image data; and using the pose data corresponding to the target image data as the pose data of the first image data.

[0038] In this technical solution, a pose determination scheme based on loop detection is given. In this scheme, by searching for the target image data that matches the first image data and using the pose data corresponding to the target image data as the pose data corresponding to the first image data, the determination of the pose can be realized.

[0039] In the above technical solution, the pose candidate set is determined by using the target object, so as to reduce the screening range of the target image data, thereby reducing the amount of data searched during the determination of the pose data, and improving the search speed of the pose.

[0040] In the above technical solution, since the target object is determined based on the first image data, the matching degree between the pose candidate set and the first image data can be ensured, thereby improving the accuracy of the determination of the pose data.

[0041] In the above technical solution, it can be understood that the first image data is the latest frame of image data in the second sliding window.

[0042] In the above technical solution, artificial intelligence can be used to recognize the first image data, so as to determine the target object contained in the first image data.

[0043] In the above technical solution, the number of target objects can be one or multiple.

[0044] In the above technical solution, the pose candidate set can be indexed by the object, so that after the target object is determined, the pose candidate set can be determined using the index.

[0045] In any of the above technical solutions, it further includes: determining the offset of the reference object and the first rotation amount corresponding to the line connecting the reference object and the first non-reference object under the two-dimensional grid corresponding to the first sliding window and the second sliding window; determining the second rotation amount corresponding to the line connecting the reference object and the second non-reference object; and storing the offset and the first rotation amount in the pose candidate set when the first rotation amount and the second rotation amount meet the second condition.

[0046] In the above technical solution, the reference object can be understood as the selected reference point, and the first non-reference object and the second non-reference object are other objects in the objects except the reference object.

[0047] In this technical solution, the construction scheme of the pose candidate set is specifically defined. In this scheme, the offset, the first rotation amount, and the second rotation amount are introduced, and the pose candidate set is determined using the offset, the first rotation amount, and the second rotation amount. In this process, since the first rotation amount and the second rotation amount meet the second condition, that is, not all calculated offsets and first rotation amounts can be stored in the pose candidate set, the number in the pose candidate set can be reduced to improve the search speed.

[0048] In the above technical solution, the offset of the reference object can be understood as follows: the coordinate obtained by mapping the reference object in the first sliding window to the two-dimensional grid is the first coordinate, and the coordinate obtained by mapping the reference object in the second sliding window to the two-dimensional grid is the second coordinate. Among them, the deviation between the first coordinate and the second coordinate is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0049] In the above technical solution, the first rotation amount can be understood as the angle value between the first line connecting the reference object and the first non-reference object in the first sliding window and the second line connecting the reference object and the first non-reference object in the second sliding window, when rotating from the first line to the second line.

[0050] Similarly, the second rotation amount can be understood as the angle value between the third line connecting the reference object and the second non-reference object in the first sliding window and the fourth line connecting the reference object and the second non-reference object in the second sliding window, when rotating from the third line to the fourth line.

[0051] In the above technical solution, the second condition may be that the difference between the second rotation amount and the first rotation amount is less than a preset angle value, or that the ratio between the second rotation amount and the first rotation amount is within a preset numerical range, such as between 0.8 and 1.2.

[0052] For example, assume that there is chair, trash can, and shoe information in windows A and B. Then align the center points of the chairs in window B with the center points of the chairs in A to calculate the offset amounts DX and DY between the two. Then calculate the angle value between the line connecting the trash can and the chair in window A and the line connecting the trash can and the chair in window B. Consider the rotational offset amount as D_theta. We further detect whether the shoe offset amount is also near D_theta. Through such a search, we finally find the candidate pose candidate set {DX, DY, D_theta} that meets the conditions.

[0053] In any of the above technical solutions, it further includes: determining the first two-dimensional grid corresponding to the second sliding window according to the candidate pose set; determining the second two-dimensional grid of the first sliding window; determining the intersection and union of the first two-dimensional grid and the second two-dimensional grid; determining the area coincidence degree according to the intersection and union; and screening the candidate pose set according to the comparison result between the area coincidence degree and the preset area coincidence degree.

[0054] In this technical solution, considering that the candidate pose set is still relatively large and the calculation speed is still relatively slow, the technical solution of the present application introduces the area coincidence degree and uses the area coincidence degree to screen the candidate pose set in order to delete the candidate pose set, thereby improving the calculation speed.

[0055] Specifically, in the case where the calculated area coincidence degree is small, the corresponding offset amount and the first rotation amount are deleted, while in the case where the calculated area coincidence degree is large, the corresponding offset amount and the first rotation amount are retained.

[0056] In the above technical solution, a calculation scheme for the area coincidence degree is also given. In this scheme, by counting the intersection and union between the first two-dimensional grid and the second two-dimensional grid, the ratio of the intersection to the union is used as the area coincidence degree.

[0057] In the above technical solution, the intersection can be understood as the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid, and the union is the difference between the total number of grids of the first two-dimensional grid and the second two-dimensional grid and the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid.

[0058] In any of the above technical solutions, the area coincidence degree is calculated by using the scan line and segment tree method.

[0059] Perform pose transformation on the grid map formed by trash cans, chairs, and shoes in the second sliding window according to DX, DY, and D_theta. Then match it with the grid map formed by trash cans, chairs, and shoes in the first sliding window, and use scan lines and segment trees to calculate the area overlap to improve the calculation speed.

[0060] In the above technical solution, the scan line can be understood as scanning in a way that a line is translated as a whole in the two-dimensional grid, and the segment tree is a binary search tree, similar to the interval tree. It divides an interval into some unit intervals, and each unit interval corresponds to a leaf node in the segment tree. Using the segment tree can quickly find the number of times a certain node appears in several line segments.

[0061] In any of the above technical solutions, the Cartographer algorithm is used to match the target image data.

[0062] In this technical solution, using the Cartographer algorithm can improve the search speed within the offset and rotation ranges.

[0063] In any of the above technical solutions, it further includes: obtaining the center points of the reference object, the first non-reference object, and the second non-reference object; determining the offset, the first rotation amount, and the second rotation amount according to the center points.

[0064] In this technical solution, using the center points to determine the offset, the first rotation amount, and the second rotation amount, it can be understood that the offset of the reference object can be understood as follows: the coordinate of the center point of the reference object in the first sliding window mapped to the two-dimensional grid is the first coordinate, and the coordinate of the center point of the reference object in the second sliding window mapped to the two-dimensional grid is the second coordinate. Among them, the deviation between the first coordinate and the second coordinate is the offset, denoted as DX and DY. Among them, DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0065] In the above technical solution, the first rotation amount can be understood as the angle value between the first connection line between the center points of the reference object and the first non-reference object in the first sliding window and the second connection line between the center points of the reference object and the first non-reference object in the second sliding window when the first connection line rotates to the second connection line.

[0066] Similarly, the second rotation amount can be understood as the angle value between the third connection line between the center points of the reference object and the second non-reference object in the first sliding window and the fourth connection line between the center points of the reference object and the second non-reference object in the second sliding window when the third connection line rotates to the fourth connection line.

[0067] In any of the above technical solutions, determining an object in the field of view based on visual data includes: identifying the visual data based on a neural network model to obtain the object in the field of view.

[0068] In this technical solution, a neural network model is used to identify the object, ensuring the credibility of the identified object.

[0069] According to the second aspect of the present invention, the present invention provides a loop detection device, including: an acquisition unit configured to periodically acquire image data in the field of view; a determination unit configured to determine a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of image data; and a judgment unit configured to determine that there is a loop between the first sliding window and the second sliding window when the first sliding window and the second sliding window meet a first condition.

[0070] The technical solution of the present application proposes a loop detection device. A mobile device with this detection device can greatly improve the robustness and reduce the probability of incorrect loop detection.

[0071] The technical solution of the present application is implemented based on the following principle. Specifically:

[0072] SLAM (simultaneous localization and mapping), also known as CML (Concurrent Mapping and Localization), is simultaneous localization and mapping, or concurrent mapping and localization. The problem can be described as: placing a robot in an unknown position in an unknown environment, is there a way for the robot to gradually draw a complete map of this environment while moving? A complete map (a consistent map) means traveling to every accessible corner of the room without obstacles.

[0073] Loop detection helps reduce the cumulative error of robot pose estimation and generate a consistent global map. In the past decade, the latest loop detection algorithms in visual and lidar SLAM mainly fall into three categories: vision-based loop detection methods, lidar-based loop detection methods, and deep learning-based loop detection methods. These methods currently compare the current single-frame data with historical data and are not very robust.

[0074] In the technical solution of the present application, by comparing a sliding window with another sliding window, since there is more than one frame of image data in the sliding window, the result after fusing multiple consecutive frames of data can be used to compare with historical data, thereby greatly improving the robustness and reducing incorrect loop detection.

[0075] In one possible design, the field of view involved in the present application can be understood as the image captured by a visual sensor.

[0076] In the above technical solution, due to the cumulative error, the map will finally drift. The loop closure can be understood as, when it is detected that the same position has been experienced, pulling the position on the map with the cumulative error back to the correct position, thereby eliminating the cumulative error.

[0077] In addition, the loop closure detection device proposed in the present application further has the following additional technical features.

[0078] In the above technical solution, the determination unit is specifically configured to: determine the objects in the field of view according to the visual data; determine the bounding boxes and categories of the objects; annotate the point cloud data according to the bounding boxes and categories of the objects; map the annotated point cloud data to a two-dimensional grid; determine that when the category and quantity of the objects remain unchanged within the first number of frames during the mapping process, record the image data of the first number of frames as a sliding window; or determine that when the cumulative number of frames of the image data is greater than or equal to the second number of frames during the mapping process, record the image data of the second number of frames as a sliding window.

[0079] In this technical solution, the visual data can be obtained by using a visual sensor, and the point cloud data can be obtained by using a TOF point cloud module. Among them, the TOF (Time of Flight) point cloud module is used for acquisition. Among them, the form of the point cloud data can be in the XYZ form.

[0080] By marking the bounding boxes for the objects, when annotating the point cloud data, the accuracy of the annotation can be ensured. At the same time, during the process of mapping the point cloud data to the two-dimensional grid, the contour of the object in the two-dimensional grid can be determined, so as to use the contour of the object in the two-dimensional grid to implement loop closure detection. In the above solution, morphological and graphics technologies are combined, which improves the speed of loop closure detection and at the same time improves the robustness of loop closure detection.

[0081] In the above technical solution, the selection scheme of the sliding window is also specifically defined. In the above solution, by using the characteristics that the category and quantity of the objects no longer change, the selection of the sliding window is realized, so as to minimize the size of the sliding window to the greatest extent, thereby reducing the computational amount of loop closure detection.

[0082] In the above technical solution, it can be understood that in the continuous image data of the first number of frames, when it is determined that the obtained category no longer changes and at the same time the quantity of the objects does not change, it is considered that the category and quantity of the objects no longer change.

[0083] In one possible design, the number of objects can be the number of objects of any category or the number of objects of all categories.

[0084] In the above technical solution, the number of objects can be statistically obtained during the process of determining the category of the objects. Specifically, the center point and the direction radius of the point cloud data of each category in the two-dimensional plane are calculated. Among them, the center point can be determined by using the depth value in the point cloud data. Specifically, the depth values are sorted, and the median point is selected as the center point of the object. At the same time, assuming that the point cloud conforms to the Gaussian distribution, we calculate the standard deviations in two directions of the two-dimensional plane as the corresponding direction radii respectively. In this way, the estimated center point and radius of each category are obtained.

[0085] For example, suppose there are two categories, chairs and trash cans. When a chair is found in the first frame, set 1 at the second bit of the binary sequence, representing 10, and record the number of different chairs at the same time. Above, we already have the center point and radius of the current chair. In the second frame, if it is still the same chair, we consider that the rectangle formed by the estimated center point and radius has a large intersection with that of the first frame. At this time, the number of chairs is still 1. At the same time, update the radius and the position of the center point of the chair according to the multiplication principle of the Gaussian distribution. If the second frame is another chair, then the rectangle formed by the estimated center point and radius has a small intersection with that of the first frame. We add 1 to the number of chairs and record it as 2.

[0086] It should be noted that the two-dimensional grid is updated not according to the estimated center point and radius of the object, but according to the projection of the actual point cloud belonging to the object category in the two-dimensional plane, so as to reflect the actual contour of the object. If a trash can is found, then set 1 at the first position of the binary sequence and record the quantity at the same time.

[0087] The scheme of determining the sliding window by the way of whether the cumulative number of frames of the image data is greater than the second number of frames ensures the integrity of the sliding window to the greatest extent.

[0088] In the above technical solution, the first number of frames is less than or equal to the second number of frames.

[0089] In the above technical solution, after the first sliding window, the same scheme is used to determine the second sliding window.

[0090] In any of the above technical solutions, the first condition includes at least one of the following: the Euclidean distance between the binary sequences of the first sliding window and the second sliding window is less than the first value, the ratio of the first quantity of the objects of the corresponding category in the second sliding window to the second quantity of the objects of the corresponding category in the first sliding window is greater than or equal to the second value, and the distances between the same or different objects are the same.

[0091] In this technical solution, the possible contents included in the first condition are specifically given. In this solution, when determining the category of an object, the first sliding window and the second sliding window are statistically analyzed, and the above statistical results are reflected in the form of a binary sequence. For example, if the object corresponds to the 12th bit in the binary sequence, if the object is detected, the 12th bit in the binary sequence is set to 1, and vice versa, it is set to 0.

[0092] By calculating the Euclidean distance between the first sliding window and the second sliding window, the similarity degree between the first sliding window and the second sliding window is calculated. When the Euclidean distance is less than the first value, it is considered that the similarity degree between the first sliding window and the second sliding window is relatively high.

[0093] Similarly, two other schemes can also be used for detection. Among them, the above three schemes can be selected according to the actual usage scenario.

[0094] In the above technical solution, the distances between the same or different objects are the same. It can be understood that the distances between the same objects in the first sliding window are the same as the distances between the corresponding same objects in the second sliding window, and the distances between different objects in the first sliding window are the same as the distances between the corresponding different objects in the second sliding window.

[0095] For example, if the central distance between the chair and the trash can in sliding window A is 1 meter, then the central distance between the chair and the trash can in sliding window B should also be around 1 meter.

[0096] In the above technical solution, when the first condition is met, it is considered that there is a loop between the first sliding window and the second sliding window.

[0097] In any of the above technical solutions, the judgment unit is further configured to: obtain the first image data in the second sliding window, where the first image data is the image data closest to the current moment in the second sliding window; determine the target object included in the first image data; determine the pose candidate set including the target object; determine the target image data in the pose candidate set that matches the first image data; and use the pose data corresponding to the target image data as the pose data of the first image data.

[0098] In this technical solution, a pose determination scheme based on loop detection is given. In this scheme, by searching for the target image data that matches the first image data, the pose data corresponding to the target image data is used as the pose data corresponding to the first image data to achieve the determination of the pose.

[0099] In the above technical solution, a pose candidate set is determined using the target object, so as to reduce the screening range of the target image data, thereby reducing the amount of data searched during the pose data determination process, and thus improving the pose search speed.

[0100] In the above technical solution, since the target object is determined based on the first image data, it is possible to ensure the matching degree between the pose candidate set and the first image data, thereby improving the accuracy of pose data determination.

[0101] In the above technical solution, it can be understood that the first image data is the latest frame of image data in the second sliding window.

[0102] In the above technical solution, artificial intelligence can be used to identify the first image data, so as to determine the target object contained in the first image data.

[0103] In the above technical solution, the number of target objects can be one or multiple.

[0104] In the above technical solution, the pose candidate set can be indexed by the object, so that after the target object is determined, the index is used to determine the pose candidate set.

[0105] In any of the above technical solutions, the judgment unit is specifically further configured to: determine the offset of the reference object and the first rotation amount corresponding to the line connecting the reference object and the first non-reference object under the two-dimensional grid corresponding to the first sliding window and the second sliding window; determine the second rotation amount corresponding to the line connecting the reference object and the second non-reference object; and store the offset and the first rotation amount in the pose candidate set when the first rotation amount and the second rotation amount meet the second condition.

[0106] In the above technical solution, the reference object can be understood as the selected reference point, and the first non-reference object and the second non-reference object are other objects except the reference object among the objects.

[0107] In this technical solution, the construction scheme of the pose candidate set is specifically defined. In this scheme, the offset, the first rotation amount, and the second rotation amount are introduced, and the offset, the first rotation amount, and the second rotation amount are used to determine the pose candidate set. During this process, since the first rotation amount and the second rotation amount meet the second condition, that is, not all calculated offsets and first rotation amounts can be stored in the pose candidate set, the number in the pose candidate set can be reduced, so as to improve the search speed.

[0108] In the above technical solution, the offset of the reference object can be understood as follows: the coordinate obtained by mapping the reference object in the first sliding window to the two-dimensional grid is the first coordinate, and the coordinate obtained by mapping the reference object in the second sliding window to the two-dimensional grid is the second coordinate. The deviation between the first coordinate and the second coordinate is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0109] In the above technical solution, the first rotation amount can be understood as the angle value between the first connection line between the reference object and the first non-reference object in the first sliding window and the second connection line between the reference object and the first non-reference object in the second sliding window when the first connection line rotates to the second connection line.

[0110] Similarly, the second rotation amount can be understood as the angle value between the third connection line between the reference object and the second non-reference object in the first sliding window and the fourth connection line between the reference object and the second non-reference object in the second sliding window when the third connection line rotates to the fourth connection line.

[0111] In the above technical solution, the second condition can be that the difference between the second rotation amount and the first rotation amount is less than a preset angle value, or that the ratio between the second rotation amount and the first rotation amount is within a preset numerical range, such as between 0.8 and 1.2.

[0112] For example, assume that there is information about chairs, trash cans, and shoes in windows A and B. Then align the center points of the chairs in window B with the center points of the chairs in A to calculate the offsets DX and DY between them. Then calculate the angle value between the connection line between the trash can and the chair in window A and the connection line between the trash can and the chair in window B. The rotational offset is considered as D_theta. We further detect whether the offset of the shoes is also near D_theta. Through such a search, we finally find the candidate pose candidate set {DX, DY, D_theta} that meets the conditions.

[0113] In any of the above technical solutions, the determination unit is specifically further configured to: determine the first two-dimensional grid corresponding to the second sliding window according to the pose candidate set; determine the second two-dimensional grid of the first sliding window; determine the intersection and union of the first two-dimensional grid and the second two-dimensional grid; determine the area coincidence degree according to the intersection and union; and screen the pose candidate set according to the comparison result between the area coincidence degree and the preset area coincidence degree.

[0114] In this technical solution, considering that the pose candidate set is still relatively large and the calculation speed is still relatively slow, the technical solution of the present application introduces the area coincidence degree and uses the area coincidence degree to screen the pose candidate set in order to delete the pose candidate set, thereby improving the calculation speed.

[0115] Specifically, in the case where the calculated area overlap is small, the corresponding offset and first rotation amount are deleted, while in the case where the calculated area overlap is large, the corresponding offset and first rotation amount are retained.

[0116] In the above technical solution, a calculation solution for the area overlap is also given. In this solution, by counting the intersection and union between the first two-dimensional grid and the second two-dimensional grid, the ratio of the intersection to the union is used as the area overlap.

[0117] In the above technical solution, the intersection can be understood as the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid, and the union is the difference between the total number of grids of the first two-dimensional grid and the second two-dimensional grid and the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid.

[0118] In any of the above technical solutions, the area overlap is calculated by using the scan line and segment tree method.

[0119] According to DX, DY, and D_theta, perform a pose transformation on the grid map formed by the trash can, chair, and shoes formed by the second sliding window. Then match it with the grid map formed by the trash can, chair, and shoes formed by the first sliding window, and use the scan line and segment tree to calculate the area overlap to improve the calculation speed.

[0120] In the above technical solution, the scan line can be understood as scanning in a way that a line is translated as a whole to the two-dimensional grid, and the segment tree is a binary search tree, similar to the interval tree. It divides an interval into some unit intervals, and each unit interval corresponds to a leaf node in the segment tree. Using the segment tree can quickly find the number of times a certain node appears in several line segments.

[0121] In any of the above technical solutions, the cartographer algorithm is used to match the target image data.

[0122] In this technical solution, using the cartographer algorithm can improve the search speed for searching within the offset and rotation ranges.

[0123] In any of the above technical solutions, the judgment unit is specifically further configured to: obtain the center points of the reference object, the first non-reference object, and the second non-reference object; determine the offset, the first rotation amount, and the second rotation amount according to the center points.

[0124] In this technical solution, the center point is used to determine the offset, the first rotation amount, and the second rotation amount. It can be understood that the offset of the reference object can be understood as follows: the coordinates obtained by mapping the center point of the reference object in the first sliding window to the two-dimensional grid are the first coordinates, and the coordinates obtained by mapping the center point of the reference object in the second sliding window to the two-dimensional grid are the second coordinates. Among them, the deviation between the first coordinates and the second coordinates is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0125] In the above technical solution, the first rotation amount can be understood as the angle value between the first connection line between the center point of the reference object and the center point of the first non-reference object in the first sliding window and the second connection line between the center point of the reference object and the center point of the first non-reference object in the second sliding window when rotating from the first connection line to the second connection line.

[0126] Similarly, the second rotation amount can be understood as the angle value between the third connection line between the center point of the reference object and the center point of the second non-reference object in the first sliding window and the fourth connection line between the center point of the reference object and the center point of the second non-reference object in the second sliding window when rotating from the third connection line to the fourth connection line.

[0127] In any of the above technical solutions, the determination unit is specifically configured to: identify the visual data based on the neural network model to obtain the objects in the field of view.

[0128] In this technical solution, using the neural network model to identify the objects ensures the credibility of the identified objects.

[0129] According to the third aspect of the present invention, the present invention provides a loop detection device, including: a controller and a memory, wherein a program or instruction is stored in the memory, and when the controller executes the program or instruction in the memory, the steps of any of the above methods are implemented.

[0130] According to the fourth aspect of the present invention, the present invention provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of any of the above methods are implemented.

[0131] According to the fifth aspect of the present invention, the present invention provides a movable device, including: any of the above loop detection devices; or the above readable storage medium.

[0132] In the above technical solution, the movable device includes a sweeping robot.

[0133] The additional aspects and advantages of the present invention will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings

[0134] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:

[0135] Figure 1 FIG. 1 shows one of the flow diagrams of the loop detection method in an embodiment of the present invention;

[0136] Figure 2 FIG. 2 shows the flow diagram of determining the first sliding window and the second sliding window in an embodiment of the present invention;

[0137] Figure 3 FIG. 3 shows another flow diagram of the loop detection method in an embodiment of the present invention;

[0138] Figure 4 FIG. 4 shows the schematic block diagram of the loop detection device in an embodiment of the present invention;

[0139] Figure 5 FIG. 5 shows the schematic block diagram of loop detection in an embodiment of the present invention;

[0140] Figure 6 FIG. 6 shows the flow diagram of the overall idea of loop detection in an embodiment of the present invention;

[0141] Figure 7 FIG. 7 shows the schematic block diagram of the loop detection device in an embodiment of the present invention. Detailed Embodiments

[0142] In order to more clearly understand the above aspects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0143] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0144] Next, reference is made to Figures 1 to 7 Describe a loop detection method, device, readable storage medium and mobile device according to some embodiments of the present invention.

[0145] As Figure 1 shown, according to the first aspect of the present invention, the present invention provides a loop detection method, including:

[0146] Step 102, periodically obtaining image data in the field of view;

[0147] Step 104: Determine a first sliding window and a second sliding window sorted in chronological order;

[0148] Step 106: When the first sliding window and the second sliding window meet the first condition, determine that there is a loop in the first sliding window and the second sliding window.

[0149] Among them, the first sliding window and the second sliding window include at least two frames of image data.

[0150] The design of this application proposes a loop detection method. By running this detection method, the robustness can be greatly improved, and the probability of errors in loop detection can be reduced.

[0151] The design of this application is implemented based on the following principles. Specifically:

[0152] SLAM (simultaneous localization and mapping), also known as CML (Concurrent Mapping and Localization), is simultaneous localization and mapping, or concurrent mapping and localization. The problem can be described as follows: When a robot is placed at an unknown location in an unknown environment, is there a way for the robot to gradually draw a complete map of this environment while moving? A so-called complete map (a consistent map) means that the robot can travel to every accessible corner of the room without being blocked.

[0153] Loop detection helps reduce the cumulative error of robot pose estimation and generate a consistent global map. In the past decade, the latest loop detection algorithms in visual and lidar SLAM mainly fall into three categories: vision-based loop detection methods, lidar-based loop detection methods, and deep learning-based loop detection methods. These methods currently compare the current single-frame data with historical data, and their robustness is not strong.

[0154] In the design of this application, sliding windows are compared with each other. Since there is more than one frame of image data in the sliding window, the result after fusing multiple consecutive frames of data can be used to compare with historical data, thereby greatly improving the robustness and reducing false loop detection.

[0155] In one possible design, the field of view involved in this application can be understood as the image captured by a visual sensor.

[0156] In the above design, due to the cumulative error, the map will finally drift. A loop can be understood as pulling the position on the map with cumulative error back to the correct position when it is detected that the same position has been experienced, thereby eliminating the cumulative error.

[0157] In the above design, the image data includes visual data and point cloud data. As Figure 2 shown, determining the first sliding window and the second sliding window specifically includes:

[0158] Step 202, determining the objects in the field of view according to the visual data;

[0159] Step 204, determining the bounding box of the object and the category of the object;

[0160] Step 206, annotating the point cloud data according to the bounding box of the object and the category of the object;

[0161] Step 208, mapping the annotated point cloud data to a two-dimensional grid;

[0162] Step 210, determining that during the mapping process, when the category and quantity of the object remain unchanged within the first number of frames, the image data of the first number of frames is recorded as a sliding window; or determining that during the mapping process, when the cumulative number of frames of the image data is greater than or equal to the second number of frames, the image data of the second number of frames is recorded as a sliding window.

[0163] In this design, the visual data can be obtained by using a visual sensor, and the point cloud data can be obtained by using a TOF (Time of Flight) point cloud module. Among them, the form of the point cloud data can be in the XYZ form.

[0164] By marking the bounding box for the object, when annotating the point cloud data, the accuracy of the annotation can be ensured. At the same time, during the process of mapping the point cloud data to the two-dimensional grid, the contour of the object in the two-dimensional grid can be determined, so as to use the contour of the object in the two-dimensional grid to implement loop detection. In the above solution, morphological and graphics technologies are combined, which improves the speed of loop detection and at the same time improves the robustness of loop detection.

[0165] In the above design, the selection scheme of the sliding window is also specifically defined. In the above solution, the characteristics that the category and quantity of the object no longer change are used to implement the selection of the sliding window, so as to minimize the size of the sliding window to the greatest extent, thereby reducing the computational amount of loop detection.

[0166] In the above design, it can be understood that in the continuous image data of the first number of frames, when the determined category no longer changes and at the same time the quantity of the object does not change, it is considered that the category and quantity of the object no longer change.

[0167] In one possible design, the quantity of the object can be the quantity of any category of objects or the quantity of all categories of objects.

[0168] In the above design, the number of objects can be obtained through statistics during the process of determining the category of the object. Specifically, the center point and the directional radius of the point cloud data of each category in the two-dimensional plane are calculated. Among them, the center point can be determined by using the depth value in the point cloud data. Specifically, the depth values are sorted, and the median point is selected as the center point of the object. At the same time, assuming that the point cloud conforms to the Gaussian distribution, we calculate the standard deviations in two directions of the two-dimensional plane as the corresponding directional radii. In this way, the estimated center point and radius of each category are obtained.

[0169] For example, suppose there are 2 categories, chairs and trash cans. When a chair is detected in the first frame, set the second bit of the binary sequence to 1, representing 10, and record the number of different chairs at the same time. We already have the center point and radius of the current chair as mentioned above. In the second frame, if it is still the same chair, we consider that the rectangle formed by the estimated center point and radius has a large intersection with that in the first frame, and the number of chairs is still 1 at this time. At the same time, update the radius and the position of the center point of the chair according to the principle of multiplying Gaussian distributions. If the second frame is another chair, then the rectangle formed by the estimated center point and radius has a small intersection with that in the first frame. We increase the number of chairs by 1 and record it as 2.

[0170] It should be noted that the two-dimensional grid is updated not according to the estimated center point and radius of the object, but according to the projection of the actual point cloud belonging to the object category on the two-dimensional plane, so as to reflect the actual contour of the object. If a trash can is detected, then set the first bit of the binary sequence to 1 and record the quantity at the same time.

[0171] The method of determining the sliding window by whether the cumulative number of frames of the image data is greater than the second number of frames ensures the integrity of the sliding window to the greatest extent.

[0172] In the above design, the first number of frames is less than or equal to the second number of frames.

[0173] In the above design, after the first sliding window, the same scheme is used to determine the second sliding window.

[0174] In any of the above designs, the first condition includes at least one of the following: the Euclidean distance between the binary sequence of the first sliding window and the binary sequence of the second sliding window is less than the first value, the ratio of the first quantity of the object of the corresponding category in the second sliding window to the second quantity of the object of the corresponding category in the first sliding window is greater than or equal to the second value, and the distances between the corresponding same or different objects are the same.

[0175] In this design, the possible contents included in the first condition are specifically given. In this solution, when determining the category of an object, the first sliding window and the second sliding window are statistically analyzed, and the above statistical results are reflected in the form of a binary sequence. For example, if the object corresponds to the 12th bit in the binary sequence, if the object is detected, the 12th bit in the binary sequence is set to 1, and vice versa, it is set to 0.

[0176] By calculating the Euclidean distance between the first sliding window and the second sliding window, the similarity degree between the first sliding window and the second sliding window is calculated. When the Euclidean distance is less than the first value, it is considered that the similarity degree between the first sliding window and the second sliding window is relatively high.

[0177] Similarly, two other solutions can also be used for detection. Among them, the use of the above three solutions can be selected according to the actual usage scenario.

[0178] In the above design, the distances between the same or different objects are the same. It can be understood that the distances between the same objects in the first sliding window are the same as the distances between the corresponding same objects in the second sliding window, and the distances between different objects in the first sliding window are the same as the distances between the corresponding different objects in the second sliding window.

[0179] For example, if the center distance between the chair and the trash can in the sliding window A is 1 meter, then the center distance between the chair and the trash can in the sliding window B should also be around 1 meter.

[0180] In the above design, when the first condition is met, it is considered that there is a loop between the first sliding window and the second sliding window.

[0181] In any of the above designs, as Figure 3 shown, it also includes:

[0182] Step 302, obtain the first image data in the second sliding window, where the first image data is the image data closest to the current moment in the second sliding window;

[0183] Step 304, determine the target object included in the first image data;

[0184] Step 306, determine the pose candidate set including the target object;

[0185] Step 308, determine the target image data in the pose candidate set that matches the first image data;

[0186] Step 310, use the pose data corresponding to the target image data as the pose data of the first image data.

[0187] In this design, a pose determination scheme based on loop detection is given. In this scheme, by searching for target image data that matches the first image data, the pose data corresponding to the target image data is used as the pose data corresponding to the first image data to achieve pose determination.

[0188] In the above design, a pose candidate set is determined using a target object to reduce the screening range of the target image data, thereby reducing the amount of data searched during the pose data determination process, and thus improving the pose search speed.

[0189] In the above design, since the target object is determined based on the first image data, the matching degree between the pose candidate set and the first image data can be ensured, thereby improving the accuracy of pose data determination.

[0190] In the above design, it can be understood that the first image data is the latest frame of image data in the second sliding window.

[0191] In the above design, artificial intelligence can be used to recognize the first image data to determine the target object contained in the first image data.

[0192] In the above design, the number of target objects can be one or multiple.

[0193] In the above design, the pose candidate set can be indexed by an object so that after the target object is determined, the index is used to determine the pose candidate set.

[0194] In any of the above designs, it further includes: determining the offset of the reference object and the first rotation amount corresponding to the line connecting the reference object and the first non-reference object under the two-dimensional grid corresponding to the first sliding window and the second sliding window; determining the second rotation amount corresponding to the line connecting the reference object and the second non-reference object; and storing the offset and the first rotation amount in the pose candidate set when the first rotation amount and the second rotation amount meet the second condition.

[0195] In the above design, the reference object can be understood as the selected reference point, and the first non-reference object and the second non-reference object are other objects in the objects except the reference object.

[0196] In this design, the construction scheme of the pose candidate set is specifically defined. In this scheme, an offset, a first rotation amount, and a second rotation amount are introduced, and the offset, the first rotation amount, and the second rotation amount are used to determine the pose candidate set. In this process, since the first rotation amount and the second rotation amount meet the second condition, that is, not all calculated offsets and first rotation amounts can be stored in the pose candidate set, the number in the pose candidate set can be reduced to improve the search speed.

[0197] In the above design, the offset of the reference object can be understood as follows: the coordinates of the reference object in the first sliding window mapped to the two-dimensional grid are the first coordinates, and the coordinates of the reference object in the second sliding window mapped to the two-dimensional grid are the second coordinates. Among them, the deviation between the first coordinates and the second coordinates is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0198] In the above design, the first rotation amount can be understood as the angle value between the first connection line between the reference object and the first non-reference object in the first sliding window and the second connection line between the reference object and the first non-reference object in the second sliding window when the first connection line rotates to the second connection line.

[0199] Similarly, the second rotation amount can be understood as the angle value between the third connection line between the reference object and the second non-reference object in the first sliding window and the fourth connection line between the reference object and the second non-reference object in the second sliding window when the third connection line rotates to the fourth connection line.

[0200] In the above design, the second condition can be that the difference between the second rotation amount and the first rotation amount is less than a preset angle value, or that the ratio between the second rotation amount and the first rotation amount is within a preset numerical range, such as between 0.8 and 1.2.

[0201] For example, assume that there is chair, trash can, and shoe information in windows A and B. Then align the center point of the chair in window B with the center point of the chair in A to calculate the offset DX and DY between the two. Then calculate the angle value between the connection line of the trash can and the chair in window A and the connection line of the trash can and the chair in window B. The rotational offset is considered as D_theta. We further detect whether the shoe offset is also near D_theta. Through such a search, we finally find the candidate pose candidate set that meets the conditions {DX, DY, D_theta}.

[0202] In any of the above designs, it further includes: determining the first two-dimensional grid corresponding to the second sliding window according to the pose candidate set; determining the second two-dimensional grid of the first sliding window; determining the intersection and union of the first two-dimensional grid and the second two-dimensional grid; determining the area coincidence degree according to the intersection and union; and screening the pose candidate set according to the comparison result between the area coincidence degree and the preset area coincidence degree.

[0203] In this design, considering that the pose candidate set is still relatively large and the calculation speed is still relatively slow, the design of this application introduces the area coincidence degree and uses the area coincidence degree to screen the pose candidate set in order to delete the pose candidate set, thereby improving the calculation speed.

[0204] Specifically, in the case where the calculated area coincidence degree is small, the corresponding offset and the first rotation amount are deleted, while in the case where the calculated area coincidence degree is large, the corresponding offset and the first rotation amount are retained.

[0205] In the above design, a calculation scheme for the area coincidence degree is also given. In this scheme, by counting the intersection and union between the first two-dimensional grid and the second two-dimensional grid, the ratio of the intersection to the union is used as the area coincidence degree.

[0206] In the above design, the intersection can be understood as the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid, and the union is the difference between the total number of grids of the first two-dimensional grid and the second two-dimensional grid and the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid.

[0207] In any of the above designs, the area coincidence degree is calculated by using the scan line and the segment tree.

[0208] According to DX, DY, and D_theta, perform a pose transformation on the grid map formed by the trash can, chair, and shoes formed by the second sliding window. Then match it with the grid map formed by the trash can, chair, and shoes formed by the first sliding window, and use the scan line and the segment tree to calculate the area coincidence degree to improve the calculation speed.

[0209] In the above design, the scan line can be understood as scanning in a way that a line is translated as a whole to the two-dimensional grid, and the segment tree is a binary search tree, similar to the interval tree. It divides an interval into some unit intervals, and each unit interval corresponds to a leaf node in the segment tree. Using the segment tree, the number of times a certain node appears in several line segments can be quickly found.

[0210] In any of the above designs, the Cartographer algorithm is used to match the target image data.

[0211] In this design, using the Cartographer algorithm can improve the search speed within the offset and rotation ranges.

[0212] In any of the above designs, it also includes: obtaining the center points of the reference object, the first non-reference object, and the second non-reference object; determining the offset, the first rotation amount, and the second rotation amount according to the center points.

[0213] In this design, the center point is used to determine the offset, the first rotation amount, and the second rotation amount. It can be understood that the offset of the reference object can be regarded as the coordinates obtained by mapping the center point of the reference object in the first sliding window to the two-dimensional grid as the first coordinates, and the coordinates obtained by mapping the center point of the reference object in the second sliding window to the two-dimensional grid as the second coordinates. Among them, the deviation between the first coordinates and the second coordinates is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0214] In the above design, the first rotation amount can be understood as the angle value of the rotation between the first connection line between the center point of the reference object and the center point of the first non-reference object in the first sliding window and the second connection line between the center point of the reference object and the center point of the first non-reference object in the second sliding window.

[0215] Similarly, the second rotation amount can be understood as the angle value of the rotation between the third connection line between the center point of the reference object and the center point of the second non-reference object in the first sliding window and the fourth connection line between the center point of the reference object and the center point of the second non-reference object in the second sliding window.

[0216] In any of the above designs, determining the objects in the field of view according to the visual data includes: identifying the visual data based on a neural network model to obtain the objects in the field of view.

[0217] In this design, using a neural network model to identify objects ensures the credibility of the identified objects.

[0218] In one of the embodiments, as Figure 4 shown, the present invention provides a loop detection device 400, including: an acquisition unit 402 for periodically acquiring image data in the field of view; a determination unit 404 for determining a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of image data; a judgment unit 406 for determining that there is a loop in the first sliding window and the second sliding window when the first sliding window and the second sliding window meet the first condition.

[0219] The design of this application proposes a loop detection device 400. A movable device with this detection device can greatly improve the robustness and reduce the probability of errors in loop detection.

[0220] The design of this application is implemented based on the following principle. Specifically:

[0221] SLAM (Simultaneous Localization and Mapping), also known as CML (Concurrent Mapping and Localization), is simultaneous localization and mapping, or concurrent mapping and localization. The problem can be described as follows: If a robot is placed at an unknown location in an unknown environment, is there a way for the robot to gradually build a complete map of this environment while moving? A complete map (a consistent map) means traveling through every accessible corner of the room without obstacles.

[0222] Loop detection helps reduce the cumulative error of robot pose estimation and generate a consistent global map. In the past decade, the latest loop detection algorithms for visual and lidar SLAM mainly fall into three categories: vision-based loop detection methods, lidar-based loop detection methods, and deep learning-based loop detection methods. These methods currently compare the current single-frame data with historical data and are not very robust.

[0223] In the design of this application, a sliding window is compared with another sliding window. Since there is more than one frame of image data in the sliding window, the result after fusing multiple consecutive frames of data can be used to compare with historical data, thereby greatly improving the robustness and reducing false loop detections.

[0224] In one possible design, the field of view involved in this application can be understood as the image captured by the visual sensor.

[0225] In the above design, due to the cumulative error, the map will eventually drift. A loop can be understood as pulling the position with cumulative error on the map back to the correct position when it is detected that the same position has been experienced, thereby eliminating the cumulative error.

[0226] In the above design, the determination unit 404 is specifically configured to: determine the objects in the field of view according to the visual data; determine the bounding box and category of the object; label the point cloud data according to the bounding box and category of the object; map the labeled point cloud data to a two-dimensional grid; determine that when the category and quantity of the object remain unchanged within the first number of frames during the mapping process, record the image data of the first number of frames as a sliding window; or determine that when the cumulative number of frames of the image data is greater than or equal to the second number of frames during the mapping process, record the image data of the second number of frames as a sliding window.

[0227] In this design, visual data can be acquired using visual sensors, and point cloud data can be obtained using a TOF point cloud module. Among them, the TOF (Time of Flight) point cloud module obtains the point cloud data, and the form of the point cloud data can be in the XYZ form.

[0228] In one possible design, while mapping to a two-dimensional grid, the corresponding semantic category and grid index are recorded. Among them, recording the corresponding semantic category can be understood as recording the category of the object into the binary sequence of the sliding window.

[0229] By marking a bounding box for the object, when annotating the point cloud data, the accuracy of the annotation can be ensured. At the same time, during the process of mapping the point cloud data to the two-dimensional grid, the contour of the object in the two-dimensional grid can be determined, so as to use the contour of the object in the two-dimensional grid to achieve loop detection. In the above scheme, morphological and graphics technologies are combined, which improves the speed of loop detection and at the same time improves the robustness of loop detection.

[0230] In the above design, the selection scheme of the sliding window is also specifically defined. In the above scheme, the characteristics that the category and quantity of the object no longer change are used to realize the selection of the sliding window, so as to minimize the size of the sliding window to the greatest extent, thereby reducing the computational amount of loop detection.

[0231] In the above design, it can be understood that in the continuous first frame number of image data, the determined category no longer changes, and at the same time, the quantity of the object also does not change, and it is considered that the category and quantity of the object no longer change.

[0232] In one possible design, the quantity of the object can be the quantity of any category of object or the quantity of all categories of objects.

[0233] In the above design, the quantity of the object can be statistically obtained during the process of determining the category of the object. Specifically, the center point and direction radius of the point cloud data of each category in the two-dimensional plane are calculated. Among them, the center point can be determined using the depth value in the point cloud data. Specifically, the depth values are sorted, and the median point is selected as the center point of the object. At the same time, assuming that the point cloud conforms to a Gaussian distribution, we calculate the standard deviations in two directions of the two-dimensional plane as the corresponding direction radii respectively. In this way, the estimated center point and radius of each category are obtained.

[0234] Among them, the process of estimating the center point first limits the height in a relatively low interval. Because we assume that the corresponding object is the nearest object seen from the viewing direction, limiting the height can exclude the interference of the subsequent object to a certain extent.

[0235] For example, assume there are two categories, chairs and trash cans. When a chair is detected in the first frame, a 1 is placed at the second position in the binary sequence, representing 10, and the number of different chairs is recorded. Above, we already have the center point and radius of the current chair. In the second frame, if it is the same chair, we consider that the rectangle formed by the estimated center point and radius has a large intersection with that in the first frame, and the number of chairs remains 1. At the same time, the radius and center point position of the chair are updated according to the principle of multiplying Gaussian distributions. If the second frame is another chair, then the rectangle formed by the estimated center point and radius has a small intersection with that in the first frame. We increment the number of chairs by 1 and record it as 2.

[0236] It should be noted that the two-dimensional grid is updated not based on the estimated center point and radius of the object, but based on the projection of the actual point cloud belonging to the object category on the two-dimensional plane, so as to reflect the actual contour of the object. If a trash can is detected, then the first position in the binary sequence is set to 1 and the quantity is recorded.

[0237] The method of determining the sliding window by whether the cumulative number of frames of the image data is greater than the second number of frames ensures the integrity of the sliding window to the greatest extent.

[0238] In the above design, the first number of frames is less than or equal to the second number of frames.

[0239] In the above design, after the first sliding window, the same scheme is used to determine the second sliding window.

[0240] In any of the above designs, the first condition includes at least one of the following: the Euclidean distance between the binary sequences of the first sliding window and the second sliding window is less than the first value, the ratio of the first quantity of the objects of the corresponding category in the second sliding window to the second quantity of the objects of the corresponding category in the first sliding window is greater than or equal to the second value, and the distances between the same or different objects are the same.

[0241] In this design, the possible contents included in the first condition are specifically given. In this scheme, when determining the category of an object, statistics are performed on the first sliding window and the second sliding window, and the above statistical results are reflected in the form of a binary sequence. For example, if the object corresponds to the 12th bit in the binary sequence, if the object is detected, the 12th position in the binary sequence is set to 1, otherwise, it is set to 0.

[0242] By calculating the Euclidean distance between the first sliding window and the second sliding window, the similarity degree between the first sliding window and the second sliding window is calculated. When the Euclidean distance is less than the first value, it is considered that the similarity degree between the first sliding window and the second sliding window is relatively high.

[0243] Similarly, two other schemes can also be adopted for detection. Among them, the use of the above three schemes can be selected according to the actual usage scenario.

[0244] In the above design, the distances between corresponding same or different objects are the same. It can be understood that the distance between the same objects in the first sliding window is the same as the distance between the corresponding same objects in the second sliding window, and the distance between different objects in the first sliding window is the same as the distance between the corresponding different objects in the second sliding window.

[0245] For example, if the distance between the centers of the chair and the trash can in sliding window A is 1 meter, then the distance between the centers of the chair and the trash can in sliding window B should also be around 1 meter.

[0246] In the above design, when the first condition is met, it is considered that there is a loop between the first sliding window and the second sliding window.

[0247] In any of the above designs, the determination unit 406 is further configured to: obtain the first image data in the second sliding window, where the first image data is the image data closest to the current moment in the second sliding window; determine the target objects included in the first image data; determine the pose candidate set including the target objects; determine the target image data in the pose candidate set that matches the first image data; and use the pose data corresponding to the target image data as the pose data of the first image data.

[0248] In this design, a pose determination scheme based on loop detection is given. In this scheme, by searching for the target image data that matches the first image data, the pose data corresponding to the target image data is used as the pose data corresponding to the first image data to achieve the determination of the pose.

[0249] In the above design, the pose candidate set is determined using the target objects to reduce the screening range of the target image data, thereby reducing the amount of data searched during the determination of the pose data, and thus improving the searching speed of the pose.

[0250] In the above design, since the target objects are determined based on the first image data, the matching degree between the pose candidate set and the first image data can be ensured, thereby improving the accuracy of the determination of the pose data.

[0251] In the above design, it can be understood that the first image data is the latest frame of image data in the second sliding window.

[0252] In the above design, artificial intelligence can be used to identify the first image data to determine the target objects included in the first image data.

[0253] In the above design, the number of target objects can be one or multiple.

[0254] In the above design, the pose candidate set can be indexed by the object, so that after the target object is determined, the index can be used to determine the pose candidate set.

[0255] In any of the above designs, the determination unit 406 is further specifically configured to: determine the offset of the reference object and the first rotation amount corresponding to the line connecting the reference object and the first non-reference object under the two-dimensional grid corresponding to the first sliding window and the second sliding window; determine the second rotation amount corresponding to the line connecting the reference object and the second non-reference object; and store the offset and the first rotation amount in the pose candidate set when the first rotation amount and the second rotation amount meet the second condition.

[0256] In the above design, the reference object can be understood as the selected reference point, and the first non-reference object and the second non-reference object are other objects in the objects except the reference object.

[0257] In this design, the construction scheme of the pose candidate set is specifically defined. In this scheme, the offset, the first rotation amount, and the second rotation amount are introduced, and the offset, the first rotation amount, and the second rotation amount are used to determine the pose candidate set. In this process, since the first rotation amount and the second rotation amount meet the second condition, that is, not all the calculated offsets and first rotation amounts can be stored in the pose candidate set, the number in the pose candidate set can be reduced to improve the search speed.

[0258] In the above design, the offset of the reference object can be understood as follows: the coordinate of the reference object in the first sliding window mapped to the two-dimensional grid is the first coordinate, and the coordinate of the reference object in the second sliding window mapped to the two-dimensional grid is the second coordinate. Among them, the deviation between the first coordinate and the second coordinate is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0259] In the above design, the first rotation amount can be understood as the angle value between the first line connecting the reference object and the first non-reference object in the first sliding window and the second line connecting the reference object and the first non-reference object in the second sliding window, when the first line rotates to the second line.

[0260] Similarly, the second rotation amount can be understood as the angle value between the third line connecting the reference object and the second non-reference object in the first sliding window and the fourth line connecting the reference object and the second non-reference object in the second sliding window, when the third line rotates to the fourth line.

[0261] In the above design, the second condition may be that the difference between the second rotation amount and the first rotation amount is less than a preset angular value, or that the ratio between the second rotation amount and the first rotation amount is within a preset numerical range, such as between 0.8 and 1.2.

[0262] For example, assume that there is information about chairs, trash cans, and shoes in windows A and B. Then, align the center points of the chairs in window B with the center points of the chairs in A to calculate the offset amounts DX and DY between them. Then, calculate the angular value between the line connecting the trash can and the chair in window A and the line connecting the trash can and the chair in window B. Consider the rotational offset amount as D_theta. We further detect whether the shoe offset amount is also near D_theta. Through such a search, we finally find the candidate pose candidate set {DX, DY, D_theta} that meets the conditions.

[0263] In any of the above designs, the determination unit 406 is specifically further configured to: determine the first two-dimensional grid corresponding to the second sliding window according to the pose candidate set; determine the second two-dimensional grid of the first sliding window; determine the intersection and union of the first two-dimensional grid and the second two-dimensional grid; determine the area coincidence degree according to the intersection and union; and screen the pose candidate set according to the comparison result between the area coincidence degree and the preset area coincidence degree.

[0264] In this design, considering that the pose candidate set is still relatively large and the calculation speed is still relatively slow, the design of this application introduces the area coincidence degree and uses the area coincidence degree to screen the pose candidate set in order to delete the pose candidate set, thereby improving the calculation speed.

[0265] Specifically, in the case where the calculated area coincidence degree is small, the corresponding offset amount and the first rotation amount are deleted, and in the case where the calculated area coincidence degree is large, the corresponding offset amount and the first rotation amount are retained.

[0266] In the above design, a calculation scheme for the area coincidence degree is also given. In this scheme, by counting the intersection and union between the first two-dimensional grid and the second two-dimensional grid, the ratio of the intersection to the union is used as the area coincidence degree.

[0267] In the above design, the intersection can be understood as the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid, and the union is the difference between the total number of grids of the first two-dimensional grid and the second two-dimensional grid and the number of overlapping grids between the first two-dimensional grid and the second two-dimensional grid.

[0268] In any of the above designs, the area coincidence degree is calculated by using the scan line and segment tree method.

[0269] Perform pose transformation on the grid maps of trash cans, chairs, and shoes formed by the second sliding window according to DX, DY, and D_theta. Then match it with the grid maps of trash cans, chairs, and shoes formed by the first sliding window, and use scan lines and segment trees to calculate the area overlap to improve the calculation speed.

[0270] In the above design, the scan line can be understood as scanning in a way that a line is translated as a whole in the two-dimensional grid, and the segment tree is a binary search tree, similar to the interval tree. It divides an interval into some unit intervals, and each unit interval corresponds to a leaf node in the segment tree. Using the segment tree can quickly find the number of times a certain node appears in several line segments.

[0271] In any of the above designs, the Cartographer algorithm is used to match the target image data.

[0272] In this design, using the Cartographer algorithm can improve the search speed within the offset and rotation ranges.

[0273] In any of the above designs, the judgment unit 406 is specifically further used for: obtaining the center points of the reference object, the first non-reference object, and the second non-reference object; determining the offset, the first rotation amount, and the second rotation amount according to the center points.

[0274] In this design, the center points are used to determine the offset, the first rotation amount, and the second rotation amount.

[0275] It can be understood that the offset of the reference object can be understood as follows: the coordinate obtained by mapping the center point of the reference object in the first sliding window to the two-dimensional grid is the first coordinate, and the coordinate obtained by mapping the center point of the reference object in the second sliding window to the two-dimensional grid is the second coordinate. Among them, the deviation between the first coordinate and the second coordinate is the offset, denoted as DX and DY, where DX represents the offset in the X-axis direction, and DY represents the offset in the Y-axis direction.

[0276] In the above design, the first rotation amount can be understood as the angle value between the first connection line between the center point of the reference object and the center point of the first non-reference object in the first sliding window and the second connection line between the center point of the reference object and the center point of the first non-reference object in the second sliding window when rotating from the first connection line to the second connection line.

[0277] Similarly, the second rotation amount can be understood as the angle value between the third connection line between the center point of the reference object and the center point of the second non-reference object in the first sliding window and the fourth connection line between the center point of the reference object and the center point of the second non-reference object in the second sliding window when rotating from the third connection line to the fourth connection line.

[0278] In any of the above designs, the determination unit 404 is specifically configured to: identify visual data based on a neural network model to obtain objects in the field of view.

[0279] In this design, a neural network model is used to identify objects, ensuring the credibility of the identified objects.

[0280] In one embodiment, as Figure 5 shown, an RGBD module is used to obtain image data. The AI object recognition module identifies objects in the current field of view based on the image and gives the BoundingBox and object category of each object. The TOF point cloud module obtains point cloud in the form of XYZ and performs preprocessing. The object semantic generation module labels the point cloud with category labels according to the results given by the AI object recognition module. The sliding window A maps the labeled point cloud to a two-dimensional grid and records its corresponding semantic category and grid index. The continuous multi-frame sliding window A performs the same operation. When it is found that the semantic categories in the window no longer increase or reach a certain frame number limit, the sliding window A is fixed and a new sliding window B is created. B also performs a similar operation. At the same time, a judgment module for loop detection condition A is started. When condition A is satisfied, it enters the loop detection and pose correction module for further judgment.

[0281] Among them, the overall idea of loop detection is as Figure 6 shown, and the overall idea of loop detection includes:

[0282] Step 602, judging whether the scenes reflected by the sliding window A and the sliding window B are the same scene according to the object semantic category and the geometric position distribution between objects;

[0283] Step 604, aligning the pose of the two-dimensional grid of the sliding window A and the two-dimensional grid of the sliding window B according to the object semantic category and the geometric position between objects;

[0284] Step 606, judging whether the intersection of the areas of the two-dimensional grids corresponding to the sliding windows A and B given by the scan line and segment tree method accounts for more than a set threshold of their union. If the judgment result is yes, execute step 608; if the judgment result is no, end;

[0285] Step 608, matching the mapping result after mapping the latest frame in the sliding window B to the two-dimensional grid with the grid map mapped by each frame in the sliding window A to obtain a matching result.

[0286] In one embodiment, as Figure 7 shown, the present invention provides a loop detection device 700, including: a controller 702 and a memory 704. Among them, a program or instruction is stored in the memory 704, and when the controller 702 executes the program or instruction in the memory 704, the steps of any of the above methods are implemented.

[0287] In one embodiment, the present invention provides a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of any one of the above methods are implemented.

[0288] In one embodiment, the present invention provides a movable device, comprising: any one of the above loop detection devices; or the above readable storage medium.

[0289] In the above design, the movable device includes a floor cleaning robot.

[0290] In the description of the present invention, the term "a plurality" means two or more, unless otherwise clearly defined. The terms "upper", "lower", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention; the terms "connection", "installation", "fixation", etc. should all be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected, or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0291] In the description of the present invention, the descriptions of the terms "one embodiment", "some embodiments", "specific embodiments", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the present invention, the schematic expressions of the above terms do not necessarily refer to the same embodiment or instance. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0292] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A loop detection method, characterized in that, Including: Periodically obtaining image data in the field of view; Determining a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of the image data; When the first sliding window and the second sliding window meet a first condition, determining that there is a loop in the first sliding window and the second sliding window; The first condition includes at least one of the following: The Euclidean distance between the binary sequence of the first sliding window and the binary sequence of the second sliding window is less than a first value, the ratio of the first number of objects of a corresponding category in the second sliding window to the second number of objects of the corresponding category in the first sliding window is greater than or equal to a second value, and the distances between corresponding same or different objects are the same.

2. The loop detection method according to claim 1, wherein The image data includes visual data and point cloud data. Determining the first sliding window and the second sliding window specifically includes: Determining the objects in the field of view according to the visual data; Determining the bounding box of the object and the category of the object; Labeling the point cloud data according to the bounding box of the object and the category of the object; Mapping the labeled point cloud data to a two-dimensional grid; Determining that when the category and quantity of the object remain unchanged within a first number of frames during the mapping process, recording the image data of the first number of frames as a sliding window; or Determining that when the cumulative number of frames of the image data is greater than or equal to a second number of frames during the mapping process, recording the image data of the second number of frames as a sliding window.

3. The loop detection method according to claim 1 or 2, characterized in that, Also including: Obtaining first image data in the second sliding window, where the first image data is the image data closest to the current moment in the second sliding window; Determining the target object included in the first image data; Determining a pose candidate set including the target object; Determining the target image data in the pose candidate set that matches the first image data; Taking the pose data corresponding to the target image data as the pose data of the first image data.

4. The loop detection method according to claim 3, wherein, Also including: Determining the offset of a reference object and the first rotation amount corresponding to the line connecting the reference object and a first non-reference object under the two-dimensional grids corresponding to the first sliding window and the second sliding window; Determining the second rotation amount corresponding to the line connecting the reference object and a second non-reference object; When the first rotation amount and the second rotation amount meet a second condition, storing the offset and the first rotation amount in the pose candidate set.

5. The loop detection method according to claim 3, characterized in that Also including: Determining a first two-dimensional grid corresponding to the second sliding window according to the pose candidate set; Determining a second two-dimensional grid of the first sliding window; Determining the intersection and union of the first two-dimensional grid and the second two-dimensional grid; Determining the area coincidence degree according to the intersection and the union; Screening the pose candidate set according to the comparison result between the area coincidence degree and a preset area coincidence degree.

6. The loop detection method according to claim 5, characterized in that Calculating the area coincidence degree by using the method of scan lines and segment trees.

7. The loop detection method according to claim 3, characterized in that, Match the target image data using the Cartographer algorithm.

8. The loop detection method according to claim 4, wherein It further includes: Obtain the center points of the reference object, the first non-reference object, and the second non-reference object; Determine the offset, the first rotation amount, and the second rotation amount according to the center points.

9. The loop detection method according to claim 2, wherein Determine the objects in the field of view according to the visual data, including: Identify the objects in the field of view based on a neural network model for the visual data.

10. A loop detection device, characterized in that, It includes: An acquisition unit for periodically acquiring image data in the field of view; A determination unit for determining a first sliding window and a second sliding window sorted in chronological order, where the first sliding window and the second sliding window include at least two frames of the image data; A judgment unit for determining that there is a loop in the first sliding window and the second sliding window when the first sliding window and the second sliding window meet a first condition; Wherein, the first condition includes at least one of the following: the Euclidean distance between the binary sequences of the first sliding window and the second sliding window is less than a first value, the ratio of the first number of objects of a corresponding category in the second sliding window to the second number of objects of the corresponding category in the first sliding window is greater than or equal to a second value, and the distances between the same or different objects are the same.

11. A loop detection device, characterized in that, It includes: A controller and a memory, where a program or instruction is stored in the memory, and the controller implements the steps of the method according to any one of claims 1 to 9 when executing the program or instruction in the memory.

12. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and the program or instruction implements the steps of the method according to any one of claims 1 to 9 when executed by a processor.

13. A mobile device, characterized in that, It includes: The loop detection device according to claim 10 or 11; Or The readable storage medium according to claim 12.

14. The mobile device according to claim 13, wherein The mobile device includes a sweeping robot.

Citation Information

Patent Citations

  • Binocular inertial navigation SLAM system carrying out feature matching based on front and rear frames

    CN110125928A

  • Method and device for constructing map based on classification detection network

    CN111339226A