Visual SLAM loopback detection method and device, equipment and storage medium
By using semantic information and feature description information in visual SLAM loopback detection to match target map points in map submap, the problem of time-consuming and low accuracy in traditional methods is solved, and more efficient and accurate loopback detection is achieved.
Patent Information
- Application Number
- CN202311678575.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional visual SLAM loopback detection finds spatial coordinate points corresponding to the keyframe that is most similar to the current keyframe in the global visual map, which takes a long time and is low in accuracy.
By obtaining the current keyframe as the matching frame, and determining the target map submap from the pre-constructed visual map containing multiple map submaps, the semantic information of the matching frame and the semantic information of the map submap filter out the target map point set that is roughly matched with the matching frame, narrowing the matching range, and then using feature description information to find the map points matching the feature points in the matching frame.
The accuracy and speed of loopback detection are improved. By increasing the number of map points in the map submap, more matching point pairs are obtained. By matching semantic information and feature description information, the matching range is narrowed and detection efficiency is improved.
Smart Images

Figure CN120125854A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time positioning and mapping, and in particular, to a visual SLAM loop detection method, device, equipment and storage medium. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) refers to a technology that, during the movement in an unknown environment, performs positioning based on its own measurement data and external observation data, and simultaneously performs incremental mapping. It is widely used in fields such as autonomous driving and intelligent robots. Loop detection in the SLAM system is extremely important, as loop detection can eliminate the cumulative error of the SLAM system. Currently, traditional visual SLAM loop detection mainly searches for the spatial coordinate points corresponding to the key frame that is most similar to the current key frame in the global visual map, which takes a relatively long time, and there are relatively few spatial coordinate points in the global visual map that can be compared with the current key frame, resulting in a low accuracy of visual SLAM loop detection. Summary of the Invention
[0003] Embodiments of the present invention provide a visual SLAM loop detection method, device, equipment and storage medium, aiming to improve the speed and accuracy of visual SLAM loop detection.
[0004] In a first aspect, embodiments of the present invention provide a visual SLAM loop detection method, including:
[0005] Obtain the current key frame as a matching frame;
[0006] Determine a target map subgraph from a visual map pre-constructed with M map subgraphs, where M is an integer greater than 0, and the map subgraph includes map points corresponding to each feature point within multiple key frames;
[0007] According to the semantic information of the matching frame and the semantic information of the target map subgraph, determine a set of target map points that match the matching frame from the target map subgraph;
[0008] According to the feature description information of the matching frame and the feature description information of the set of target map points, determine the target map points corresponding to at least some feature points in the matching frame from the set of target map points.
[0009] In a second aspect, embodiments of the present invention further provide a visual SLAM loop detection device, including:
[0010] An acquisition module, configured to obtain the current key frame as a matching frame;
[0011] A map sub - graph determination module, configured to determine a target map sub - graph from a pre - constructed visual map including M map sub - graphs, where M is an integer greater than 0, and the map sub - graph includes map points corresponding to respective feature points in a plurality of key frames;
[0012] A semantic matching module, configured to determine a set of target map points matching the matching frame from the target map sub - graph according to the semantic information of the matching frame and the semantic information of the target map sub - graph;
[0013] A feature matching module, configured to determine target map points corresponding to at least some feature points in the matching frame from the set of target map points according to the feature description information of the matching frame and the feature description information of the set of target map points.
[0014] In a third aspect, an embodiment of the present invention further provides a movable electronic device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the visual SLAM loop detection method described in the first aspect is implemented.
[0015] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer - readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the visual SLAM loop detection method described in the first aspect.
[0016] An embodiment of the present invention provides a visual SLAM loop detection method, device, equipment and storage medium. The map sub - graph in the embodiment of the present invention includes map points corresponding to respective feature points in a plurality of key frames, that is, there are more map points in the map sub - graph, while there are fewer spatial coordinate points in the traditional loop detection. Therefore, in the embodiment of the present invention, by matching the matching frame with the map sub - graph, more matching point pairs (the matching point pairs include feature points and map points matching the feature points) can be obtained, improving the accuracy of loop detection. And when the matching frame is matched with the map sub - graph, the semantic information of the matching frame and the semantic information of the map sub - graph are used to screen out a set of target map points that are roughly matched with the matching frame from the map sub - graph, narrowing the matching range. Then, the feature description information is used to find map points that match at least some feature points in the matching frame from the set of target map points (the narrowed matching range), improving the speed of loop detection. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a visual SLAM loop detection method provided by an embodiment of the present invention;
[0019] Figure 2 is Figure 1 a schematic sub-step flowchart of the visual SLAM loop detection method in
[0020] Figure 3 It is a schematic flowchart of another visual SLAM loop detection method provided by an embodiment of the present invention;
[0021] Figure 4 It is a schematic flowchart of another visual SLAM loop detection method provided by an embodiment of the present invention;
[0022] Figure 5 It is a schematic block diagram of the structure of a visual SLAM loop detection device provided by an embodiment of the present invention;
[0023] Figure 6 It is a schematic block diagram of the structure of a movable electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0025] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0026] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0027] It should be noted that Simultaneous Localization and Mapping (SLAM) refers to the process in which a moving object calculates its own position while constructing an environmental map based on sensor information, solving the problems of localization and map construction for a movable electronic device when moving in an unknown environment. SLAM is mainly applied to fields such as robots, drones, autonomous driving, AR glasses, VR glasses, etc. SLAM technology is divided into two categories: if the sensor is a lidar, it is called lidar SLAM; if the sensor is a camera, it is called visual SLAM.
[0028] Visual SLAM can be divided into four parts: the front end, the back end, mapping, and loop detection. The main function of the front end is to calculate the relative position relationship of the corresponding camera poses between image frames. The front end includes parts such as feature extraction, feature matching, and calculating the pose using the matched features. The function of the back end is mainly to optimize the output result of the front end to obtain the optimal pose estimation. Loop detection, also known as closed-loop detection, refers to the ability of a robot to recognize a scene it has reached. Loop detection provides the association between current data and all historical data. After the tracking algorithm is lost, loop detection can also be used for relocalization. Currently, loop detection in visual SLAM mainly searches for the spatial coordinate points corresponding to the key frame most similar to the current key frame in the global visual map, which takes a lot of time, and there are relatively few spatial coordinate points in the global visual map that can be compared with the current key frame, so the accuracy of loop detection in visual SLAM is relatively low.
[0029] To solve the above problems, the embodiments of the present invention provide a method, device, equipment, and storage medium for visual SLAM loop detection. The map subgraph in the embodiments of the present invention includes map points corresponding to each feature point in multiple key frames, that is, there are more map points in the map subgraph, while there are relatively few spatial coordinate points in traditional loop detection. Therefore, by matching the matching frame with the map subgraph in the embodiments of the present invention, more matching point pairs (matching point pairs include feature points and map points matching the feature points) can be obtained, improving the accuracy of loop detection. When the matching frame is matched with the map subgraph, the semantic information of the matching frame and the semantic information of the map subgraph are used to screen out a set of target map points that are roughly matched with the matching frame from the map subgraph, narrowing the matching range, and then the feature description information is used to find map points that match at least some of the feature points in the matching frame from the set of target map points (the narrowed matching range), improving the speed of loop detection.
[0030] An embodiment of the present invention provides a visual SLAM loop detection method that can be used in a movable electronic device, which can be a device such as a mobile phone, a tablet computer, a laptop computer, smart glasses, a robot, etc. Of course, the visual SLAM loop detection method can also be applied to a server. The movable electronic device collects key frames and sends the collected key frames to the server, and the server executes the visual SLAM loop detection method provided by the embodiment of the present invention according to the key frames. Among them, the server can be a separate server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0031] The following will describe in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0032] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a visual SLAM loop detection method provided by an embodiment of the present invention.
[0033] As Figure 1 shown, the visual SLAM loop detection method includes steps S101 to S103.
[0034] Step S101: Obtain the current key frame as a matching frame.
[0035] In this embodiment, the current key frame is a key frame collected by a visual sensor on the movable electronic device at the current moment. A key frame refers to a key image frame in a video, which can be used to represent important events, actions, and transitions in the video. Among them, the visual sensor mainly includes several parts such as optical devices, image sensors, and digital signal processors.
[0036] In some embodiments, obtaining the current key frame as a matching frame may include: obtaining the previous trigger time of loop detection and obtaining the number of key frames detected after the previous trigger time; when the number of detected key frames reaches a preset number threshold, obtaining the current key frame as a matching frame. Among them, the preset number threshold can be set based on the actual situation, and the embodiment of the present invention does not make specific limitations on this. For example, the preset number threshold is any integer from 2 to 10. In this embodiment, by triggering the visual SLAM loop detection when the number of key frames detected after the previous trigger time of loop detection reaches the set number threshold, the visual SLAM loop detection can be performed periodically, improving the accuracy of the visual SLAM loop detection.
[0037] In some embodiments, obtaining the current key frame as the matching frame may include: obtaining the previous trigger time of loop detection and the current system time; when the absolute value of the difference between the previous trigger time and the current system time reaches a preset duration threshold, obtaining the current key frame as the matching frame. Among them, the preset duration threshold can be set based on the actual situation, and the embodiments of the present invention do not make specific limitations on this. For example, the preset duration threshold is 5 seconds or 8 seconds, etc. In this embodiment, by triggering the visual SLAM loop detection when the absolute value of the difference between the previous trigger time of loop detection and the current system time reaches the preset duration threshold, the visual SLAM loop detection can be performed periodically, improving the accuracy of the visual SLAM loop detection.
[0038] In some embodiments, obtaining the current key frame as the matching frame may include: obtaining N consecutive key frames as the matching frame, where the N consecutive key frames at least include the current key frame, and N is an integer greater than or equal to 2. For example, taking the current key frame and the previous key frame, a total of two key frames as the matching frame, or taking the current key frame and the previous two key frames, a total of three key frames as the matching frame. In this embodiment, by obtaining multiple consecutive key frames as the matching frame, and the multiple key frames at least include the current key frame, more feature points are included in the matching frame. In this way, when matching with the map subgraph subsequently, more matching point pairs (the matching point pairs include feature points and map points matching the feature points) can be obtained, further improving the accuracy of loop detection.
[0039] Step S102, determine the target map subgraph from the visual map including M map subgraphs pre-constructed.
[0040] In this embodiment, M is an integer greater than 0, that is, the visual map includes at least one map subgraph, and the map subgraph includes map points corresponding to each feature point in multiple key frames.
[0041] In some embodiments, determining the target map subgraph from the visual map including M map subgraphs pre-constructed may include: randomly selecting a map subgraph from the visual map including M map subgraphs pre-constructed as the target map subgraph, where the map subgraph selected each time loop detection is performed is different, and after M times of loop detection, each of the M map subgraphs is selected once. In this embodiment, by selecting different map subgraphs each time loop detection is performed, it is ensured that each map subgraph in the visual map can be selected once within a period of time, making the loop detection more comprehensive and further improving the accuracy of loop detection.
[0042] In some embodiments, determining a target map sub - graph from a pre - constructed visual map containing M map sub - graphs may include: obtaining the target pose of the matching frame, and selecting, from the pre - constructed visual map containing M map sub - graphs, the map sub - graph that matches the target pose as the target map sub - graph. In this embodiment, by selecting the map sub - graph that matches the target pose of the matching frame as the target map sub - graph, the relevance between the target map sub - graph and the matching frame can be ensured. In this way, after subsequent matching between the matching frame and the target map sub - graph, more matching point pairs (matching point pairs include feature points and map points matching the feature points) can be obtained, further improving the accuracy of loop detection.
[0043] In some embodiments, obtaining the target pose of the matching frame may include: when the matching frame only includes the current key frame, determining the pose of the current key frame as the target pose; when the matching frame includes multiple key frames, calculating the average pose according to the poses of each key frame included in the matching frame, and determining the average pose as the target pose. Here, the pose refers to the position coordinates and attitude angles of an object, robot, or person in a specified coordinate system. The attitude angles may include the pitch angle, heading angle, and / or roll angle of the object, robot, or person in the specified coordinate system. In practical applications, a movable electronic device equipped with a visual sensor can control the visual sensor to capture the external environment from different angles and positions during movement to obtain multiple key frames with different poses.
[0044] In some embodiments, calculating the average pose according to the poses of each key frame included in the matching frame may include: calculating the average position coordinates according to the position coordinates of each key frame included in the matching frame; calculating the average attitude angle according to the attitude angles of each key frame included in the matching frame; and determining the average position coordinates and the average attitude angle as the average pose.
[0045] In some embodiments, selecting, from the pre - constructed visual map containing M map sub - graphs, the map sub - graph that matches the target pose as the target map sub - graph may include: determining the similarity between the target pose and the pose corresponding to each map sub - graph in the visual map, and selecting the map sub - graph corresponding to the maximum similarity in the visual map as the target map sub - graph.
[0046] Step S103: Determine a set of target map points that match the matching frame from the target map sub - graph according to the semantic information of the matching frame and the semantic information of the target map sub - graph.
[0047] In this embodiment, by using the semantic information of the matching frame and the semantic information of the target map sub - graph, a set of target map points that are roughly matched with the matching frame is filtered out from the target map sub - graph, narrowing the matching range. This facilitates subsequent searching for map points that match each feature point in the matching frame within the narrowed matching range, improving the speed of loop detection.
[0048] In some embodiments, the semantic information of the matching frame may include first semantic labels of respective feature points within the matching frame, and the semantic information of the target map subgraph may include second semantic labels of respective map points within the target map subgraph. Determining a set of target map points matching the matching frame from the target map subgraph according to the semantic information of the matching frame and the semantic information of the target map subgraph may include: comparing the first semantic labels of the respective feature points within the matching frame with the second semantic labels of the respective map points in the target map subgraph to obtain a semantic label comparison result; and determining a set of target map points matching the matching frame from the target map subgraph according to the semantic label comparison result, wherein the second semantic label of each map point in the set of target map points is the same as the first semantic label. It can be understood that the first semantic label and the second semantic label can be set based on actual situations, and the embodiments of the present invention do not make specific limitations thereto. For example, in a home scenario, the first semantic label and / or the second semantic label may include a table, a chair, a book, a trash, a person, a pet, a trash can, a door, a lamp, etc.
[0049] In some embodiments, as Figure 2 shown, step S103 includes: sub-steps S1031 to S1032.
[0050] Sub-step S1031, aligning the map points corresponding to the respective feature points within the matching frame into the target map subgraph according to the rotation matrix corresponding to the target map subgraph to obtain a first set of map points.
[0051] In this embodiment, each map subgraph has a corresponding coordinate system, different map subgraphs correspond to different rotation matrices, and the rotation matrix is the rotation matrix of the inertial navigation unit. For example, through the rotation matrix corresponding to the target map subgraph, the first spatial coordinates of the map points corresponding to the respective feature points within the matching frame can be converted into second spatial coordinates in the coordinate system corresponding to the target map subgraph, and at this time, the map points corresponding to each second spatial coordinate constitute the first set of map points.
[0052] Sub-step S1032, determining a set of target map points matching the matching frame from the first set of map points according to the semantic information of the matching frame and the semantic information of the first set of map points.
[0053] In this embodiment, through the map points corresponding to the respective feature points within the matching frame and the rotation matrix corresponding to the target map subgraph, a set of map points roughly matching the matching frame can be screened out from the target map subgraph, and then through the semantic information of the matching frame and the semantic information of the set of map points obtained by the rough matching, a set of target map points with a smaller range can be further matched from the set of map points obtained by the rough matching, which can improve the speed of loop detection while ensuring the accuracy of loop detection.
[0054] In some embodiments, the semantic information of the matching frame includes the first semantic labels of the respective feature points within the matching frame, and the semantic information of the first map point set includes the second semantic labels of the respective map points within the first map point set. For example, the first semantic labels of the respective feature points within the matching frame are compared with the second semantic labels of the respective map points within the first map point set to obtain a semantic label comparison result; based on this semantic label comparison result, a target map point set that matches the matching frame is determined from the first map point set, where the second semantic label of each map point in the target map point set is the same as the first semantic label.
[0055] In some embodiments, the semantic label comparison result includes the comparison result between the second semantic labels of the respective map points within the first map point set and the first semantic labels of the respective feature points within the matching frame, and this comparison result includes a first comparison result or a second comparison result. The first comparison result is that the second semantic label of the map point is the same as the first semantic label of the feature point, and the second comparison result is that the second semantic label of the map point is different from the first semantic label of the feature point. Based on this semantic label comparison result, determining a target map point set that matches the matching frame from the first map point set may include: screening out the map points corresponding to the first comparison result from the first map point set to form a target map point set that matches the matching frame.
[0056] In some embodiments, determining a target map point set that matches the matching frame from the first map point set according to the semantic information of the matching frame and the semantic information of the first map point set may include: obtaining the map points located within the viewing angle range corresponding to the matching frame from the first map point set to form a second map point set; and determining a target map point set that matches the matching frame from the second map point set according to the semantic information of the matching frame and the semantic information of the second map point set. In this embodiment, by obtaining the map points located within the viewing angle range corresponding to the matching frame from the first map point set, a second map point set with a smaller range can be obtained, and then by using the semantic information of the matching frame and the semantic information of the second map point set, a target map point set with a further reduced range can be further matched from the second map point set, which can improve the speed of loop detection while ensuring the accuracy of loop detection.
[0057] In some embodiments, when the matching frame only includes the current key frame, the viewing angle range corresponding to the matching frame is the viewing angle range corresponding to the current key frame; when the matching frame includes multiple key frames, the average viewing angle range is calculated according to the viewing angle range corresponding to each key frame included in the matching frame, and this average viewing angle range is determined as the viewing angle range corresponding to the matching frame. Among them, the viewing angle range corresponding to the matching frame includes at least one of the heading angle range, roll angle range, and pitch angle range corresponding to the matching frame.
[0058] Step S104: Determine the target map points corresponding to at least some of the feature points in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set.
[0059] In the embodiments of the present invention, the map sub-graph includes the map points corresponding to the feature points in multiple key frames, that is, there are more map points in the map sub-graph, while there are fewer spatial coordinate points in the traditional loop detection. Therefore, by matching the matching frame with the map sub-graph in the embodiments of the present invention, more matching point pairs (the matching point pairs include feature points and the map points matching the feature points) can be obtained, improving the accuracy of loop detection. When matching the matching frame with the map sub-graph, the semantic information of the matching frame and the semantic information of the map sub-graph are used to screen out the target map point set that is roughly matched with the matching frame from the map sub-graph, narrowing the matching range. Then, the feature description information is used to find the map points that match at least some of the feature points in the matching frame from the target map point set (the narrowed matching range), thus improving the speed of loop detection.
[0060] In some embodiments, the feature description information of the matching frame includes the brief descriptors of the feature points in the matching frame, and the feature description information of the target map point set includes the brief descriptors of the feature points corresponding to the map points in the target map point set. Determining the target map points corresponding to at least some of the feature points in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set may include: for each feature point in the matching frame, determine the similarity between the feature point and the feature points corresponding to the map points in the target map point set according to the brief descriptor of the feature point and the brief descriptors of the feature points corresponding to the map points in the target map point set. If the maximum similarity is greater than or equal to the preset similarity threshold, then determine the map point corresponding to the maximum similarity as the target map point matching the feature point. If the maximum similarity is less than the preset similarity threshold, then determine that there is no map point in the target map point set that matches the feature point.
[0061] In some embodiments, based on a preset nearest neighbor matching algorithm, determine the target map points corresponding to at least some of the feature points in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set. Among them, the preset nearest neighbor matching algorithm may include the Kt-Tree nearest neighbor matching algorithm, the DBoW algorithm, the DBoW2 algorithm, the DBoW3 algorithm, or the FBoW algorithm, etc. In this embodiment, the nearest neighbor matching algorithm is used to match the feature points in the matching frame with the map points in the target map point set, and it can quickly determine the target map points corresponding to at least some of the feature points in the matching frame from the target map point set, further improving the speed of loop detection.
[0062] In some embodiments, such as Figure 3 shown, after step S104, it further includes:
[0063] Step S105, optimize the pose of the current key frame according to each feature point in the matching frame and the target map points corresponding to at least some of the feature points in the matching frame.
[0064] For example, obtain the feature points bound to each target map point, and determine the difference between the feature points bound to each target map point and the corresponding feature points of the target map point in the matching frame; compare each of these differences with each other to determine the largest difference, and correct the largest difference through posegraph (graph optimization) to optimize the pose of the current key frame. This embodiment improves the accuracy of pose optimization.
[0065] In some embodiments, such as Figure 4 shown, before step S101, it further includes:
[0066] Step S106, obtain the image frames captured by the visual sensor at each pose.
[0067] In this embodiment, the visual sensor is a sensor that can simulate the human visual system. It can convert optical signals into digital signals, thereby realizing functions such as object recognition, tracking, and measurement. The visual sensor mainly includes several parts such as optical devices, image sensors, and digital signal processors. In addition, the visual sensor is an important device for visual SLAM map construction. Pose refers to the position and orientation of an object, robot, or person in a specified coordinate system. Obviously, the positions and / or orientations corresponding to different poses are also different. In practical applications, a movable electronic device equipped with a visual sensor can control the visual sensor to capture images of the external environment from different angles and positions to obtain multiple image frames with different poses.
[0068] Step S107, divide the multiple image frames into at least one image group according to the poses corresponding to each of the multiple image frames, and the ranges where the poses corresponding to the image frames in each image group are located satisfy a preset condition.
[0069] In this embodiment, the image frames include key frames. A key frame refers to a key image frame in a video, which can be used to represent important events, actions, and transitions in the video. An image frame serving as a key frame generally has a corresponding pose. It can be understood that the preset conditions satisfied by the ranges where the poses corresponding to the image frames in each image group are located can be the same or different, and the embodiments of the present invention do not make specific limitations on this.
[0070] For example, among multiple image frames obtained by indoor shooting, multiple first image frames correspond to the pantry on the right, multiple second image frames correspond to the corridor in front, and multiple third image frames correspond to the office area on the left. Based on this, classifying image frames whose poses are within a range that meets a preset condition into the same image group may include: classifying multiple first image frames belonging to the pantry into one image group, classifying multiple second image frames belonging to the corridor into one image group, and classifying multiple third image frames belonging to the office area into one image group. Another example is that the range where the poses corresponding to the image frames in each image group meet the preset condition includes: the pose deviation between any two image frames in each image group is within a preset pose deviation range. Among them, the preset pose deviation range corresponding to each image group is the same, and the pose deviation between any two image frames is the absolute value of the difference between the poses of any two image frames.
[0071] It can be understood that the pose corresponding to the image frame, that is, the attitude and the position, can be regarded as an area, and the range of the area where the poses of each image frame in the same image group are located should meet the preset condition. Specifically, the indoor area includes multiple rooms at different positions, such as rooms in the four directions of southeast, northwest. Therefore, the poses corresponding to the image frames obtained by shooting different rooms are generally different. At this time, the spatial range of the room can be used as the preset condition. If the range where the pose corresponding to the image frame is located meets this preset condition, the image frames that meet this preset condition are classified into the same image group.
[0072] Step S108: Construct a map subgraph corresponding to the image group according to the poses and feature points corresponding to the image frames in the image group. The visual map includes at least one map subgraph corresponding to the image group.
[0073] In this embodiment, each map subgraph has a corresponding coordinate system. During the process of constructing the map subgraph, the three-dimensional space corresponding to the image frame can be determined according to the pose corresponding to each image frame, that is, the position and the attitude. Then, the spatial coordinates of the feature points corresponding to the feature points on each image frame in the three-dimensional space, that is, the map points, can be determined according to the brief descriptors corresponding to the feature points on each image frame. Thus, the map subgraph can be constructed according to the three-dimensional space corresponding to the pose of the image frame and the map points corresponding to the feature points of the image frame. Based on this, the feature points on the image frames in the same image group are bound to the corresponding feature points of the map points in the map subgraph, that is, the feature points on the image frames are bound to the corresponding map points in the map subgraph, so that the map subgraph corresponding to the same image group can be constructed. Finally, the map subgraphs corresponding to multiple image groups can form the overall visual map.
[0074] For example, after constructing different map subgraphs, more feature points of the image frame can be extracted to determine map points based on the brief descriptors corresponding to these feature points, and the map points can be bound to the corresponding map subgraphs to increase the map points of each map subgraph. Since the number of map points increases, there are more map points of the map subgraphs that can be compared during loop detection, thus making the constructed map subgraphs more accurate and improving the accuracy of pose optimization.
[0075] In some embodiments, after step S108, it further includes: determining the matching degree between any two map subgraphs in the visual map according to the semantic information of each map subgraph in the visual map; if the matching degree between at least two map subgraphs is greater than a preset matching degree threshold, then deleting at least one of the at least two map subgraphs that is earlier in the construction time sequence from the visual map. In this embodiment, by retaining one of the similar multiple map subgraphs in the visual map and deleting other map subgraphs with a matching degree greater than the preset matching degree threshold, the data redundancy brought by the similar map subgraphs is reduced, the operating memory and storage space of the device are improved, so as to ensure the smooth operation of the device, and the deleted map subgraph is the one earlier in the construction time sequence, and one map subgraph with the latest construction time sequence is retained, so as to improve the accuracy of constructing the visual map on the premise of ensuring the smooth operation of the device.
[0076] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of the structure of a visual SLAM loop detection device provided by an embodiment of the present invention.
[0077] As Figure 5 shown, the visual SLAM loop detection device 100 includes:
[0078] An acquisition module 110, configured to acquire the current key frame as a matching frame;
[0079] A map subgraph determination module 120, configured to determine a target map subgraph from a pre-constructed visual map including M map subgraphs, where M is an integer greater than 0, and the map subgraph includes map points corresponding to each feature point in multiple key frames;
[0080] A semantic matching module 130, configured to determine a target map point set matching the matching frame from the target map subgraph according to the semantic information of the matching frame and the semantic information of the target map subgraph;
[0081] A feature matching module 140, configured to determine the target map points corresponding to each feature point in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set.
[0082] In some embodiments, the semantic matching module 130 includes:
[0083] A map point alignment sub-module, configured to align the map points corresponding to the feature points in the matching frame into the target map sub-graph according to the rotation matrix corresponding to the target map sub-graph, so as to obtain a first map point set;
[0084] A semantic matching sub-module, configured to determine a target map point set that matches the matching frame from the first map point set according to the semantic information of the matching frame and the semantic information of the first map point set.
[0085] In some embodiments, the semantic matching sub-module is further configured to:
[0086] Obtain the map points located within the viewing range corresponding to the matching frame from the first map point set to form a second map point set;
[0087] Determine a target map point set that matches the matching frame from the second map point set according to the semantic information of the matching frame and the semantic information of the second map point set.
[0088] In some embodiments, the map sub-graph determination module 120 is further configured to:
[0089] Obtain the target pose of the matching frame, and select the map sub-graph that matches the target pose from the pre-constructed visual map including M map sub-graphs as the target map sub-graph;
[0090] Alternatively, randomly select one of the map sub-graphs from the pre-constructed visual map including M map sub-graphs as the target map sub-graph, where the map sub-graphs selected each time for loop detection are different, and after M times of loop detection, the M map sub-graphs are all selected once.
[0091] In some embodiments, the acquisition module 110 is further configured to:
[0092] Acquire N consecutive key frames as matching frames, where the N key frames at least include the current key frame, and N is an integer greater than or equal to 2.
[0093] In some embodiments, the visual SLAM loop detection device 100 further includes:
[0094] An image frame acquisition module, configured to acquire the image frames captured by the visual sensor at each pose;
[0095] An image division module, configured to divide the multiple image frames into at least one image group according to the poses corresponding to the multiple image frames, and the ranges where the poses corresponding to the image frames in each image group are located meet preset conditions;
[0096] A map construction module, configured to construct a map subgraph corresponding to the image group according to the poses and feature points corresponding to each image frame in the image group, where the visual map includes map subgraphs corresponding to the at least one image group.
[0097] In some embodiments, the visual SLAM loop detection device 100 further includes:
[0098] A matching degree determination module, configured to determine the matching degree between any two map subgraphs in the visual map according to the semantic information of each map subgraph in the visual map;
[0099] A deletion module, configured to, if the matching degree between at least two map subgraphs is greater than a preset matching degree threshold, delete at least one of the at least two map subgraphs with an earlier construction time sequence from the visual map.
[0100] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described visual SLAM loop detection device can refer to the corresponding process in the foregoing embodiment of the visual SLAM loop detection method, and will not be described herein again.
[0101] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of the structure of a movable electronic device provided by an embodiment of the present invention.
[0102] As Figure 6 shown, the movable electronic device 200 includes a processor 201 and a memory 202, and the processor 201 and the memory 202 are connected through a bus 203, and the bus is, for example, an I2C (Inter-integrated Circuit) bus.
[0103] Specifically, the processor 201 is used to provide computing and control capabilities to support the operation of the entire mobile electronic device. The processor 201 can be a Central Processing Unit (CPU), and it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0104] Specifically, the memory 202 can be a Flash chip, Read-Only Memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.
[0105] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the embodiment of the present invention, and does not constitute a limitation on the mobile electronic device to which the solution of the embodiment of the present invention is applied. The specific mobile electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0106] Among them, the processor 201 is used to run the computer program stored in the memory 202 and implement any one of the visual SLAM loop detection methods provided by the embodiments of the present invention when executing the computer program.
[0107] In one embodiment, the processor 201 is used to run the computer program stored in the memory and implement the following steps when executing the computer program:
[0108] Obtain the current key frame as the matching frame;
[0109] Determine the target map subgraph from the pre-constructed visual map including M map subgraphs, where M is an integer greater than 0, and the map subgraph includes map points corresponding to each feature point in multiple key frames;
[0110] According to the semantic information of the matching frame and the semantic information of the target map subgraph, determine the set of target map points matching the matching frame from the target map subgraph;
[0111] Determine the target map points corresponding to at least some of the feature points in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set.
[0112] In some embodiments, when the processor 201 implements determining a target map point set matching the matching frame from the target map sub-graph according to the semantic information of the matching frame and the semantic information of the target map sub-graph, it is configured to implement:
[0113] Align the map points corresponding to the respective feature points within the matching frame into the target map sub-graph according to the rotation matrix corresponding to the target map sub-graph to obtain a first map point set;
[0114] Determine a target map point set matching the matching frame from the first map point set according to the semantic information of the matching frame and the semantic information of the first map point set.
[0115] In some embodiments, when the processor 201 implements determining a target map point set matching the matching frame from the first map point set according to the semantic information of the matching frame and the semantic information of the first map point set, it is configured to implement:
[0116] Obtain the map points within the viewing range corresponding to the matching frame from the first map point set to form a second map point set;
[0117] Determine a target map point set matching the matching frame from the second map point set according to the semantic information of the matching frame and the semantic information of the second map point set.
[0118] In some embodiments, when the processor 201 implements determining a target map sub-graph from a pre-constructed visual map including M map sub-graphs, it is configured to implement any one of the following:
[0119] Obtain the target pose of the matching frame, and select the map sub-graph matching the target pose from the pre-constructed visual map including M map sub-graphs as the target map sub-graph;
[0120] Randomly select one of the map sub-graphs from the pre-constructed visual map including M map sub-graphs as the target map sub-graph, where the selected map sub-graph is different each time loop detection is performed, and after M times of loop detection, each of the M map sub-graphs is selected once.
[0121] In some embodiments, when the processor 201 implements obtaining the current key frame as the matching frame, it is configured to implement:
[0122] Obtain N consecutive key frames as matching frames, where the N key frames at least include the current key frame, and N is an integer greater than or equal to 2.
[0123] In some embodiments, before the processor 201 implements obtaining the current key frame as a matching frame, it is further configured to implement:
[0124] Obtain the image frames captured by the visual sensor at each pose;
[0125] According to the poses corresponding to the multiple image frames, divide the multiple image frames into at least one image group, and the range where the poses corresponding to the image frames in each image group are located satisfies a preset condition;
[0126] Construct a map subgraph corresponding to the image group according to the poses and feature points corresponding to the image frames in the image group, and the visual map includes the map subgraphs corresponding to the at least one image group.
[0127] In some embodiments, after the processor 201 implements constructing the map subgraph corresponding to the image group according to the poses and feature points corresponding to the image frames in the image group, it is further configured to implement:
[0128] Determine the matching degree between any two map subgraphs in the visual map according to the semantic information of each map subgraph in the visual map;
[0129] If the matching degree between at least two map subgraphs is greater than a preset matching degree threshold, then delete at least one of the at least two map subgraphs with an earlier construction time sequence from the visual map.
[0130] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described movable electronic device can refer to the corresponding process in the embodiment of the foregoing visual SLAM loop detection method, and will not be repeated here.
[0131] The embodiment of the present invention further provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any visual SLAM loop detection method provided in the specification of the embodiment of the present invention.
[0132] Among them, the storage medium may be the internal storage unit of the removable electronic device described in the foregoing embodiments, such as the hard disk or memory of the removable electronic device. The storage medium may also be an external storage device of the removable electronic device, such as a plug-in hard disk equipped on the removable electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0133] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.
[0134] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or system. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or system comprising the element.
[0135] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A visual SLAM loop detection method, characterized in that, it includes: Obtain the current key frame as the matching frame; Determine the target map sub - graph from a pre - constructed visual map containing M map sub - graphs, where M is an integer greater than 0, and the map sub - graph includes map points corresponding to each feature point in multiple key frames; According to the semantic information of the matching frame and the semantic information of the target map sub - graph, determine a set of target map points in the target map sub - graph that match the matching frame; According to the feature description information of the matching frame and the feature description information of the set of target map points, determine the target map points corresponding to at least some feature points in the matching frame from the set of target map points.
2. The visual SLAM loop detection method according to claim 1, characterized in that, The step of determining a set of target map points in the target map sub - graph that match the matching frame according to the semantic information of the matching frame and the semantic information of the target map sub - graph includes: According to the rotation matrix corresponding to the target map sub - graph, align the map points corresponding to each feature point in the matching frame into the target map sub - graph to obtain a first set of map points; According to the semantic information of the matching frame and the semantic information of the first set of map points, determine a set of target map points in the first set of map points that match the matching frame.
3. The visual SLAM loop detection method according to claim 2, characterized in that, The step of determining a set of target map points in the first set of map points that match the matching frame according to the semantic information of the matching frame and the semantic information of the first set of map points includes: Obtain the map points within the viewing angle range corresponding to the matching frame from the first set of map points to form a second set of map points; According to the semantic information of the matching frame and the semantic information of the second set of map points, determine a set of target map points in the second set of map points that match the matching frame.
4. The visual SLAM loop detection method according to claim 1, characterized in that, The step of determining the target map sub - graph from a pre - constructed visual map containing M map sub - graphs includes any one of the following: Obtain the target pose of the matching frame, and select the map sub - graph in the pre - constructed visual map containing M map sub - graphs that matches the target pose as the target map sub - graph; Randomly select one of the map sub - graphs in the pre - constructed visual map containing M map sub - graphs as the target map sub - graph, where the selected map sub - graph is different each time loop detection is performed, and after M times of loop detection, each of the M map sub - graphs is selected once.
5. The visual SLAM loop detection method according to claim 1, characterized in that, The step of obtaining the current key frame as the matching frame includes: Obtain N consecutive key frames as the matching frame, where the N key frames at least include the current key frame, and N is an integer greater than or equal to 2.
6. The visual SLAM loop detection method according to any one of claims 1 - 5, characterized in that, Before obtaining the current key frame as the matching frame, it further includes: Obtain the image frames captured by the visual sensor at each pose; According to the poses corresponding to each of the multiple image frames, divide the multiple image frames into at least one image group, and the range where the poses corresponding to the image frames in each image group are located satisfies a preset condition; According to the poses and feature points corresponding to the image frames in the image group, construct a map sub-graph corresponding to the image group, and the visual map includes the map sub-graphs corresponding to the at least one image group.
7. The visual SLAM loop detection method according to claim 6, wherein, after constructing the map sub-graph corresponding to the image group according to the poses and feature points corresponding to the image frames in the image group, further includes: Determine the matching degree between any two map sub-graphs in the visual map according to the semantic information of each map sub-graph in the visual map; If the matching degree between at least two map sub-graphs is greater than a preset matching degree threshold, delete at least one of the at least two map sub-graphs with an earlier construction time sequence from the visual map.
8. A visual SLAM loop detection device, wherein, includes: An acquisition module, configured to acquire the current key frame as a matching frame; A map sub-graph determination module, configured to determine a target map sub-graph from a pre-constructed visual map including M map sub-graphs, where M is an integer greater than 0, and the map sub-graph includes map points corresponding to each feature point in multiple key frames; A semantic matching module, configured to determine a target map point set matching the matching frame from the target map sub-graph according to the semantic information of the matching frame and the semantic information of the target map sub-graph; A feature matching module, configured to determine the target map points corresponding to the feature points in the matching frame from the target map point set according to the feature description information of the matching frame and the feature description information of the target map point set.
9. A movable electronic device, wherein, the movable electronic device includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing the connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, it realizes the visual SLAM loop detection method according to any one of claims 1 to 7.
10. A storage medium for computer-readable storage, wherein, the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the visual SLAM loop detection method according to any one of claims 1 to 7.