Loop closure detection method and computer device
By acquiring feature matching and pose optimization from image frames in commercial service robots, the accuracy problem of loop closure detection was solved, improving the accuracy of map construction and the precision of robot navigation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHENZHEN PUDU TECH CO LTD
- Filing Date
- 2025-12-12
- Publication Date
- 2026-07-30
AI Technical Summary
In complex environments, commercial service robots may fail to accurately detect loopback poses or detect incorrect loopback poses, resulting in significant errors in map building and impacting the robot's subsequent operations.
By acquiring multiple image frames collected by the robot, the feature matching degree between the last image frame and its preceding image frames is determined. The image frame with the highest matching degree is selected as the target for loop closure detection. Loop closure detection is performed in combination with distance conditions, and the pose of the image frames is optimized to improve detection accuracy.
It significantly improves the accuracy of loop closure detection, ensures the accuracy of map construction, reduces cumulative errors, and improves the precision of robot obstacle avoidance and navigation.
Smart Images

Figure CN2025142130_30072026_PF_FP_ABST
Abstract
Description
Loop closure detection methods and computer equipment
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202510124895.8, filed on January 24, 2025, entitled "Loop Detection Method and Computer Device", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of robotics, and in particular to a loop closure detection method and computer equipment. Background Technology
[0004] The statements herein are provided only as background information in connection with this application and do not necessarily constitute exemplary technology.
[0005] With the rapid development of robotics technology, Visual Simultaneous Localization and Mapping (VSLAM) technology is being used more and more widely in commercial service robots. Consequently, VSLAM technology is encountering more and more challenges in its application. Visual loop closure detection is an essential task in order for robots to build accurate, low-error visual maps.
[0006] In indoor operating scenarios, commercial service robots are typically equipped with monocular cameras that point vertically upwards to capture images of the ceiling above the robot. In complex environments, such as large factories and office buildings, the ceiling environments often exhibit high similarity. Therefore, if loop closure poses cannot be accurately detected or incorrect loop closure poses are detected in these complex environments, the map constructed by the robot will have significant errors. This will undoubtedly have a serious adverse impact on the subsequent business operations performed by the commercial service robot. Summary of the Invention
[0007] According to various embodiments of this application, a loop closure detection method, apparatus, computer device, computer-readable storage medium, and computer program product are provided.
[0008] Firstly, this application provides a loop closure detection method, including:
[0009] Acquire multiple image frames collected by the robot, and determine the last image frame and a first preset number of first preceding image frames among the multiple image frames;
[0010] Determine the feature matching degree between each first preceding image frame and the last image frame, and determine the first target image frame corresponding to the last image frame among a first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree with the last image frame among the first preset number of first preceding image frames.
[0011] Among multiple image frames, determine the second target image frame corresponding to the first target image frame corresponding to the last image frame; the second target image frame corresponding to the last image frame is the image frame among multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies the first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame.
[0012] Based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, loop closure detection is performed in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
[0013] Secondly, this application also provides a loop closure detection device, comprising:
[0014] The acquisition module is used to acquire multiple image frames collected by the robot, and to determine the last image frame and a first preset number of first preceding image frames among the multiple image frames.
[0015] The first determining module is used to determine the feature matching degree between each first preceding image frame and the last image frame, and to determine the first target image frame corresponding to the last image frame among a first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree with the last image frame among the first preset number of first preceding image frames.
[0016] The second determining module is used to determine, among multiple image frames, the second target image frame corresponding to the first target image frame corresponding to the last image frame; the second target image frame corresponding to the last image frame is the image frame among the multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies the first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame.
[0017] The detection module is used to perform loop closure detection in multiple image frames based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
[0018] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement some or all of the steps described in any method of the first aspect of the embodiments of this application.
[0019] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.
[0020] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.
[0021] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained from these drawings without creative effort.
[0023] Figure 1 is a flowchart illustrating a loop closure detection method in one embodiment;
[0024] Figure 2 is a structural block diagram of a loop closure detection device in one embodiment;
[0025] Figure 3 is an internal structure diagram of a computer device in one embodiment;
[0026] Figure 4 is an internal structure diagram of a computer device in one embodiment. Detailed Implementation
[0027] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings. Preferred embodiments of this application are shown in the drawings. However, this application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of this application.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. The terminology used herein in the specification of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0029] This application provides a loop closure detection method, illustrated by its application to a computer device, which can be a terminal or a server. It is understood that this method can also be applied to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various robots, autonomous vehicles, automated delivery vehicles, personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. Robots can be cleaning robots, delivery robots, logistics robots, inspection robots, and disinfection robots, etc. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0030] In an exemplary embodiment, as shown in FIG1, a loop closure detection method is provided. Taking the application of this method to a robot as an example, the method includes the following steps 102 to 108. Wherein:
[0031] Step 102: Acquire multiple image frames collected by the robot, and determine the last image frame and a first preset number of first preceding image frames among the multiple image frames.
[0032] The robot is equipped with an image acquisition device, which uses the images acquired to build a map and determine its location. The image acquisition device can capture images from at least one location, including above, to the side, in front, behind, and on the ground.
[0033] An image frame is a single image within a video stream or image sequence acquired by the robot. In VSLAM technology, image frames are key data units that constitute the environment map and the robot's path. Easily understood, each image frame corresponds to an acquisition time point and an initial pose. The initial pose of each image frame can be obtained through a wheel odometry system, other sensors mounted on the robot to acquire its pose, or by fusion of multiple sensors.
[0034] In an exemplary embodiment, the acquisition of multiple image frames collected by the robot includes: acquiring multiple image frames collected by the robot through an image acquisition device; wherein the lens of the image acquisition device is facing upwards, and the image frame is an image above the environment in which the robot is located; when the robot is in an indoor scene, the image frame may be a ceiling image; when the robot is in an outdoor scene, the image frame may be an image of the sky, buildings, or trees above the robot.
[0035] Optionally, the number of multiple image frames can be 5, 10, 20 or other numbers, depending on the actual algorithm calculation requirements, and is not specifically limited here.
[0036] Alternatively, the image acquisition device may be a monocular camera, a binocular camera, or a depth camera, etc.
[0037] Optionally, the robot may be, but is not limited to, various industrial robots (such as handling robots, palletizing robots, painting robots, etc.), service robots (such as cleaning robots, delivery robots, lawnmowing robots, guiding robots, building robots, etc.) or special robots (firefighting robots, underwater robots, security robots, etc.) that require autonomous movement for map building.
[0038] The last image frame refers to the image frame whose acquisition time node is last among multiple image frames; that is, the last image frame is the last image frame acquired by the robot. The first preceding image frames refer to the first preset number of image frames whose acquisition time node precedes the last image frame. In simple terms, the last image frame in the first preceding image frames is adjacent to the acquisition time node of the last image frame. For example, if the robot acquires 100 image frames, the last image frame is the 100th image frame acquired by the robot. Assuming the first preset number is 10, the first preceding image frames are the 90th, 91st, ..., 99th image frames acquired by the robot.
[0039] Optionally, the first preset quantity can be 5, 10, 15 or other quantities.
[0040] Step 104: Determine the feature matching degree between each first preceding image frame and the last image frame, and determine the first target image frame corresponding to the last image frame from the first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree between it and the last image frame from the first preset number of first preceding image frames.
[0041] Among them, the first target image frame corresponding to the last image frame has the highest feature matching degree with the last image frame among the first preset number of first preceding image frames. That is to say, the first target image frame is the image frame with the highest visual similarity to the last image frame among the first preset number of first preceding image frames. This means that the first target image frame corresponding to the last image frame and the last image frame may have been collected by the robot at the same or similar positions.
[0042] Optionally, to ensure the accuracy of the robot's loop closure detection, while ensuring that the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree with the last image frame, it can also be further ensured that the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame meets the first matching degree condition. The first matching degree condition can be a feature matching degree greater than or equal to 50%, or greater than 60%, etc., which can be set according to actual needs.
[0043] In an exemplary embodiment, the determination of the feature matching degree between each first preceding image frame and the last image frame includes: extracting feature points of each image frame, and performing pairwise matching of each image frame based on the feature points of each image frame to determine the feature matching relationship between each image frame; and determining the feature matching degree between each first preceding image frame and the last image frame based on the feature matching relationship between each image frame.
[0044] The feature matching relationship between each image frame includes the feature matching relationship between any two image frames in a plurality of image frames, representing the correspondence of feature points between each image frame. Therefore, based on the feature matching relationship between each image frame, the feature matching relationship between any two image frames in a plurality of image frames can be determined.
[0045] Optionally, the feature points of each image frame can be Oriented Fast and Rotated BRIEF (ORB) feature points, Scale-Invariant Feature Transform (SIFT) feature points, or other deep learning feature points.
[0046] Step 106: Among multiple image frames, determine the second target image frame corresponding to the first target image frame corresponding to the last image frame; the second target image frame corresponding to the last image frame is the image frame among multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies the first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame.
[0047] The first distance condition can be a distance within a radius of 3m, 4m, 5m, 6m or other distances; for example, when the first distance condition is a distance within a radius of 5m, the second target image frame corresponding to the last image frame is a portion of the multiple image frames whose distance is within a radius of 5m of the first target image frame corresponding to the last image frame.
[0048] The second target image frame corresponding to the last image frame satisfies the first distance condition with the first target image frame corresponding to the last image frame, and has the highest feature matching degree with the first target image frame corresponding to the last image frame. In other words, the second target image frame corresponding to the last image frame is the image frame that is closest to the first target image frame corresponding to the last image frame among multiple image frames, and has the highest visual similarity to the first target image frame corresponding to the last image frame. This means that the second target image frame corresponding to the last image frame and the first target image frame corresponding to the last image frame may have been acquired by the robot at the same or similar locations. Furthermore, since the first target image frame corresponding to the last image frame has the highest feature matching degree with the last image frame representing the robot's current location, the second target image frame corresponding to the last image frame is an image frame from a historical location that may have the highest similarity to the robot's current location.
[0049] Specifically, the feature matching degree between the second target image frame corresponding to the last image frame and the first target image frame corresponding to the last image frame can also be determined based on the feature matching relationship between each image frame.
[0050] Step 108: Based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, perform loop closure detection in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
[0051] Loop closure detection, also known as loop shut-off detection, refers to a robot's ability to recognize areas it has previously visited, thus closing the map loop. Simply put, when a robot turns left or right during mapping, it recognizes that it has been to a certain location and matches the newly generated map with the current one. Loop closure detection is challenging because successful detection significantly reduces accumulated errors, helping the robot perform obstacle avoidance and navigation more accurately and quickly. Incorrect detection results can severely degrade the map, impacting subsequent obstacle avoidance and navigation. Therefore, loop closure detection is essential for building large-area, large-scene maps.
[0052] Optionally, at least one loop closure frame of the first target image frame corresponding to the last image frame obtained by loop closure detection can be one image frame that meets the loop closure detection conditions, or multiple image frames that meet the loop closure detection conditions.
[0053] Specifically, when a loop closure frame is detected, it indicates that the robot has returned to a previously visited location. At this point, pose optimization can be performed on multiple image frames to enable the robot to construct a map with higher accuracy. It's easy to understand that if loop closure frames are mistakenly identified as present when they don't actually exist, performing pose optimization on these frames under such erroneous circumstances will lead to significant errors in the robot's map construction. Therefore, the robot needs to ensure accurate loop closure detection to avoid this undesirable situation that results in inaccurate map construction.
[0054] It is easy to understand that when at least one loopback frame is found among multiple image frames, the pose of the first target image frame corresponding to the last image frame is the loopback pose.
[0055] In the above loop closure detection method, multiple image frames collected by the robot are acquired, and a last image frame and a first preset number of first preceding image frames are determined among the multiple image frames; the feature matching degree between each first preceding image frame and the last image frame is determined, and a first target image frame corresponding to the last image frame is determined among the first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree among the first preset number of first preceding image frames; among the multiple image frames, a second target image frame corresponding to the first target image frame corresponding to the last image frame is determined; the second target image frame corresponding to the last image frame is the image frame among the multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies a first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame; based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, loop closure detection is performed in the multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame. The loop closure detection method provided in this application determines a first target image frame in the first preceding image frame that has the highest feature matching degree with the last image frame. Then, it determines a second target image frame in multiple image frames that has the highest feature matching degree with the last image frame and whose distance from the first target image frame satisfies a first distance condition. Based on the feature matching degree between the first target image frame and the second target image frame, loop closure detection is performed in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame. Obviously, the loop closure detection method provided in this embodiment can accurately detect loop closures in multiple image frames using the first target image frame corresponding to the last image frame to obtain loop closure frames with high accuracy, which can significantly improve the accuracy of loop closure detection. Furthermore, with high accuracy in loop closure detection, the accuracy of map construction can be further ensured.
[0056] In an exemplary embodiment, when at least one loopback frame to the first target image frame corresponding to the last image frame is determined, the above method further includes:
[0057] Based on the initial pose of each loopback frame, the calculated pose of each loopback frame, and the first objective function, the initial pose of each image frame is optimized to determine the first optimized pose of each image frame.
[0058] In an exemplary embodiment, the method further includes: determining the calculated pose of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame; after determining at least one loopback frame of the first target image frame corresponding to the last image frame, the method further includes: determining the calculated pose of each loopback frame in the calculated pose of each image frame.
[0059] The calculated pose of each image frame can be used to optimize the pose of each image frame, enabling the robot to build a more accurate map. By optimizing the pose of each image frame, the cumulative error of the robot in the map building process can be reduced, thereby improving the consistency between the robot's overall motion trajectory and the actual map.
[0060] The first optimized pose of each image frame is used to enable the robot to optimize the constructed map, reduce the accumulated error in the mapping process, and build a map with higher accuracy and higher consistency with the actual map.
[0061] In an exemplary embodiment, the above-described determination of the calculated pose of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame includes: determining the three-dimensional coordinates of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame; and determining the calculated pose of each image frame based on the feature matching relationship between each image frame and the three-dimensional coordinates of each image frame.
[0062] For example, taking the determination of the calculated pose of the last image frame as an example, firstly, based on the feature matching relationship between the first target image frame of the last image frame and other image frames, and the initial pose of each image frame, the three-dimensional coordinates of the first target image frame of the last image frame can be calculated. Then, based on the feature matching relationship between the last image frame and the first target image frame of the last image frame, and the three-dimensional coordinates of the first target image frame of the last image frame, the calculated pose of the last image frame can be calculated. The calculation principle for the calculated poses of other image frames is similar, so it will not be repeated here.
[0063] Optionally, based on the feature matching relationships between image frames and the 3D coordinates of each image frame, the calculated pose of each image frame can be determined using the Perspective-n-Point (PnP) algorithm. Specifically, in VSLAM technology, the PnP algorithm is typically used to recover the robot's pose in a 2D image frame from the correspondence between feature points of a set of matched 2D image frames and their corresponding 3D feature points.
[0064] Specifically, the first optimized pose of each image frame is the initial pose of each image frame based on each loop-loop frame, the calculated pose of each loop-loop frame, and the pose that the first objective function can solve to minimize the first objective function.
[0065] For example, the number of multiple image frames is represented as n, and the first optimized pose of each image frame is represented as T. i The initial pose of each loop frame is represented as T' k The calculated pose obtained from each loopback frame based on the initial pose is represented as T'. cal_k Then the first objective function can be in,
[0066] In one embodiment, the initial pose of each loopback frame is input into a first objective function, and the initial pose that minimizes the first objective function is the first optimized pose of each image frame.
[0067] In this embodiment, given at least one loopback frame of the first target image frame corresponding to the last image frame, a first optimized pose for each image frame is determined based on the initial pose of each loopback frame, the calculated pose of each loopback frame, and the first objective function. Therefore, based on accurate loopback frames, the accuracy of the determined first optimized pose for each image frame can be improved, thereby further ensuring the accuracy of map construction.
[0068] In an exemplary embodiment, after determining the first target image frame corresponding to the last image frame among a first preset number of first preceding image frames, the above method further includes:
[0069] Based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, the initial pose of each image frame is optimized to determine the second optimized pose of each image frame.
[0070] In an exemplary embodiment, the above-mentioned determination of the calculated pose of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame includes: determining the calculated pose of each image frame based on the feature matching relationship between each image frame and the second optimized pose of each image frame.
[0071] The second optimized pose of each image frame is obtained by correcting the trajectory drift of each image frame using the last image frame, so that the loop closure pose can be accurately determined in the subsequent loop closure detection process.
[0072] Specifically, the second optimized pose of each image frame is the initial pose of each image frame based on the last image frame, the calculated pose of the last image frame, and the pose that the second objective function can solve to minimize the second objective function.
[0073] For example, the second optimized pose of each image frame is represented as T k The initial pose of the last image frame is represented as T. end The calculated pose of the last image frame is represented as T. cal_end Then the second objective function can be in,
[0074] In an exemplary embodiment, the initial pose of each image frame is input into the second objective function, and the initial pose that minimizes the second objective function is the second optimized pose of each image frame.
[0075] In an exemplary embodiment, the second optimized pose of each image frame is optimized based on the second optimized pose of each loopback frame, the calculated pose of each loopback frame, and the first objective function to determine the first optimized pose of each image frame. Furthermore, the calculated pose of each loopback frame is calculated based on the second optimized pose of each image frame. In this case, the calculated pose of each image frame based on the second optimized pose is represented as T. cal_k Then the first objective function is in,
[0076] In this embodiment, based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, the second optimized pose of each image frame is determined. The trajectory drift of each image frame is initially corrected by the last image frame. Thus, the calculated pose of each image frame and the first optimized pose of each image frame are obtained based on the second optimized pose of each image frame with higher accuracy after the initial correction of trajectory drift. In this way, the accuracy of loop closure detection can be further improved.
[0077] In one exemplary embodiment, the method further includes:
[0078] If the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame satisfies the first matching degree condition, the three-dimensional coordinates of the first target image frame corresponding to the last image frame are determined based on the feature matching relationship between the first target image frame corresponding to the last image frame and each image frame and the initial pose of each image frame.
[0079] The calculated pose of the last image frame is determined based on the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame, as well as the three-dimensional coordinates of the first target image frame corresponding to the last image frame.
[0080] The three-dimensional coordinates of the first target image frame corresponding to the last image frame can be determined by triangulation or other three-dimensional reconstruction techniques, based on the feature matching relationship between the first target image frame corresponding to the last image frame and each image frame, as well as the initial pose of each image frame.
[0081] Optionally, the calculated pose of the last image frame can be determined by using the PnP algorithm, based on the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame, as well as the three-dimensional coordinates of the first target image frame corresponding to the last image frame.
[0082] In this embodiment, by using the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame, as well as the three-dimensional coordinates of the first target image frame corresponding to the last image frame, the calculated pose of the last image frame can be determined based on the correspondence between the two-dimensional feature matching relationship and the three-dimensional coordinates. Thus, by ensuring the accuracy of the calculated pose of the last image frame, it is possible to further ensure that the second optimized pose of each image frame also has a high accuracy, thereby further improving the accuracy of loop closure detection.
[0083] In an exemplary embodiment, if the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame does not meet the first matching degree condition, the initial pose of the first target image frame corresponding to the last image frame is determined as the calculated pose of the last image frame.
[0084] In this embodiment, if the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame does not meet the first matching degree condition, it means that it is difficult to accurately determine the calculated pose of the last image frame by calculation. Since the robot's initial pose and last pose are usually ensured to be the same during manual intervention or operation, the initial pose of the first target image frame corresponding to the last image frame is directly determined as the calculated pose of the last image frame to avoid calculating an incorrect pose, which would lead to a large error in the map constructed by the robot.
[0085] In an exemplary embodiment, determining the second target image frame corresponding to the first target image frame corresponding to the last image frame among multiple image frames includes:
[0086] Among multiple image frames, at least one surrounding image frame whose distance to the first target image frame corresponding to the last image frame satisfies the first distance condition is determined.
[0087] In at least one surrounding image frame, the second target image frame corresponding to the last target image frame that has the highest feature matching degree with the first target image frame corresponding to the last image frame is determined.
[0088] In this process, the acquisition time of each peripheral image frame is earlier than the acquisition time of the first target image frame corresponding to the last image frame. Furthermore, when the direction of the peripheral image frame is the same as or opposite to the direction of the first target image frame corresponding to the last image frame, the time difference between the acquisition time of the peripheral image frame and the acquisition time of the first target image frame corresponding to the last image frame satisfies the preset time condition.
[0089] Specifically, the acquisition time of each peripheral image frame is earlier than the acquisition time of the first target image frame corresponding to the last image frame. That is to say, each peripheral image frame was acquired before the first target image frame corresponding to the last image frame was acquired.
[0090] Optionally, the preset time condition can be a time difference greater than or equal to 5 seconds, 10 seconds, 15 seconds, or other time differences.
[0091] Optionally, the orientation of each peripheral image frame can be determined by performing pose matching between the calculated pose of each peripheral image frame and the calculated pose of the first target image frame corresponding to the last image frame. This can be done to determine whether the orientation of each peripheral image frame is the same as or opposite to the orientation of the first target image frame corresponding to the last image frame.
[0092] Specifically, in order to perform loop closure detection across multiple image frames, the acquisition time of at least one surrounding image frame used to determine the second target image frame corresponding to the last image frame should be earlier than the acquisition time of the first target image frame corresponding to the last image frame. Meanwhile, to avoid excessive computational power consumption during loop closure detection, a certain acquisition time difference exists between the surrounding image frames and the first target image frame corresponding to the last image frame, provided their directions are the same or opposite. Clearly, by limiting the acquisition time of the image frames, the accuracy of the second target image frame corresponding to the last image frame determined by at least one surrounding image frame is ensured, while also ensuring sufficient efficiency for subsequent loop closure detection.
[0093] In an exemplary embodiment, the method of determining the second target image frame corresponding to the last target image frame with the highest feature matching degree among at least one surrounding image frame includes: rotating the first target image frame corresponding to the last image frame to the same direction as each surrounding image frame to obtain a plurality of rotated first target image frames, wherein the rotated first target image frames correspond one-to-one with the surrounding image frames; determining the feature matching relationship between each rotated first target image frame and each surrounding image frame; and determining the second target image frame corresponding to the last target image frame with the highest feature matching degree among at least one surrounding image frame based on the feature matching relationship between each rotated first target image frame and each surrounding image frame.
[0094] In this embodiment, at least one surrounding image frame with a small distance to the first target image frame corresponding to the last image frame is identified among multiple image frames. This at least one surrounding image frame is a potential candidate for loop closure poses. Then, among the at least one surrounding image frame, a second target image frame corresponding to the last image frame with the highest feature matching degree is identified. This second target image frame corresponding to the last image frame is a historical position image frame with the highest similarity to the robot's current position. Furthermore, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, loop closure poses are obtained by performing loop closure detection on the second target image frame across multiple image frames. Clearly, the loop closure detection method provided in this embodiment can accurately detect loop closures on the first target image frame across multiple image frames to obtain loop closure poses with high accuracy, thereby significantly improving the accuracy of loop closure detection. Simultaneously, due to the constraints on the acquisition time of the surrounding image frames, the efficiency of loop closure detection is also improved while increasing the accuracy.
[0095] In an exemplary embodiment, the above-mentioned method of performing loop closure detection in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, includes:
[0096] If the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the second matching degree condition.
[0097] Determine the first feature matching degree between the image frame preceding the first target image frame corresponding to the last image frame and the image frame preceding the second target image frame corresponding to the last image frame.
[0098] Determine the first relative position difference between the first target image frame corresponding to the last image frame and the image frame preceding the first target image frame corresponding to the last image frame.
[0099] Determine the first pose difference between the second target image frame corresponding to the last image frame and the image frame preceding the second target image frame corresponding to the last image frame.
[0100] And, determine the second pose difference between the previous image frame of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame.
[0101] If at least one of the following conditions is met: the first feature matching degree satisfies the third matching degree condition; the distance difference between the first relative position difference and the first position pose difference satisfies the second distance condition; and the distance difference between the first relative position difference and the second position pose difference satisfies the third distance condition, then the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second target image frame corresponding to the last image frame.
[0102] The second matching degree condition can be a feature matching degree greater than or equal to 50%, 60%, 80%, or other feature matching degrees.
[0103] Optionally, the third matching degree condition can be a feature matching degree greater than or equal to 50%, 60%, 80%, or other feature matching degrees. Optionally, the second matching degree condition can be the same as the third matching degree condition.
[0104] Optionally, the second distance condition can be a distance difference of less than 15cm, 20cm, 25cm, or other distance differences.
[0105] Optionally, the third distance condition can be a distance difference of less than 15cm, 20cm, 25cm, or other distance differences. Optionally, the second distance condition can be the same as the third distance condition.
[0106] The preceding image frame of the first target image frame corresponding to the last image frame is the image frame whose acquisition time node is the previous acquisition time node of the first target image frame corresponding to the last image frame. In other words, among multiple consecutive image frames with consecutive acquisition time nodes, the preceding image frame of the first target image frame corresponding to the last image frame is an image frame that is adjacent to the first target image frame corresponding to the last image frame and whose acquisition time node is earlier than that of the first target image frame corresponding to the last image frame.
[0107] The preceding image frame of the second target image frame corresponding to the last image frame is the image frame whose acquisition time node is the previous acquisition time node of the second target image frame corresponding to the last image frame. In other words, among multiple consecutive image frames with consecutive acquisition time nodes, the preceding image frame of the second target image frame corresponding to the last image frame is an image frame that is adjacent to the second target image frame corresponding to the last image frame and whose acquisition time node is earlier than that of the second target image frame corresponding to the last image frame.
[0108] Specifically, the first relative position difference between the first target image frame corresponding to the last image frame and the previous image frame corresponding to the last image frame is determined by the straight-line distance between the position of the first target image frame corresponding to the last image frame and the position of the previous image frame corresponding to the last image frame.
[0109] Specifically, the first pose difference between the second target image frame corresponding to the last image frame and the image frame preceding the second target image frame corresponding to the last image frame is determined by the relative change between the pose of the second target image frame corresponding to the last image frame and the pose of the image frame preceding the second target image frame corresponding to the last image frame. The second pose difference between the image frames preceding the first target image frame corresponding to the last image frame and the image frames preceding the second target image frame corresponding to the last image frame is similar to the first pose difference, and therefore will not be described again here.
[0110] Specifically, in VSLAM technology, position refers to a specific point on the robot in three-dimensional space. Position does not involve the robot's orientation information, but only its translation information. That is, position describes the robot's linear movement in three-dimensional space, without any rotation or change of orientation. In VSLAM technology, pose refers to the combination of the robot's position and its orientation. Orientation refers to the robot's direction or orientation. Therefore, pose describes not only the robot's specific point in three-dimensional space, but also its direction or orientation. In other words, pose involves the robot's orientation and translation information.
[0111] Specifically, satisfying at least one of the following conditions—that the first feature matching degree satisfies the third matching degree condition, the distance difference between the first relative position difference and the first position pose difference satisfies the second distance condition, and the distance difference between the first relative position difference and the second position pose difference satisfies the third distance condition—can mean satisfying any one of the three conditions, any two of the three conditions, or all of the three conditions.
[0112] It is understandable that if the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the second matching degree condition, it indicates that the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame may have been acquired at similar or the same location. In order to more accurately determine the loop closure, it is possible to further determine whether the first feature matching degree satisfies the third matching degree condition, and / or whether the distance difference between the first relative position difference and the first pose difference satisfies the second distance condition, and / or whether the distance difference between the first relative position difference and the second pose difference satisfies the third distance condition.
[0113] In one specific embodiment, if the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the second matching degree condition, it indicates that the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame may have been acquired at similar or identical locations. To further confirm whether the second target image frame corresponding to the last image frame forms a loop with the first target image frame corresponding to the last image frame—that is, to further confirm whether the looped frame of the first target image frame corresponding to the last image frame is the second target image frame corresponding to the last image frame—it is necessary to confirm the last image frame... Whether the previous image frame of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame were also acquired at similar or the same positions, and whether there is a certain degree of motion continuity and motion similarity between the first target image frame corresponding to the last image frame and the previous image frame of the first target image frame corresponding to the last image frame, between the second target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame, and between the previous image frame of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame. Therefore, when all three conditions are met simultaneously—the first feature matching degree satisfies the third matching degree condition, the distance difference between the first relative position difference and the first pose difference satisfies the second distance condition, and the distance difference between the first relative position difference and the second pose difference satisfies the third distance condition—it can be considered that the robot has returned to the previously visited position. At this time, the pose of the second target image frame corresponding to the last image frame is determined to form a loop with the first target image frame corresponding to the last image frame. That is to say, at this time, the loop frame of the first target image frame corresponding to the last image frame is determined to be the second target image frame corresponding to the last image frame.
[0114] It can be seen that when there is a high similarity between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, by setting specific loop closure conditions to find the loop closure pose, it can be ensured that the second target image frame corresponding to the last image frame is the loop closure frame of the first target image frame corresponding to the last image frame, thereby improving the accuracy of loop closure detection.
[0115] In an exemplary embodiment, the above-mentioned method of performing loop closure detection in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, includes:
[0116] When the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition.
[0117] In a plurality of image frames, a second preset number of consecutive candidate image frames and a third preset number of second preceding image frames for each candidate image frame are determined;
[0118] The feature matching degree between each second preceding image frame and each candidate image frame is determined respectively, and the first target image frame corresponding to each candidate image frame is determined in the third preset number of second preceding image frames of each candidate image frame, and the second target image frame corresponding to the first target image frame corresponding to each candidate image frame is determined in the multiple image frames respectively.
[0119] Loop closure detection is performed by considering the relationship between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame, so as to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
[0120] The first target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among the third preset number of second preceding image frames; the second target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among the multiple image frames whose distance to the first target image frame corresponding to the candidate image frame satisfies the first distance condition and whose distance to the first target image frame corresponding to the candidate image frame satisfies the first distance condition.
[0121] The second preset number of consecutive candidate image frames refers to a second preset number of candidate image frames that are continuous and uninterrupted at the acquisition time node. The second preceding image frame refers to a third preset number of image frames whose acquisition time node precedes each candidate image frame among multiple image frames. It is easy to understand that the image frame whose acquisition time node is the last in the second preceding image frame is adjacent to the acquisition time node of the candidate image frame corresponding to the second preceding image frame.
[0122] In an exemplary embodiment, the above-described loop closure detection based on the relationship between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame, to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, includes:
[0123] Determine the second feature matching degree between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame.
[0124] Determine the second relative position difference between the first target image frame corresponding to each candidate image frame and the previous image frame corresponding to the first target image frame.
[0125] In addition, the second pose difference between the second target image frame corresponding to each candidate image frame and the previous image frame of the second target image frame corresponding to each candidate image frame is determined.
[0126] When the second feature matching degree satisfies the fifth matching degree condition, and the distance difference between the second relative position difference and the second pose difference satisfies at least one of the fourth distance conditions, the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second preset number of candidate image frames.
[0127] The fourth matching degree condition can be a feature matching degree greater than 20% and less than 50%, greater than 10% and less than 50%, greater than 10% and less than 60%, or other feature matching degrees.
[0128] Optionally, the fifth matching degree condition can be a feature matching degree greater than 20% and less than 50%, greater than 10% and less than 50%, greater than 10% and less than 60%, or other feature matching degrees. Optionally, the fourth matching degree condition can be the same as the fifth matching degree condition.
[0129] Optionally, the fourth distance condition can be a distance difference of less than 15cm, 20cm, 25cm, or other distance differences. Optionally, the second, third, and fourth distance conditions can be the same.
[0130] Optionally, the second preset quantity can be 5, 10, 15, or other quantities. Optionally, the first preset quantity can be the same as the second preset quantity.
[0131] Optionally, the third preset quantity can be 5, 10, 15, or other quantities. Optionally, the second preset quantity can be the same as the third preset quantity.
[0132] Specifically, satisfying multiple second feature matching degrees satisfies the fifth matching degree condition, and satisfying at least one of the fourth distance conditions in the distance difference between the first relative position difference and the first pose difference can mean satisfying either of the two conditions or satisfying all of the two conditions.
[0133] It is understandable that the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition, indicating that the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame have a certain similarity, but it cannot be completely determined whether they were collected at similar or the same position. In order to confirm whether the similarity between the two is sufficient to form a loop closure for more accurate loop closure judgment, more image frames are needed to assist in loop closure detection. Therefore, it is possible to further determine whether the second feature matching degree satisfies the fifth matching degree condition, and / or whether the distance difference between the second relative position difference and the second pose difference satisfies the fourth distance condition.
[0134] In one specific embodiment, if the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition, it indicates that the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame may have been acquired at similar or identical locations. However, since the feature matching degree is insufficient to completely determine this, in order to further confirm whether the robot has returned to a previously visited location, it is necessary to confirm whether there are motion continuity and motion similarity among the second preset number of consecutive candidate image frames acquired by the robot within the historical acquisition time nodes that satisfy the loopback frame judgment. Therefore, when both the second feature matching degree satisfies the fifth matching degree condition and the distance difference between the second relative position difference and the second pose difference satisfies the fourth distance condition, it can be considered that the robot has returned to a previously visited location. In this case, the loopback frame of the first target image frame corresponding to the last image frame is determined to be one of the second preset number of candidate image frames.
[0135] It can be seen that when the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame have a certain degree of similarity but the similarity is not high enough, by setting specific loop closure conditions to find loop closure frames, it is possible to ensure that the second preset number of candidate image frames are loop closure frames of the first target image frame corresponding to the last image frame, thus avoiding missed detection of loop closure frames and improving the accuracy of loop closure detection.
[0136] The application process of the above loop closure detection method is illustrated below with a detailed embodiment:
[0137] (1) Acquisition process of multiple image frames
[0138] The robot acquires multiple image frames and determines the last image frame and a first preset number of first preceding image frames among the multiple image frames.
[0139] (2) Determination process of the first target image frame
[0140] Determine the feature matching degree between each first preceding image frame and the last image frame, and determine the first target image frame corresponding to the last image frame among a first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree with the last image frame among the first preset number of first preceding image frames.
[0141] (3) The process of determining the second optimized pose for each image frame
[0142] If the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame satisfies the first matching degree condition, the three-dimensional coordinates of the first target image frame corresponding to the last image frame are determined based on the feature matching relationship between the first target image frame corresponding to the last image frame and each image frame and the initial pose of each image frame; the calculated pose of the last image frame is determined based on the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame and the three-dimensional coordinates of the first target image frame corresponding to the last image frame.
[0143] If the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame does not meet the first matching degree condition, the initial pose of the first target image frame corresponding to the last image frame is determined as the calculated pose of the last image frame.
[0144] Based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, the second optimized pose of each image frame is determined.
[0145] (4) The process of determining the loop pose
[0146] Among multiple image frames, at least one surrounding image frame is identified whose distance to the first target image frame corresponding to the last image frame satisfies a first distance condition; among the at least one surrounding image frame, a second target image frame corresponding to the last image frame with the highest feature matching degree is identified; if the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies a second matching degree condition, a first feature matching degree is determined between the previous image frame of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame, and the first target image frame corresponding to the last image frame is determined to be the first target image frame corresponding to the last image frame. The first relative position difference between the previous image frames of the last image frame is used to determine the first pose difference between the second target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame, and the second pose difference between the previous image frames of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame; when at least one of the following conditions is met: the first feature matching degree satisfies the third matching degree condition, the distance difference between the first relative position difference and the first pose difference satisfies the second distance condition, and the distance difference between the first relative position difference and the second pose difference satisfies the third distance condition, the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second target image frame corresponding to the last image frame;
[0147] When the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition, a second preset number of consecutive candidate image frames and a third preset number of second preceding image frames are determined from multiple image frames. The feature matching degree between each second preceding image frame and each candidate image frame is determined respectively. The first target image frame corresponding to each candidate image frame is determined from the third preset number of second preceding image frames of each candidate image frame. Furthermore, the second target image frame corresponding to the first target image frame corresponding to each candidate image frame is determined from multiple image frames. Thus, the first target image frame corresponding to each candidate image frame is determined. The second feature matching degree between the target image frame and the second target image frame corresponding to each candidate image frame is determined, the second relative position difference between the first target image frame corresponding to each candidate image frame and the previous image frame of the first target image frame corresponding to each candidate image frame is determined, and the second pose difference between the second target image frame corresponding to each candidate image frame and the previous image frame of the second target image frame corresponding to each candidate image frame is determined. When at least one of the following conditions is met, the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second preset number of candidate image frames.
[0148] In this process, the acquisition time of each peripheral image frame is earlier than the acquisition time of the first target image frame corresponding to the last image frame. Furthermore, when the direction of the peripheral image frame is the same as or opposite to the direction of the first target image frame corresponding to the last image frame, the time difference between the acquisition time of the peripheral image frame and the acquisition time of the first target image frame corresponding to the last image frame satisfies the preset time condition.
[0149] (5) The process of determining the pose of each image frame
[0150] When at least one loopback frame of the first target image frame corresponding to the last image frame is determined, the calculated pose of each image frame is determined based on the feature matching relationship between each image frame and the second optimized pose of each image frame.
[0151] (6) The process of determining the first optimized pose for each image frame
[0152] Based on the second optimized pose of each image frame, the calculated pose of each image frame, and the first objective function, the first optimized pose of each image frame is determined.
[0153] In this embodiment, by determining the first target image frame corresponding to the last image frame with the highest feature matching degree with the last image frame, the calculated pose of the last image frame can be determined through the first target image frame corresponding to the last image frame. Based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, the second optimized pose of each image frame is determined. The trajectory drift of each image frame is initially corrected through the last image frame. Furthermore, the second target image frame corresponding to the last image frame is determined through the first target image frame corresponding to the last image frame. The second target image frame corresponding to the last image frame is close in distance and has a high similarity to the last image frame. Then, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, loop closure detection is performed in multiple image frames to obtain at least one loop closure frame of the first target image frame corresponding to the last image frame. The accuracy of loop closure detection can be improved by ensuring the accuracy of the loop closure frame. As can be seen, the loop closure detection method provided in this embodiment enables the robot to perform accurate loop closure detection in complex environments. Thus, the robot can distinguish between correct loop closures and false loop closures in complex environments. Furthermore, in complex environments with similar ceiling features, it can avoid the undesirable situation where the robot detects incorrect loop closures, leading to large map construction errors. At the same time, it can also improve the accuracy of loop closure detection in environments with dissimilar features, thereby improving the accuracy of map construction.
[0154] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0155] Based on the same inventive concept, this application also provides a loop closure detection device for implementing the loop closure detection method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more loop closure detection device embodiments provided below can be found in the limitations of the loop closure detection method described above, and will not be repeated here.
[0156] In an exemplary embodiment, as shown in FIG2, a loop closure detection device is provided, comprising: an acquisition module 202, a first determination module 204, a second determination module 206, and a detection module 208, wherein:
[0157] The acquisition module 202 is used to acquire multiple image frames collected by the robot, and determine the last image frame and a first preset number of first preceding image frames among the multiple image frames.
[0158] The first determining module 204 is used to determine the feature matching degree between each first preceding image frame and the last image frame, and to determine the first target image frame corresponding to the last image frame among a first preset number of first preceding image frames; the first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree with the last image frame among the first preset number of first preceding image frames.
[0159] The second determining module 206 is used to determine, among multiple image frames, the second target image frame corresponding to the first target image frame corresponding to the last image frame; the second target image frame corresponding to the last image frame is the image frame among the multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies the first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame.
[0160] The detection module 208 is used to perform loop closure detection in multiple image frames based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
[0161] In an exemplary embodiment, the detection module 208 is further configured to optimize the initial pose of each image frame based on the initial pose of each loopback frame, the calculated pose of each loopback frame, and the first objective function, and determine the first optimized pose of each image frame.
[0162] In an exemplary embodiment, the first determining module 204 is further configured to optimize the initial pose of each image frame based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, and determine the second optimized pose of each image frame.
[0163] In an exemplary embodiment, the first determining module 204 is further configured to, when the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame satisfies the first matching degree condition, determine the three-dimensional coordinates of the first target image frame corresponding to the last image frame based on the feature matching relationship between the first target image frame corresponding to the last image frame and each image frame and the initial pose of each image frame; and determine the calculated pose of the last image frame based on the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame and the three-dimensional coordinates of the first target image frame corresponding to the last image frame.
[0164] In an exemplary embodiment, the first determining module 204 is further configured to determine the initial pose of the first target image frame corresponding to the last image frame as the calculated pose of the last image frame when the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame does not meet the first matching degree condition.
[0165] In an exemplary embodiment, the second determining module 206 is further configured to: determine at least one surrounding image frame whose distance to the first target image frame corresponding to the last image frame satisfies a first distance condition among multiple image frames; and determine, among the at least one surrounding image frame, the second target image frame corresponding to the last image frame that has the highest feature matching degree with the first target image frame corresponding to the last image frame; wherein, the acquisition time node of each surrounding image frame is earlier than the acquisition time node of the first target image frame corresponding to the last image frame, and, when the direction of the surrounding image frame is the same as or opposite to the direction of the first target image frame corresponding to the last image frame, the time difference between the acquisition time node of the surrounding image frame and the acquisition time node of the first target image frame corresponding to the last image frame satisfies a preset time condition.
[0166] In an exemplary embodiment, the detection module 208 is further configured to, when the feature matching degree between the first target image frame corresponding to the last image frame and the previous image frame corresponding to the second target image frame satisfies the second matching degree condition, determine the first feature matching degree between the first target image frame corresponding to the last image frame and the previous image frame corresponding to the second target image frame, determine the first relative position difference between the first target image frame corresponding to the last image frame and the previous image frame corresponding to the first target image frame, and determine the relationship between the second target image frame corresponding to the last image frame and the last image frame. The first pose difference between the previous image frames of the corresponding second target image frame, and the second pose difference between the previous image frames of the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame; when at least one of the following conditions is met: the first feature matching degree satisfies the third matching degree condition, the distance difference between the first relative position difference and the first pose difference satisfies the second distance condition, and the distance difference between the first relative position difference and the second pose difference satisfies the third distance condition, the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second target image frame corresponding to the last image frame.
[0167] In an exemplary embodiment, the detection module 208 is further configured to, when the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition, determine a second preset number of consecutive candidate image frames and a third preset number of second preceding image frames for each candidate image frame in a plurality of image frames; determine the feature matching degree between each second preceding image frame and each candidate image frame respectively, and determine the first target image frame corresponding to each candidate image frame in the third preset number of second preceding image frames of each candidate image frame, and determine the first target image frame corresponding to each candidate image frame in a plurality of image frames respectively. The corresponding second target image frame; loop closure detection is performed based on the relationship between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame to determine at least one loop closure frame of the first target image frame corresponding to the last image frame; wherein, the first target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among the second preset number of second preceding image frames; the second target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among the multiple image frames whose distance to the first target image frame corresponding to the candidate image frame satisfies the first distance condition.
[0168] In an exemplary embodiment, the detection module 208 is further configured to determine a second feature matching degree between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame, determine a second relative position difference between the first target image frame corresponding to each candidate image frame and the previous image frame of the first target image frame corresponding to each candidate image frame, and determine a second pose difference between the second target image frame corresponding to each candidate image frame and the previous image frame of the second target image frame corresponding to each candidate image frame; when at least one of the following conditions is met: the second feature matching degree satisfies the fifth matching degree condition, and the distance difference between the second relative position difference and the second pose difference satisfies the fourth distance condition, the loopback frame of the first target image frame corresponding to the last image frame is determined to be a second preset number of candidate image frames.
[0169] Each module in the aforementioned loop closure detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0170] In an exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 3. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores image frame data. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a loop closure detection method.
[0171] In an exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram is shown in Figure 4. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a loop closure detection method. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0172] Those skilled in the art will understand that the structure shown in Figure 4 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.
[0173] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0174] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.
[0175] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0177] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0178] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0179] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A loop closure detection method, the method comprising: Acquire multiple image frames collected by the robot, and determine the last image frame and a first preset number of first preceding image frames among the multiple image frames; Determine the feature matching degree between each of the first preceding image frames and the last image frame, and determine the first target image frame corresponding to the last image frame from a first preset number of first preceding image frames; The first target image frame corresponding to the last image frame is the image frame with the highest feature matching degree among the first preset number of first preceding image frames; Among multiple image frames, a second target image frame corresponding to the first target image frame corresponding to the last image frame is determined; the second target image frame corresponding to the last image frame is an image frame among multiple image frames whose distance to the first target image frame corresponding to the last image frame satisfies a first distance condition and has the highest feature matching degree with the first target image frame corresponding to the last image frame; Based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, loop closure detection is performed in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame.
2. The method according to claim 1, characterized in that, If at least one loopback frame is determined to be connected to the first target image frame corresponding to the last image frame, the method further includes: Based on the initial pose of each loopback frame, the calculated pose of each loopback frame, and the first objective function, the initial pose of each image frame is optimized to determine the first optimized pose of each image frame.
3. The method according to claim 1, characterized in that, After determining the first target image frame corresponding to the last image frame in a first preset number of first preceding image frames, the above method further includes: Based on the initial pose of the last image frame, the calculated pose of the last image frame, and the second objective function, the initial pose of each image frame is optimized to determine the second optimized pose of each image frame.
4. The method according to claim 3, characterized in that, The method further includes: When the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame satisfies the first matching degree condition, the three-dimensional coordinates of the first target image frame corresponding to the last image frame are determined based on the feature matching relationship between the first target image frame corresponding to the last image frame and each of the image frames and the initial pose of each of the image frames. Based on the feature matching relationship between the first target image frame corresponding to the last image frame and the last image frame, as well as the three-dimensional coordinates of the first target image frame corresponding to the last image frame, the calculated pose of the last image frame is determined.
5. The method according to claim 4, characterized in that, If the feature matching degree between the first target image frame corresponding to the last image frame and the last image frame does not meet the first matching degree condition, the initial pose of the first target image frame corresponding to the last image frame is determined as the calculated pose of the last image frame.
6. The method according to claim 1, characterized in that, Determining the second target image frame corresponding to the first target image frame corresponding to the last image frame among multiple image frames includes: Among multiple image frames, at least one surrounding image frame whose distance to the first target image frame corresponding to the last image frame satisfies a first distance condition is determined. In at least one surrounding image frame, a second target image frame is identified that has the highest feature matching degree with the first target image frame corresponding to the last image frame. Wherein, the acquisition time node of each of the surrounding image frames is earlier than the acquisition time node of the first target image frame corresponding to the last image frame, and when the direction of the surrounding image frame is the same as or opposite to the direction of the first target image frame corresponding to the last image frame, the time difference between the acquisition time node of the surrounding image frame and the acquisition time node of the first target image frame corresponding to the last image frame satisfies a preset time condition.
7. The method according to claim 1, characterized in that, The step of performing loop closure detection in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, includes: When the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the second matching degree condition. Determine the first feature matching degree between the image frame preceding the first target image frame corresponding to the last image frame and the image frame preceding the second target image frame corresponding to the last image frame. Determine the first relative position difference between the first target image frame corresponding to the last image frame and the image frame preceding the first target image frame corresponding to the last image frame. Determine the first pose difference between the second target image frame corresponding to the last image frame and the image frame preceding the second target image frame corresponding to the last image frame. And, determine the second pose difference between the previous image frame of the first target image frame corresponding to the last image frame and the previous image frame of the second target image frame corresponding to the last image frame; If at least one of the following conditions is met: the first feature matching degree satisfies the third matching degree condition; the distance difference between the first relative position difference and the first pose difference satisfies the second distance condition; and the distance difference between the first relative position difference and the second pose difference satisfies the third distance condition, then the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second target image frame corresponding to the last image frame.
8. The method according to claim 1, characterized in that, The step of performing loop closure detection in multiple image frames to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, based on the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame, includes: When the feature matching degree between the first target image frame corresponding to the last image frame and the second target image frame corresponding to the last image frame satisfies the fourth matching degree condition. In a plurality of image frames, a second preset number of consecutive candidate image frames and a third preset number of second preceding image frames for each of the candidate image frames are determined; The feature matching degree between each second preceding image frame and each candidate image frame is determined respectively, and the first target image frame corresponding to each candidate image frame is determined in the third preset number of second preceding image frames of each candidate image frame, and the second target image frame corresponding to the first target image frame corresponding to each candidate image frame is determined in the multiple image frames respectively. Loop closure detection is performed by considering the relationship between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame, so as to determine at least one loop closure frame of the first target image frame corresponding to the last image frame. Wherein, the first target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among the third preset number of second preceding image frames; the second target image frame corresponding to the candidate image frame is the image frame with the highest feature matching degree among multiple image frames whose distance to the first target image frame corresponding to the candidate image frame satisfies the first distance condition and whose distance to the first target image frame corresponding to the candidate image frame satisfies the first distance condition.
9. The method according to claim 8, characterized in that, The step of performing loop closure detection based on the relationship between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame, to determine at least one loop closure frame of the first target image frame corresponding to the last image frame, includes: Determine the second feature matching degree between the first target image frame corresponding to each candidate image frame and the second target image frame corresponding to each candidate image frame. Determine the second relative position difference between the first target image frame corresponding to each candidate image frame and the previous image frame corresponding to the first target image frame. In addition, the second pose difference between the second target image frame corresponding to each candidate image frame and the previous image frame of the second target image frame corresponding to each candidate image frame is determined. When the second feature matching degree satisfies the fifth matching degree condition, and the distance difference between the second relative position difference and the second pose difference satisfies the fourth distance condition, the loopback frame of the first target image frame corresponding to the last image frame is determined to be the second preset number of candidate image frames.
10. The method according to claim 1, characterized in that, Determining the feature matching degree between each of the first preceding image frames and the last image frame includes: Feature points of each image frame are extracted, and pairwise matching is performed on each image frame based on the feature points of each image frame to determine the feature matching relationship between each image frame; Based on the feature matching relationship between each of the image frames, the feature matching degree between each of the first preceding image frames and the last image frame is determined.
11. The method according to claim 1, characterized in that, The method further includes: The calculated pose of each image frame is determined based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame. After determining at least one loopback frame of the first target image frame corresponding to the last image frame, the method further includes: The calculated pose of each loopback frame is determined from the calculated pose of each of the image frames.
12. The method according to claim 11, characterized in that, The step of determining the calculated pose of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame includes: Based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame, the three-dimensional coordinates of each image frame are determined. The calculated pose of each image frame is determined based on the feature matching relationship between each image frame and the three-dimensional coordinates of each image frame.
13. The method according to claim 11, characterized in that, The step of determining the calculated pose of each image frame based on the feature matching relationship between each image frame and the initial pose corresponding to each image frame includes: The calculated pose of each image frame is determined based on the feature matching relationship between each image frame and the second optimized pose of each image frame.
14. The method according to claim 6, characterized in that, The step of determining the second target image frame corresponding to the last target image frame, which has the highest feature matching degree with the first target image frame corresponding to the last image frame, from at least one surrounding image frame includes: The first target image frame corresponding to the last image frame is rotated to the same direction as each of the surrounding image frames to obtain a plurality of rotated first target image frames, and the rotated first target image frames correspond one-to-one with the surrounding image frames. The feature matching relationship between each of the rotated first target image frames and each of the surrounding image frames is determined respectively; Based on the feature matching relationship between each first target image frame after rotation and each of the surrounding image frames, the second target image frame corresponding to the last image frame with the highest feature matching degree with the first target image frame corresponding to the last image frame is determined in at least one of the surrounding image frames.
15. The method according to claim 1, characterized in that, The robot is equipped with an image acquisition device, and the acquisition of multiple image frames captured by the robot includes: The robot acquires multiple image frames through the image acquisition device.
16. The method according to claim 15, characterized in that, The image acquisition device is a monocular camera, a binocular camera, or a depth camera.
17. The method according to claim 2, characterized in that, The first objective function is: Where n represents the number of the multiple image frames, T i T' represents the first optimized pose of each of the aforementioned image frames. k T' represents the initial pose of each loopback frame. cal_k This represents the calculated pose of each loopback frame based on the initial pose.
18. The method according to claim 17, characterized in that, The initial pose that minimizes the first objective function is the first optimized pose of each image frame.
19. The method according to claim 3, characterized in that, The second objective function is: Among them, T k T represents the second optimized pose of each of the aforementioned image frames. end T represents the initial pose of the last image frame. cal_end This represents the calculated pose of the last image frame.
20. A computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method according to any one of claims 1 to 19.