Robot pose determination method and apparatus, robot, and storage medium

By acquiring top-view images using an RGB camera on the top of the robot and matching the images with the highest similarity in a visual map information database, and using the two-dimensional coordinates of candidate feature points and the three-dimensional coordinates of reference feature points, the PNP algorithm is employed for precise positioning. This solves the problem that traditional indoor positioning is easily affected by the movement of natural features, and achieves higher positioning accuracy and stability.

CN116168080BActive Publication Date: 2026-06-02SHENZHEN PUDU TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN PUDU TECH CO LTD
Filing Date
2022-12-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In traditional indoor positioning technology, robot positioning is easily affected by the movement of natural features in front, leading to positioning failure or drift and affecting the positioning effect.

Method used

The robot acquires top-view images using an RGB camera on top, matches the images with the highest similarity in a visual map database, determines the robot's pose using the two-dimensional coordinates of candidate feature points and the three-dimensional coordinates of reference feature points, and employs the PNP algorithm for precise localization.

Benefits of technology

It improves the accuracy of robot positioning, avoids positioning failure or drift, and enhances the stability and precision of positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168080B_ABST
    Figure CN116168080B_ABST
Patent Text Reader

Abstract

The application relates to a robot pose determination method and device, a robot and a storage medium. The method comprises the following steps: collecting a top-view image in a driving area by an RGB camera arranged on the top of the robot, taking the top-view image as a current positioning image, obtaining a target top-view image with the maximum similarity to the current positioning image from a visual map information library, and determining a current robot pose when the robot collects the current positioning image according to two-dimensional coordinates of candidate feature points in the current positioning image and three-dimensional coordinates of reference feature points in the target top-view image. The visual map information library comprises a plurality of top-view images collected by the robot at different poses in the driving area. The current robot pose can be determined based on the top-view image by the above method. Most of the features in the top-view image are static and cannot be moved, so that positioning failure or drift is avoided, and the accuracy of robot positioning is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and in particular to a method, apparatus, robot, and storage medium for determining the pose of a robot. Background Technology

[0002] With the advancement of technology, more and more robots are being used in people's daily lives. Among these applications, accurate localization of robots within a known map is fundamental for robots to perform complex tasks.

[0003] In traditional indoor positioning technology, the robot is usually positioned by using forward-looking lasers to obtain natural features in front of it. These natural features are usually objects such as tables, chairs, and sofas on the ground. These objects may be moved, which may cause the positioning to fail or drift, thus affecting the positioning effect. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, robot, and storage medium for determining the pose of a robot to address the aforementioned technical problems.

[0005] In a first aspect, this application provides a method for determining the pose of a robot, comprising:

[0006] The current positioning image is acquired through an RGB camera; the current positioning image is a top-view image collected by the robot in the driving area.

[0007] Obtain the top-view image of the target with the highest similarity to the current positioning image from the visual map information database; the visual map information database includes multiple top-view images collected by the robot in different poses in the driving area;

[0008] Based on the two-dimensional coordinates of candidate feature points in the current localization image and the three-dimensional coordinates of reference feature points in the top view image of the target, the current robot pose when the robot acquires the current localization image is determined.

[0009] In one embodiment, both candidate feature points and reference feature points include multiple points; the current robot pose when the robot acquires the current positioning image is determined based on the two-dimensional coordinates of the candidate feature points in the current positioning image and the three-dimensional coordinates of the reference feature points in the target top view image, including:

[0010] Obtain the feature vectors of each candidate feature point in the current localization image and the feature vectors of each reference feature point in the top view image of the target;

[0011] Based on the feature vectors of each candidate feature point and the feature vectors of each reference feature point, determine the matching feature point pairs between the current localization image and the top view image of the target;

[0012] The robot pose when acquiring the current localization image is determined by the two-dimensional coordinates of the candidate feature points in the matched feature point pair in the current localization image, and the three-dimensional coordinates of the reference feature points in the matched feature point pair.

[0013] In one embodiment, determining the matching feature point pairs between the current localization image and the target top view image based on the feature vectors of each candidate feature point and each reference feature point includes:

[0014] Obtain the cosine value between the feature vector of each candidate feature point and the feature vector of each reference feature point;

[0015] Candidate feature points and reference feature points corresponding to cosine values ​​that satisfy the cosine value constraint are determined as matching feature point pairs.

[0016] In one embodiment, the method further includes:

[0017] Multiple first measurement poses are obtained by acquiring the robot's measurement pose at different times through preset sensors;

[0018] The second measurement pose corresponding to the current location image is obtained based on the two first measurement poses before and after the current location image.

[0019] The robot pose at each moment corresponding to the first measurement pose is determined based on the current robot pose and the second measurement pose.

[0020] In one embodiment, determining the robot pose at each moment corresponding to the first measured pose based on the current robot pose and the second measured pose includes:

[0021] Obtain the relative pose between the current robot pose and the second measured pose;

[0022] When the time of the first measured pose is after the time of the current positioning image and before the time of the next positioning image, the robot pose at the time corresponding to the first measured pose is obtained by using the first measured pose and the relative pose.

[0023] Secondly, this application also provides a method for establishing a visual map information database for a robot, which is applied in the pose determination method described in any of the above claims. The method includes:

[0024] Multiple top-view images of the driving area are acquired using an RGB camera. Each top-view image includes multiple reference feature points.

[0025] The robot's measurement poses at different times are obtained by pre-set sensors, resulting in multiple third measurement poses. A fourth measurement pose corresponding to the top view image is then determined based on the third measurement poses.

[0026] A visual map information database is constructed based on each top-view image, the feature points in each top-view image, and the fourth measurement pose corresponding to each top-view image.

[0027] In one embodiment, a visual map information database is constructed based on each top-view image, feature points in each top-view image, and a fourth measurement pose corresponding to each top-view image, including:

[0028] The three-dimensional coordinates of the reference feature points are determined based on the top-view images corresponding to two adjacent fourth measurement poses.

[0029] Loop closure processing is performed on multiple fourth measurement poses to correct them, resulting in corrected image poses.

[0030] A preset optimization algorithm is used to optimize the poses of multiple images and the 3D coordinates of reference feature points, resulting in multiple optimized image poses and multiple optimized 3D coordinates.

[0031] By associating each top-view image, the corresponding reference feature points of the top-view image, the optimized 3D coordinates of the reference feature points, and the poses of multiple optimized images, a visual map information database is formed.

[0032] Thirdly, this application also provides a robot pose determination device, comprising:

[0033] The first acquisition module is used to acquire the current positioning image through an RGB camera; wherein, the current positioning image is a top-view image collected by the robot in the driving area;

[0034] The second acquisition module is used to acquire the top view image of the target with the highest similarity to the current positioning image from the visual map information database; wherein, the visual map information database includes multiple top view images collected by the robot in different poses in the driving area;

[0035] The pose determination module is used to determine the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image.

[0036] Fourthly, this application also provides a robot, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0037] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0038] In the aforementioned robot pose determination method, visual map information database establishment method, device, robot, and storage medium, an RGB camera mounted on top of the robot acquires a top-view image in the driving area, which serves as the current positioning image. The visual map information database then retrieves a target top-view image with the highest similarity to the current positioning image. Based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top-view image, the current robot pose at the time the current positioning image was acquired is determined. The visual map information database includes multiple top-view images acquired by the robot in different poses within the driving area. This method enables the determination of the current robot pose based on top-view images. Since the top-view images primarily contain static features that are not moved, positioning failures or drifts are avoided, significantly improving the accuracy of robot positioning. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating a robot pose determination method in one embodiment;

[0040] Figure 2 This is a flowchart illustrating a robot pose determination method in another embodiment;

[0041] Figure 3 This is a flowchart illustrating the process of determining matching feature point pairs in one embodiment;

[0042] Figure 4 This is a flowchart illustrating a robot pose determination method in another embodiment;

[0043] Figure 5 This is a schematic diagram illustrating the correspondence between a positioning image and a measured pose in one embodiment;

[0044] Figure 6 This is a flowchart illustrating the process of determining the robot pose at each corresponding moment of the first measurement pose in one embodiment.

[0045] Figure 7 This is a schematic diagram illustrating the process of constructing a visual map information database in one embodiment;

[0046] Figure 8 This is a flowchart illustrating the process of constructing a visual map information database in another embodiment;

[0047] Figure 9 This is a structural block diagram of a robot pose determination device in one embodiment;

[0048] Figure 10 This is a structural block diagram of a device for establishing a visual map information database for a robot in one embodiment;

[0049] Figure 11 This is a diagram of the internal structure of a robot in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] In one embodiment, such as Figure 1 As shown, a method for determining the pose of a robot is provided. This embodiment illustrates the application of this method to a robot, and includes the following steps:

[0052] S110: Acquire the current location image through an RGB camera.

[0053] The current positioning image is a top-view image captured by the robot within its travel area. This image is used to determine the robot's current pose when the image was captured. The robot's travel area refers to the ground region it traverses, and the top-view image is the image of the corresponding top region within that ground region.

[0054] In an optional embodiment, the RGB camera is specifically positioned on the top of the robot, with its main viewing axis pointing upwards. Specifically, it can be vertically upwards or tilted upwards, thereby obtaining a top-view image as the current positioning image.

[0055] In the optional scenarios, if the driving area is an indoor scene, its top-view image can specifically be an image of the indoor ceiling. If the driving area is an outdoor scene, its top-view image includes objects above the robot, such as the sky, leaves, eaves, etc., without limitation.

[0056] In other embodiments, the RGB camera can also be set in other places on the robot, as long as it can acquire top-view images; no specific limitation is made here.

[0057] Optionally, the current positioning image can be a top-view image acquired when the robot first determines its pose in the driving area, or a top-view image acquired when the robot's pose is not determined for the first time.

[0058] In an optional embodiment, the currently located image may specifically be an RGB image.

[0059] Optionally, an RGB camera is installed on the top of the robot. While the robot is moving within the designated area, the RGB camera can capture a top-view image at the current moment, which serves as the aforementioned current positioning image. The RGB camera can acquire top-view images when triggered by the operator, or it can automatically acquire top-view images periodically. Optionally, the robot can also be equipped with other image acquisition devices besides the RGB camera to acquire the current positioning image; this embodiment does not specifically limit the type of image acquisition device.

[0060] S120. Obtain the top-view image of the target that has the highest similarity to the current positioning image from the visual map information database.

[0061] The visual map information database includes multiple top-view images captured by the robot in different poses within the driving area.

[0062] Optionally, the robot uses an image similarity algorithm to calculate the similarity between the current localization image and each frame of the top-view image in the visual map information database, and selects the top-view image with the highest similarity as the target top-view image.

[0063] Optionally, the image similarity algorithm described above can be a structural similarity (SSIM) algorithm, which determines the similarity between images by comparing brightness, contrast, and structure between the current localized image and the top-view image. Alternatively, the image similarity algorithm can be a vector similarity algorithm, which represents the current localized image and the top-view image as vectors and calculates the cosine value between the two vectors as the image similarity. Another option is a histogram similarity algorithm, which calculates the histograms of the current localized image and the top-view image, and then calculates the normalized correlation coefficient (such as Bach distance, histogram intersection distance, etc.) between the two histograms to obtain the image similarity. Finally, the image similarity algorithm can be a hash similarity algorithm, which determines the image similarity by calculating the hash values ​​of the current localized image and the top-view image.

[0064] S130. Based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image, determine the current robot pose when the robot acquires the current positioning image.

[0065] Here, candidate feature points are feature points in the current localization image, and reference feature points are feature points in the top-view image of the target. Optionally, the aforementioned candidate and reference feature points can be ORB (Oriented FAST and Rotated BRIEF) points, which can be detected based on the FAST (features from accelerated segment test) algorithm, or they can be other feature points, such as SIFT points or other deep learning points, and obtained using the corresponding detection algorithm.

[0066] Optionally, the visual map information database also includes the three-dimensional coordinates of reference feature points in each top-view image in the real scene. After the robot determines the target top-view image in the visual map information database, it can determine the current robot pose when acquiring the current positioning image based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top-view image.

[0067] It should be noted that the reference feature points in the top-view image are obtained after global optimization (Global BA) to minimize the reprojection error of each feature point in the top-view image and then remove feature points whose reprojection error is greater than the error threshold. That is, the top-view image includes multiple feature points, and some feature points are selected as reference feature points in the top-view image based on the reprojection error of each feature point.

[0068] Optionally, the robot can extract candidate feature points in the current localization image through a feature point recognition model, and then obtain the two-dimensional coordinates of each candidate feature point in the current localization image. The robot further determines matching feature point pairs between candidate feature points and reference feature points, and uses the PNP (Perspective-n-Point) algorithm to determine the current robot pose when acquiring the current localization image based on the two-dimensional coordinates of the candidate feature points and the three-dimensional coordinates of the reference feature points in the matching feature point pairs.

[0069] In this embodiment, the robot acquires top-view images of the driving area using an RGB camera mounted on its top, which serves as the current positioning image. It then retrieves the target top-view image with the highest similarity to the current positioning image from a visual map information database. Based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top-view image, the robot's current pose when acquiring the current positioning image is determined. The visual map information database includes multiple top-view images acquired by the robot in different poses within the driving area. This method enables the determination of the current robot pose based on top-view images. Since top-view images primarily contain static features that are not moved, positioning failures or drifts are avoided, significantly improving the accuracy of robot positioning. Furthermore, the three-dimensional coordinates of feature points in the image accurately reflect their spatial location and do not change with the robot's pose. Therefore, determining the current robot pose when acquiring the current positioning image based on the two-dimensional coordinates of the feature points in the current positioning image and their actual three-dimensional coordinates further improves the accuracy of the determined current robot pose, thereby enhancing the positioning effect.

[0070] In practical applications, both candidate feature points and reference feature points include multiple points, such as... Figure 2 As shown, S130 above, determining the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image, includes:

[0071] S210. Obtain the feature vectors of each candidate feature point in the current localization image and the feature vectors of each reference feature point in the target top view image.

[0072] The feature vector is used to characterize the orientation and position of the feature point in the image. Optionally, any position in the current localization image can be used as a reference point, and the vector formed by the reference point and the candidate feature point is the feature vector of the candidate feature point in the current localization image. Similarly, any position in the top-view image of the target can be used as a reference point, and the vector formed by the reference point and the reference feature point is the feature vector of the reference feature point in the top-view image of the target.

[0073] It should be noted that the reference point in the current positioning image and the reference point in the target top-view image are at the same position. For example, the reference point in the current positioning image is the center point of the current positioning image, while the reference point in the target top-view image is the center point of the target top-view image.

[0074] Optionally, the robot vectorizes the candidate feature points in the current localization image and the reference feature points in the target top view image, determines the feature vector of each candidate feature point through the reference points in the current localization image, and determines the feature vector of each reference feature point through the reference points in the target top view image.

[0075] S220. Based on the feature vectors of each candidate feature point and the feature vectors of each reference feature point, determine the matching feature point pairs between the current localization image and the top view image of the target.

[0076] Optionally, the robot can determine matching feature point pairs based on the similarity between each candidate feature point and each reference feature point. In this embodiment, the robot vectorizes the candidate and reference feature points to determine the similarity between each candidate and reference feature point based on the feature vectors of each candidate and reference feature points. Then, based on the obtained similarity, the robot determines the matching feature point pairs between the current localization image and the target top view image among multiple candidate and reference feature points.

[0077] S230. Determine the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of the candidate feature points in the matched feature point pair in the current positioning image and the three-dimensional coordinates of the reference feature points in the matched feature point pair.

[0078] Optionally, after determining the matching feature point pairs, the robot can use the PNP algorithm to determine the robot pose when acquiring the current localization image based on the two-dimensional coordinates of the candidate feature points and the three-dimensional coordinates of the reference feature points in the matching feature point pairs. The matching feature point pairs must include at least three pairs.

[0079] In the above embodiments, the cosine value between vectors can be used to characterize the similarity between each candidate feature point and each reference feature point, such as... Figure 3 As shown, S220 above, based on the feature vectors of each candidate feature point and the feature vectors of each reference feature point, determines the matching feature point pairs between the current localization image and the target top view image, including:

[0080] S310. Obtain the cosine value between the feature vector of each candidate feature point and the feature vector of each reference feature point.

[0081] Optionally, the robot uses the cosine calculation formula between vectors to calculate the cosine value between the feature vector of each candidate feature point and the feature vector of each reference feature point, thereby obtaining the cosine value between the feature vector of each candidate feature point and the feature vector of any reference feature point.

[0082] S320. The candidate feature points and reference feature points corresponding to the cosine values ​​that satisfy the cosine value constraint are determined as the matching feature point pairs.

[0083] Among them, the cosine value constraint condition includes a cosine value greater than a cosine value threshold.

[0084] Optionally, the robot compares the obtained cosine value with a preset cosine value threshold, and determines the candidate feature point and reference feature point corresponding to the cosine value greater than the threshold as a matching feature point pair. For example, if the current localization image includes candidate feature points A, B, C, and D, and the center point O of the current localization image is used as the reference point, the corresponding feature vector is obtained. as well as The top-view image of the target includes reference feature points a, b, and c. Taking the center point o of the top-view image of the target as the reference point, the corresponding feature vector is obtained. After calculating the cosine values ​​between the vectors, we obtain and The cosine value Cosθ1 is greater than the cosine threshold Cosθ0. and The cosine value Cosθ2 is greater than the cosine threshold Cosθ0. and If the cosine value Cosθ3 between the two points is greater than the cosine value threshold Cosθ0, the robot can determine that candidate feature point A and reference feature point a are a matching feature point pair, candidate feature point B and reference feature point b are a matching feature point pair, and candidate feature point C and reference feature point c are a matching feature point pair.

[0085] Optionally, if there exist cosine values ​​between the feature vector of the same reference feature point and the feature vectors of at least two candidate feature points that are all greater than a cosine threshold, then the reference feature point and candidate feature point corresponding to the maximum cosine value are determined as a matched feature point pair. For example, continuing the above example, if we obtain... and The cosine value Cosθ1 between them is greater than the cosine threshold Cosθ0, and and If the cosine value Cosθ4 between the two points is also greater than the cosine threshold Cosθ0, where Cosθ1>Cosθ4, the robot can determine that the candidate feature point A and the reference feature point a are a matching feature point pair.

[0086] In this embodiment, both candidate feature points and reference feature points include multiple points. The robot acquires the feature vectors of each candidate feature point in the current positioning image and the feature vectors of each reference feature point in the target top-view image. Then, based on the feature vectors of each candidate feature point and each reference feature point, it determines the matching feature point pairs between the current positioning image and the target top-view image. The robot's current pose is determined based on the two-dimensional coordinates of the candidate feature points in the matching feature point pairs in the current positioning image and the three-dimensional coordinates of the reference feature points in the matching feature point pairs. Specifically, this is based on vectorization processing, using the cosine value between the feature vectors of each candidate feature point and each reference feature point to characterize the similarity between vectors, thereby determining the matching feature point pairs. This method improves the matching accuracy of feature point pairs. Based on the accurately matched feature point pairs, the PNP algorithm is used to solve for the current robot pose, thus improving the accuracy of the determined current robot pose.

[0087] In applications where robots periodically and automatically acquire ceiling images for localization, to suit real-time positioning, such as... Figure 4 As shown, the above method also includes:

[0088] S410. Obtain the robot's measurement pose at different times through preset sensors to obtain multiple first measurement poses.

[0089] Among them, the preset sensor can directly acquire the robot's pose, and the robot's pose acquired by the preset sensor is called the measured pose. Optionally, the preset sensor can be an odometer installed on the robot, such as a wheeled odometer, or an IMU (Inertial Measurement Unit).

[0090] Optionally, while the robot is driving in the driving area, the robot can continuously acquire the robot's measurement pose at different times through preset sensors, and use it as the first measurement pose to obtain multiple first measurement poses.

[0091] S420. Based on the two first measurement poses before and after the current location image, obtain the second measurement pose corresponding to the current location image.

[0092] The current location image is located at the current moment of sampling. It should be noted that the sampling frequencies of the preset sensor and the RGB camera are different. Generally, the sampling frequency of the preset sensor for pose sampling is higher than that of the RGB camera. In order to determine the current location image and the measured pose at the same moment, it is necessary to obtain the second measured pose corresponding to the current location image moment based on the two first measured poses before and after the current location image moment.

[0093] Optionally, among multiple first measurement poses, the robot obtains two first measurement poses whose measurement time is adjacent to or before the time of the current localization image (i.e., the current time). These two first measurement poses can be interpolated to obtain a second measurement pose corresponding to the time of the current localization image. For example... Figure 5 As shown, for the current positioning image A, the first measurement pose Pose1 and Pose2, which are adjacent to the current positioning image A at the measurement time, are interpolated to obtain the second measurement pose PoseA corresponding to the time of the current positioning image A. Optionally, the second measurement pose PoseA can be obtained by using common interpolation algorithms, which are not limited here.

[0094] For the current positioning image B, interpolation is performed on the first measurement poses Pose4 and Pose5, which are adjacent to the current positioning image B at the measurement time, to obtain the second measurement pose PoseB corresponding to the time of the current positioning image B. Optionally, the second measurement pose PoseB can be obtained using common interpolation algorithms, which are not limited here.

[0095] S430. Determine the robot pose at each moment corresponding to the first measurement pose based on the current robot pose and the second measurement pose.

[0096] Optionally, after obtaining the current robot pose and the second measured pose at the current moment, the robot can determine the robot pose at the corresponding moment of each first measured pose based on the relative relationship between the current robot pose and the second measured pose.

[0097] In an alternative embodiment, such as Figure 6 As shown, S430 above, determining the robot pose at each moment corresponding to the first measurement pose based on the current robot pose and the second measurement pose, includes:

[0098] S610. Obtain the relative pose between the current robot pose and the second measured pose.

[0099] Optionally, the robot can calculate the relative pose dT between the current robot pose T0 and the second measured pose T_o0 according to the following formula (1):

[0100] dT = T_o0 * T0_inv (Formula 1)

[0101] Where T_o0 represents the second measured pose, and T0_inv represents the inverse matrix of the current robot pose.

[0102] It should be noted that the current positioning image is updated periodically, and correspondingly, the relative pose dT is also updated as the current positioning image is updated.

[0103] S620. When the time of the first measured pose is after the time of the current positioning image and before the time of the next positioning image, the robot pose at the time corresponding to the first measured pose is obtained by using the first measured pose and the relative pose.

[0104] Optionally, for multiple first measurement poses acquired by the robot, if the time of the first measurement pose is after the time of the current positioning image and before the time of the next positioning image, that is, for the first measurement poses acquired during the period when the robot has acquired the current positioning image but has not yet acquired the next positioning image, the robot pose at the corresponding time of each first measurement pose is obtained by comparing each first measurement pose with the relative pose. Continuing from the above... Figure 5 For example, when positioning image A is the current positioning image, positioning image B is the next positioning image of the current positioning image A. For each first measurement pose between the current positioning image A and positioning image B, namely Pose2, Pose3 and Pose4, the robot can take Pose2 as T_o0, and substitute Pose2 as T_o0 and the previously calculated relative pose dT into formula (1) to obtain the robot pose inverse matrix at the time corresponding to Pose2, and then obtain the robot pose Pose2' at the corresponding time. Similarly, the robot poses Pose3', Pose4' and Pose5' at the corresponding times are obtained respectively.

[0105] In this embodiment, the robot acquires its measured poses at different times using preset sensors, obtaining multiple first measured poses. Based on the two first measured poses preceding and following the current location image, a second measured pose corresponding to the current location image is obtained. Then, the robot pose at each time corresponding to the first measured pose is determined based on the current robot pose and the second measured pose. Specifically, the relative pose between the current robot pose and the second measured pose can be obtained. When the time of the first measured pose is after the time of the current location image and before the time of the next location image, the robot pose at the time corresponding to the first measured pose is obtained through the first measured pose and the relative pose. This method achieves the determination of the robot pose at the corresponding time between the current location image and the next location image, ensuring a high positioning frequency while improving positioning accuracy, making it suitable for real-time positioning.

[0106] This application also provides a method for constructing a visual map information database for a robot, which is applied to the pose determination method described in any of the above embodiments. For example... Figure 7 As shown, the method includes:

[0107] The S710 acquires multiple top-view images of the driving area using an RGB camera. These top-view images include multiple reference feature points.

[0108] Optionally, when the robot constructs the visual map information database corresponding to the driving area, in an optional embodiment, when the driving area is a circular area, it can drive around the driving area clockwise once and counterclockwise once. When the driving area is a non-circular area, it can drive from the starting point of its driving area to the destination and then return from the destination to the starting point.

[0109] If there is a side lane, the player should first push into the side lane and then return.

[0110] During the robot's movement, multiple top-view images of the driving area, including multiple reference feature points, are acquired using an RGB camera.

[0111] S720: Obtain the robot's measurement pose at different times through preset sensors, obtain multiple third measurement poses, and determine the fourth measurement pose corresponding to the top view image based on the third measurement poses.

[0112] Optionally, during the process of the robot moving around the driving area, the robot can obtain the robot's measurement pose at different times through preset sensors, which can be used as the third measurement pose to obtain multiple third measurement poses.

[0113] As mentioned earlier, the sampling frequency of the preset sensor used for pose sampling is usually higher than that of the RGB camera. In order to determine the top view image and the measured pose at the same moment, it is necessary to obtain the fourth measured pose corresponding to the moment of the top view image based on the two third measured poses before and after the moment of the top view image (such as by interpolation).

[0114] S730. Construct a visual map information database based on each top view image, the feature points in each top view image, and the fourth measurement pose corresponding to each top view image.

[0115] In an alternative embodiment, such as Figure 8 As shown, the above-mentioned S730, which constructs a visual map information database based on each top-view image, the feature points in each top-view image, and the fourth measurement pose corresponding to each top-view image, includes:

[0116] S810: Determine the three-dimensional coordinates of the reference feature point based on the top view images corresponding to two adjacent fourth measurement poses.

[0117] Optionally, after determining the fourth measurement pose corresponding to the time of each top view image, the robot further determines the three-dimensional coordinates of the reference feature point in the real scene in the top view image based on the top view images corresponding to two adjacent fourth measurement poses. Specifically, the PNP algorithm can be used to obtain this coordinates, which will not be elaborated here.

[0118] S820: Perform loop closure processing on multiple fourth measurement poses to correct the multiple fourth measurement poses, and obtain multiple corrected image poses.

[0119] Optionally, during the process of acquiring top-view images while the robot is circling the driving area, the robot can use a "word band" model for loop closure detection. If a loop is detected, it indicates that the robot has completed a full circle of the driving area and returned to its starting position. At this point, the robot can use a pose correction algorithm to correct pose drift and obtain corrected poses for multiple images.

[0120] S830: A preset optimization algorithm is used to optimize the poses of multiple images and the three-dimensional coordinates of reference feature points, resulting in multiple optimized image poses and multiple optimized three-dimensional coordinates.

[0121] Optionally, after obtaining the corrected image pose, the robot uses the corrected image pose to recalculate the 3D coordinates of all feature points in the top-view image and performs global optimization to minimize the reprojection error and remove feature points whose reprojection error exceeds the error threshold. After the above series of corrections and optimizations, multiple optimized image poses and multiple optimized 3D coordinates are obtained.

[0122] S840. Associate each top view image, the corresponding reference feature point of the top view image, the optimized 3D coordinates of the reference feature point, and the poses of multiple optimized images to form a visual map information database.

[0123] Optionally, the robot associates each top-view image, the corresponding reference feature point of the top-view image, the optimized 3D coordinates of the reference feature point, and multiple optimized image poses to form a visual map information library. That is, the visual map information library includes each frame of top-view image acquired by the robot, as well as the corresponding reference feature point in the top-view image, the 3D coordinates of the reference feature point, and the image pose when the top-view image of that frame was acquired.

[0124] Optionally, when the visual map information database includes the image pose corresponding to each frame of top view image, when determining the current robot pose of the current localization image, the robot can obtain the image pose corresponding to the target top view image with the highest similarity to the current localization image from the visual map information database, and use it as the current robot pose.

[0125] In the above embodiments, before constructing the visual map information database, it is also necessary to calibrate the parameters of the RGB camera and preset sensors (such as wheeled odometers) integrated / installed on the robot. For example, the parameters to be calibrated include: the camera intrinsic parameters of the camera device (K = [fx, fy, cx, cy], distortion coefficient d = [d1, d2, d3, d4]), the wheeled odometer extrinsic parameters (left wheel radius r1, right wheel radius r2, diameter D between the two wheels), and the extrinsic parameters between the RGB camera and the wheeled odometer.

[0126] In this embodiment, a visual map information database is pre-established so that the robot can directly locate itself based on the data in the visual map information database, thereby further improving the robot's pose determination efficiency.

[0127] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0128] In one embodiment, such as Figure 9 As shown, this application provides a robot pose determination device, including: a first acquisition module 901, a second acquisition module 902, and a pose determination module 903, wherein:

[0129] The first acquisition module 901 is used to acquire the current positioning image through an RGB camera; wherein, the current positioning image is a top-view image collected by the robot in the driving area;

[0130] The second acquisition module 902 is used to acquire the top view image of the target with the highest similarity to the current positioning image from the visual map information database; wherein, the visual map information database includes multiple top view images collected by the robot in different poses in the driving area;

[0131] The pose determination module 903 is used to determine the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image.

[0132] In one embodiment, both candidate feature points and reference feature points include multiple points; the pose determination module 903 is specifically used for:

[0133] Obtain the feature vectors of each candidate feature point in the current localization image and the feature vectors of each reference feature point in the target top view image; determine the matching feature point pairs between the current localization image and the target top view image based on the feature vectors of each candidate feature point and each reference feature point; determine the current robot pose when the robot acquires the current localization image based on the two-dimensional coordinates of the candidate feature points in the matching feature point pairs in the current localization image and the three-dimensional coordinates of the reference feature points in the matching feature point pairs.

[0134] In one embodiment, the pose determination module 903 is specifically used for:

[0135] Obtain the cosine value between the feature vector of each candidate feature point and the feature vector of each reference feature point; determine the candidate feature point and the reference feature point corresponding to the cosine value that satisfies the cosine value constraint as the matching feature point pair.

[0136] In one embodiment, the above-described apparatus further includes a real-time positioning module, used for:

[0137] The robot's measurement poses at different times are obtained by using preset sensors to obtain multiple first measurement poses; a second measurement pose corresponding to the current location image is obtained based on the two first measurement poses before and after the current location image; the robot pose at the time corresponding to each first measurement pose is determined according to the current robot pose and the second measurement pose.

[0138] In one embodiment, the real-time positioning module is specifically used for:

[0139] Obtain the relative pose between the current robot pose and the second measured pose; when the time of the first measured pose is after the time of the current positioning image and before the time of the next positioning image, obtain the robot pose at the time corresponding to the first measured pose through the first measured pose and the relative pose.

[0140] In one embodiment, such as Figure 10 As shown, this application provides an apparatus for establishing a visual map information database for a robot, which is applied in the pose determination method described in any of the above claims. The establishment method includes: an acquisition module 1001, a measurement module 1002, and a construction module 1003, wherein:

[0141] Image acquisition module 1001 is used to acquire multiple top-view images of the driving area through an RGB camera. The top-view images include multiple reference feature points.

[0142] The pose measurement module 1002 is used to acquire the robot's measurement pose at different times through preset sensors, obtain multiple third measurement poses, and determine the fourth measurement pose corresponding to the top view image based on the third measurement poses.

[0143] The map construction module 1003 is used to construct a visual map information database based on each top view image, the feature points in each top view image, and the fourth measurement pose corresponding to each top view image.

[0144] In one embodiment, the map building module 1003 is specifically used for:

[0145] The 3D coordinates of reference feature points are determined based on the top-view images corresponding to two adjacent fourth measurement poses; loop closure processing is performed on multiple fourth measurement poses to correct them, resulting in multiple corrected image poses; a preset optimization algorithm is used to optimize the multiple image poses and the 3D coordinates of the reference feature points, resulting in multiple optimized image poses and multiple optimized 3D coordinates; each top-view image, the reference feature points of the corresponding top-view image, the optimized 3D coordinates of the reference feature points, and the multiple optimized image poses are associated to form a visual map information database.

[0146] In one embodiment, a robot is provided whose internal structure diagram can be as follows: Figure 11 As shown, the computer device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for determining the pose of a robot and a method for establishing a visual map information database for the robot. The robot's display screen can be an LCD screen or an e-ink screen. The robot's input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the robot's shell, or an external keyboard, touchpad, or mouse.

[0147] Those skilled in the art will understand that Figure 11The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0148] In one embodiment, a robot is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the methods described above.

[0149] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0150] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0151] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0152] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0153] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining the pose of a robot, characterized in that, The method includes: The current positioning image is acquired using an RGB camera; wherein, the current positioning image is a top-view image collected by the robot in the driving area; Obtain the top-view image of the target with the highest similarity to the current positioning image from the visual map information database; wherein, the visual map information database includes multiple top-view images collected by the robot in different poses in the driving area; The current robot pose when the robot acquires the current positioning image is determined based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image. Multiple first measurement poses are obtained by acquiring the robot's measurement pose at different times through preset sensors; A second measurement pose corresponding to the current location of the image is obtained based on the two first measurement poses before and after the current location image. Obtain the relative pose between the current robot pose and the second measured pose; When the time of the first measured pose is after the time of the current positioning image and before the time of the next positioning image, the robot pose at the time corresponding to the first measured pose is obtained by comparing the first measured pose with the relative pose.

2. The method according to claim 1, characterized in that, Both the candidate feature points and the reference feature points include multiple points; determining the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of the candidate feature points in the current positioning image and the three-dimensional coordinates of the reference feature points in the target top view image includes: Obtain the feature vectors of each candidate feature point in the current positioning image and the feature vectors of each reference feature point in the target top view image; Based on the feature vectors of each candidate feature point and the feature vectors of each reference feature point, a matching pair of feature points is determined between the current positioning image and the target top view image; The current robot pose when the robot acquires the current positioning image is determined based on the two-dimensional coordinates of the candidate feature points in the matched feature point pair in the current positioning image and the three-dimensional coordinates of the reference feature points in the matched feature point pair.

3. The method according to claim 2, characterized in that, The step of determining the matching feature point pairs between the current positioning image and the target top view image based on the feature vectors of each candidate feature point and the feature vectors of each reference feature point includes: Obtain the cosine value between the feature vector of each candidate feature point and the feature vector of each reference feature point; Candidate feature points and reference feature points corresponding to cosine values ​​that satisfy the cosine value constraint are determined as matching feature point pairs.

4. A method for establishing a visual map information database for a robot, characterized in that, The visual map information database is applied in the pose determination method according to any one of claims 1-3, the method comprising: The RGB camera acquires multiple top-view images of the driving area, and each top-view image includes multiple reference feature points. The robot's measurement poses at different times are obtained by using preset sensors to obtain multiple third measurement poses, and a fourth measurement pose corresponding to the top view image is determined based on the third measurement poses. The visual map information database is constructed based on each of the top-view images, the feature points in each of the top-view images, and the fourth measurement pose corresponding to each of the top-view images.

5. The method according to claim 4, characterized in that, The step of constructing the visual map information database based on each of the top-view images, the feature points in each of the top-view images, and the fourth measurement pose corresponding to each of the top-view images includes: The three-dimensional coordinates of the reference feature point are determined based on the top-view images corresponding to two adjacent fourth measurement poses. Loop closure processing is performed on multiple fourth measurement poses to correct them, resulting in multiple corrected image poses. A preset optimization algorithm is used to optimize the multiple image poses and the three-dimensional coordinates of the reference feature points to obtain multiple optimized image poses and multiple optimized three-dimensional coordinates; The visual map information database is formed by associating each of the top-view images, the corresponding reference feature points of the top-view images, the optimized 3D coordinates of the reference feature points, and the multiple optimized image poses.

6. A robot pose determination device, characterized in that, The device includes: The first acquisition module is used to acquire the current positioning image through an RGB camera set on the top of the robot; wherein, the current positioning image is a top-view image collected by the robot in the driving area; The second acquisition module is used to acquire the top view image of the target with the highest similarity to the current positioning image from the visual map information database; wherein, the visual map information database includes multiple top view images collected by the robot in different poses in the driving area; The pose determination module is used to determine the current robot pose when the robot acquires the current positioning image based on the two-dimensional coordinates of candidate feature points in the current positioning image and the three-dimensional coordinates of reference feature points in the target top view image. The real-time positioning module is used to acquire the robot's measurement pose at different times through preset sensors to obtain multiple first measurement poses; to obtain a second measurement pose corresponding to the current positioning image based on two first measurement poses before and after the current positioning image; to obtain the relative pose between the current robot pose and the second measurement pose; and to obtain the robot pose at the time corresponding to the first measurement pose when the time of the first measurement pose is after the current positioning image and before the time of the next positioning image through the first measurement pose and the relative pose.

7. A device for establishing a visual map information database for a robot, characterized in that, The visual map information database is applied in the pose determination method according to any one of claims 1-3, and the apparatus includes: The image acquisition module is used to acquire multiple top-view images of the driving area through the RGB camera, wherein the top-view images include multiple reference feature points; The pose measurement module is used to acquire the robot's measurement pose at different times through preset sensors, obtain multiple third measurement poses, and determine a fourth measurement pose corresponding to the top view image based on the third measurement poses. The map building module is used to build the visual map information database based on each of the top view images, the feature points in each of the top view images, and the fourth measurement pose corresponding to each of the top view images.

8. A robot comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 5.