Method, apparatus, and autonomous robot for global localization, and storage medium

CN117288184BActive Publication Date: 2026-09-18BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210701625.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2026-09-18
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

[0004]然而,GPS容易受卫星状况、电离层状况影响,在隧道、高楼密集区域定位效果较差

Benefits of technology

[0047]In the global positioning method provided in this application embodiment, the pose information estimated by VIO and the global feature point map are combined, and the autonomous robot is positioned by filtering and updating. The positioning process does not require GPS participation, which avoids the problem that GPS is affected by interference in scenarios such as tunnels and densely built-up areas when GPS and VIO are combined, resulting in poor global positioning effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117288184B_ABST
    Figure CN117288184B_ABST
Patent Text Reader

Abstract

The application discloses a global positioning method and device, an autonomous robot and a storage medium, and belongs to the technical field of positioning. The method comprises the following steps: acquiring positioning information, acquiring a scene image and inertial measurement unit (IMU) information, extracting a plurality of two-dimensional feature points of the scene image, determining estimated pose information based on the plurality of two-dimensional feature points and the IMU information, matching the plurality of two-dimensional feature points with a global feature point map to obtain three-dimensional feature points matched by the plurality of two-dimensional feature points in the global feature point map, and filtering and updating the positioning information according to the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point and the estimated pose information. The application can avoid the problem that GPS cannot effectively position due to serious interference in some scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of positioning technology, and in particular to a method, apparatus, autonomous robot, and storage medium for global positioning. Background Technology

[0002] Autonomous robots are a key research area in the field of robotics, and can include drones, among others. Autonomous robots can move autonomously to designated locations to perform tasks through localization and navigation. Therefore, accurate localization is crucial for autonomous robots.

[0003] Currently, autonomous robots typically use the Global Positioning System (GPS) to achieve autonomous positioning.

[0004] However, GPS is easily affected by satellite conditions and ionospheric conditions, and its positioning performance is poor in tunnels and densely populated areas with tall buildings. Summary of the Invention

[0005] This application provides a method, apparatus, autonomous robot, and storage medium for global positioning, achieving positioning without using GPS, thus avoiding the problem of GPS inaccurate positioning in certain locations. The technical solution is as follows:

[0006] Firstly, a global positioning method is provided, the method comprising:

[0007] Obtain location information;

[0008] Acquire scene images and inertial measurement unit (IMU) information;

[0009] Extract multiple two-dimensional feature points from the scene image;

[0010] Based on the multiple two-dimensional feature points and the IMU information, the estimated pose information is determined;

[0011] The multiple two-dimensional feature points are matched with the global feature point map to obtain the three-dimensional feature points matched by the multiple two-dimensional feature points in the global feature point map;

[0012] The positioning information is filtered and updated based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information.

[0013] In one possible implementation, the step of filtering and updating the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information includes:

[0014] If the two-dimensional feature point is a visual inertial odometry (VIO) feature point, and the VIO state variables include the feature point information of the two-dimensional feature point, then the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point, and the estimated pose information are input into the EKF to filter and update the positioning information. The feature point information of the two-dimensional feature point includes the coordinates of the two-dimensional feature point, the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, and the identifier of the two-dimensional feature point.

[0015] In one possible implementation, the method further includes:

[0016] If the two-dimensional feature point is a VIO feature point, and the state variable of the VIO does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0017] In one possible implementation, the step of filtering and updating the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information includes:

[0018] If the two-dimensional feature point is not a VIO feature point, feature tracking is performed on the two-dimensional feature point in the latest acquired scene image. If there is a target two-dimensional feature point in the latest acquired scene image that matches the two-dimensional feature point, the measurement noise is calculated based on the coordinates of the three-dimensional feature point that matches the two-dimensional feature point.

[0019] The coordinates of the two-dimensional feature points, the coordinates of the three-dimensional feature points matched by the two-dimensional feature points, the initial measurement variance corresponding to the three-dimensional feature points matched by the two-dimensional feature points, the estimated pose information, and the measurement noise are input into the EKF to update the positioning information.

[0020] In one possible implementation, the method further includes:

[0021] If the two-dimensional feature point is not a VIO feature point, and the target two-dimensional feature point does not exist in the latest acquired scene image, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0022] In one possible implementation, extracting multiple two-dimensional feature points from the scene image includes:

[0023] Extract a first number of two-dimensional feature points from the scene image, and add VIO feature identifiers to the first number of two-dimensional feature points as the VIO feature points;

[0024] Extract a second number of two-dimensional feature points from the scene image.

[0025] Secondly, a global positioning device is provided, the device comprising:

[0026] The acquisition module is used to acquire location information, scene images, and IMU information;

[0027] An extraction module is used to extract multiple two-dimensional feature points from the scene image;

[0028] The determination module is used to determine the estimated pose information based on the plurality of two-dimensional feature points and the IMU information;

[0029] The matching module is used to match the plurality of two-dimensional feature points with a global feature point map to obtain three-dimensional feature points that are matched by the plurality of two-dimensional feature points in the global feature point map.

[0030] The update module is used to filter and update the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information.

[0031] In one possible implementation, the update module is configured to:

[0032] If the two-dimensional feature point is a VIO feature point, and the state variables of the VIO include the feature point information of the two-dimensional feature point, then the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point, and the estimated pose information are input into the EKF to filter and update the positioning information of this device. The feature point information of the two-dimensional feature point includes the coordinates of the two-dimensional feature point, the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, and the identifier of the two-dimensional feature point.

[0033] In one possible implementation, the update module is further configured to:

[0034] If the two-dimensional feature point is a VIO feature point, and the state variable of the VIO does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0035] In one possible implementation, the update module is configured to:

[0036] If the two-dimensional feature point is not a VIO feature point, feature tracking is performed on the two-dimensional feature point in the latest acquired scene image. If there is a target two-dimensional feature point in the latest acquired scene image that matches the two-dimensional feature point, the measurement noise is calculated based on the coordinates of the second three-dimensional feature point that matches the two-dimensional feature point.

[0037] The coordinates of the two-dimensional feature points, the coordinates of the three-dimensional feature points matched by the two-dimensional feature points, the initial measurement variance corresponding to the three-dimensional feature points matched by the two-dimensional feature points, the estimated pose information, and the measurement noise are input into the EKF to update the positioning information of this device.

[0038] In one possible implementation, the update module is configured to:

[0039] If the two-dimensional feature point is not a VIO feature point, and the target two-dimensional feature point does not exist in the latest acquired scene image, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0040] In one possible implementation, the extraction module is configured to:

[0041] Extract a first number of two-dimensional feature points from the scene image, and add VIO feature identifiers to the first number of two-dimensional feature points as the VIO feature points;

[0042] Extract a second number of two-dimensional feature points from the scene image.

[0043] Thirdly, an autonomous robot is provided, the autonomous robot including a processor and a memory, the memory storing at least one instruction, the instruction being loaded and executed by the processor to perform the operation performed by the global positioning method as described in the first aspect above.

[0044] Fourthly, a computer-readable storage medium is provided, the storage medium storing at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed by the global positioning method as described in the first aspect above.

[0045] Fifthly, a computer program product is provided, the computer program product including at least one instruction, the instruction being loaded and executed by a processor to perform the operations performed by the global positioning method as described in the first aspect above.

[0046] The beneficial effects of the technical solutions provided in this application are:

[0047] In the global positioning method provided in this application embodiment, the pose information estimated by VIO and the global feature point map are combined, and the autonomous robot is positioned by filtering and updating. The positioning process does not require GPS participation, which avoids the problem that GPS is affected by interference in scenarios such as tunnels and densely built-up areas when GPS and VIO are combined, resulting in poor global positioning effect. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart of a VIO method provided in an embodiment of this application;

[0050] Figure 2 This is a flowchart of a global positioning method provided in an embodiment of this application;

[0051] Figure 3 This is a flowchart of a global positioning method provided in an embodiment of this application;

[0052] Figure 4 This is a schematic diagram of a global positioning device structure provided in an embodiment of this application;

[0053] Figure 5 This is a schematic diagram of the structure of an autonomous robot provided in an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0055] To facilitate understanding of this application, the terms used in this application will be introduced below.

[0056] Visual-inertial odometry (VIO):

[0057] VIO combines scene images captured by a visual sensor (camera) with IMU information detected by an inertial measurement unit (IMU) to achieve simultaneous localization and mapping (SLAM).

[0058] The general principle of VIO is as follows: Figure 1As shown, feature extraction and matching are performed on the scene image captured by the vision sensor to obtain two-dimensional feature points, and the IMU information detected by the IMU is integrated. Then, a Multi-State Constraint Kalman Filter (MSCKF) is used to process the integrated results of the two-dimensional feature points and IMU information to obtain the estimated pose information.

[0059] This application provides a global positioning method that can be implemented by an autonomous robot, such as a drone or unmanned delivery vehicle. In this global positioning method, the pose information estimated by VIO (Vehicle Identification and Positioning) and a global feature point map are combined, and a filtering update method is used to achieve the positioning of the autonomous robot. The positioning process does not require GPS involvement, avoiding the problem of poor global positioning performance caused by GPS interference in scenarios such as tunnels and densely populated areas with tall buildings when GPS and VIO are combined.

[0060] The global positioning method provided in the embodiments of this application will be described below. Figure 2 This is a flowchart of the method. See also... Figure 2 The processing flow of this method may include the following steps:

[0061] Step 101: Obtain scene images and IMU information.

[0062] In implementation, the visual sensor acquires scene images at a preset frequency, and the IMU detects IMU information at a preset frequency. This IMU information may include the autonomous robot's acceleration and angular velocity information. After acquiring the scene images, the visual sensor sends them to the processor, and after detecting the IMU information, the IMU sends its information to the processor. Thus, the processor can acquire both the scene images acquired by the visual sensor and the IMU information detected by the IMU.

[0063] Among them, the visual sensor and the IMU can form a VIO, also known as a visual-inertial system (VINS).

[0064] Step 102: Extract multiple two-dimensional feature points from the scene image.

[0065] In implementation, after acquiring the scene image, the processor can extract feature points from the scene image. Specifically, the Feature Points from Accelerated Segment Test (FAST) method can be used to extract feature points from the scene image.

[0066] In addition, the processor can extract more two-dimensional feature points from each preset number of scene images.

[0067] For example, by taking 98 scene images at intervals, a large number of two-dimensional feature points can be extracted. Specifically, for the first scene image, a first number of two-dimensional feature points can be extracted and labeled with VIO feature point identifiers. Then, a second number of two-dimensional feature points can be extracted; these second number of feature points do not need to be labeled with VIO feature point identifiers. For the second to the 99th scene images, the first number of two-dimensional feature points can be extracted and labeled with VIO feature point identifiers for each image. For the 100th scene image, the first number of two-dimensional feature points can be extracted and labeled with VIO feature point identifiers. Then, a second number of two-dimensional feature points can be extracted; these second number of feature points do not need to be labeled with VIO feature point identifiers, and so on. The two-dimensional feature points labeled with VIO feature point identifiers are called VIO feature points.

[0068] Step 103: Determine the estimated pose information based on multiple two-dimensional feature points and IMU information.

[0069] In implementation, for each scene image, the two-dimensional coordinates of the acquired two-dimensional feature points in the camera coordinate system and the currently acquired IMU information can be input into the MSCKF to obtain the current estimated pose information of the device. The estimated pose information includes estimated position information and estimated orientation information. The position information is the coordinates of the autonomous robot in the world coordinate system, and the orientation information is the orientation of the autonomous robot in the world coordinate system.

[0070] Step 104: Match multiple two-dimensional feature points with the global feature point map to obtain the three-dimensional feature points matched by the multiple two-dimensional feature points in the feature point map.

[0071] The global feature point map can be an offline-constructed global sparse feature point map, and the coverage of the global sparse feature point map includes at least the area where the autonomous robot may travel.

[0072] In implementation, during feature point extraction in step 103, a descriptor corresponding to each feature point can also be obtained, and each three-dimensional feature point in the feature point map also has a corresponding descriptor. Furthermore, for each two-dimensional feature point, the Euclidean distance between the descriptor of the two-dimensional feature point and the descriptor of each three-dimensional feature point can be calculated, and the three-dimensional feature point corresponding to the minimum Euclidean distance is determined as the three-dimensional feature point that matches the two-dimensional feature point.

[0073] In one possible implementation, when matching two-dimensional feature points with a global feature point map, the estimated pose information determined in step 103 can also be used for matching.

[0074] Specifically, a matching range can be determined in the global feature point map based on the estimated location information. Then, the Euclidean distance between the descriptors of the two-dimensional feature points and the descriptors of all three-dimensional feature points within the matching range is calculated. The three-dimensional feature point with the smallest corresponding Euclidean distance is selected as the matching three-dimensional feature point. This reduces the matching range, decreases the computational load, and thus improves matching efficiency.

[0075] Step 105: Obtain positioning information and update the positioning information according to the type of each two-dimensional feature point and the coordinates of the three-dimensional feature points matched by each two-dimensional feature point.

[0076] Among them, the types of two-dimensional feature points include VIO feature points and other feature points. Two-dimensional feature points with VIO feature point labels are VIO feature points, while two-dimensional feature points without VIO feature point labels are other feature points.

[0077] In implementation, for each two-dimensional feature point acquired in a scene image, it is determined whether the two-dimensional feature point is a VIO feature point. Specifically, the determination method can be as follows:

[0078] Determine whether a 2D feature point has a VIO feature point identifier. If a 2D feature point has a VIO feature point identifier, then the 2D feature point is determined to be a VIO feature point; if a 2D feature point does not have a VIO feature point identifier, then the 2D feature point is determined not to be a VIO feature point.

[0079] If the 2D feature point is a VIO feature point, then it is further determined whether the VIO state variables include the feature point information of the 2D feature point. The feature point information includes the coordinates of the 2D feature point, the coordinates of the matched 3D feature point, and the identifier of the 2D feature point. Specifically, the determination method can be as follows:

[0080] In step 103, during feature point extraction, the identifier of each two-dimensional feature point can also be obtained. Therefore, when determining whether the VIO state variable includes the feature point information of the two-dimensional feature point, it is only necessary to determine whether the VIO state variable includes the identifier of the two-dimensional feature point. If the VIO state variable includes the identifier of the two-dimensional feature point, it is determined that the VIO state variable includes the feature point information of the two-dimensional feature point; if the VIO state variable does not include the identifier of the two-dimensional feature point, it is determined that the VIO state variable does not include the feature point information of the two-dimensional feature point.

[0081] For a two-dimensional feature point in a scene image that is both a VIO feature point and included in the VIO state variables, update strategy one can be executed.

[0082] Update Strategy 1:

[0083] The coordinates of the matched 3D feature points, the initial measurement variance of the matched 3D feature points, and the estimated pose information of the corresponding scene image are input into an Extended Kalman Filter (EKF). Using Equation 1 as the observation equation, the previous-moment localization information of the device (autonomous robot) in the state variables is filtered and updated. Here, "two-dimensional feature points" refers to two-dimensional feature points that are both VIO feature points and included in the VIO's state variables within a scene image. "Previous-moment" refers to the autonomous robot's localization information before the filtering update, i.e., the localization information recorded in the VIO's state variables before the filtering update. The localization information can include the autonomous robot's position (coordinates), orientation (direction angle), and velocity in the world coordinate system. The initial measurement variance can be the reprojection error during feature point map construction and is preset.

[0084] ( G p f ) m = G T i * i p f Formula 1

[0085] In this system, the superscript indicates the coordinate system, G is the world coordinate system, and i is the body IMU coordinate system. M represents the measurement value, p f Represents the coordinates of three-dimensional feature points. In summary, ( G p f ) m This represents the coordinates of the measured three-dimensional feature point in the world coordinate system, which is the coordinates of the first three-dimensional feature point matched by the two-dimensional feature point mentioned above. G T i This indicates the transformation relationship between the machine's IMU coordinate system and the world coordinate system. i p f This represents the coordinates of a 3D feature point in the IMU coordinate system.

[0086] The EKF filtering process can be roughly divided into two stages. The first stage estimates the latest location information based on the previous location information, and the second stage filters and updates the estimated location information.

[0087] If the two-dimensional feature point is a VIO feature point, and the VIO state variable does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable, and the initial measurement variance of the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0088] If the 2D feature point is not a VIO feature point, then feature tracking is performed on the most recently acquired scene image to determine if a target 2D feature point matching the 2D feature point exists. The specific method is as follows:

[0089] After determining that the two-dimensional feature point is not a VIO feature point, it is determined whether there is a target two-dimensional feature point with a target identifier among the two-dimensional feature points extracted from the target scene image. The target scene image is the latest scene image received from the vision sensor, and the target identifier is the identifier of the two-dimensional feature point that was determined not to be a VIO feature point.

[0090] If a target 2D feature point with a target identifier exists among the 2D feature points extracted from the target scene image, then it is determined that the 2D feature point has a matching target 2D feature point in the latest acquired scene image; if no target 2D feature point with a target identifier exists among the 2D feature points extracted from the target scene image, then it is determined that the 2D feature point does not have a matching 2D feature point in the latest acquired scene image.

[0091] For a two-dimensional feature point that is not a VIO feature point in a scene image, but has a matching target two-dimensional feature point in the latest acquired scene image, update strategy two can be executed.

[0092] Update Strategy Two:

[0093] Based on the coordinates of the matched 3D feature points of these 2D feature points, the measurement noise of each 2D feature point is calculated. These 2D feature points refer to those that are not VIO feature points in a given scene image but have a matching target 2D feature point in the most recently acquired scene image. The coordinates of this 2D feature point in the camera coordinate system, the coordinates of the matched 3D feature point, the initial measurement variance corresponding to the matched 3D feature point, the estimated pose information of the scene image, and the measurement noise are input into the EKF. The local positioning information recorded in the VIO state variables is updated using the following Equation 2 as the observation equation.

[0094] z = ([u, v]) T ) m -h( c T i * i T G *( G p f )m ) Formula 2

[0095] Among them, ([u, v] T ) m This represents the coordinates of the measured two-dimensional feature points in the camera coordinate system, where T represents the transpose. h( c T i * i T G *( G p f ) m This means that the coordinates of the measured 3D feature points (the second 3D feature points mentioned above) are first transformed from the world coordinate system to the body IMU coordinate system, and then from the body IMU coordinate system to the camera coordinate system. i T G This indicates the transformation relationship from the world coordinate system to the machine's IMU coordinate system. c T i This represents the transformation relationship from the IMU coordinate system to the camera coordinate system. One possible way to calculate h() is as shown in Formula 3.

[0096]

[0097] Among them, c x c y c z This represents the three coordinate values ​​of a 3D feature point in the world coordinate system.

[0098] The calculation formula for the above-mentioned noise measurement is as follows: Formula 4.

[0099]

[0100] Where A is the detection noise of the two-dimensional feature points, which is a preset value and can be in the form of a 2*2 diagonal matrix; B is the global position noise of the feature point map, which is a preset value and can be in the form of a 3*3 matrix. h is the same as Formula 3 above.

[0101] If a two-dimensional feature point is not a VIO feature point, and there is no matching target two-dimensional feature point for that two-dimensional feature point in the latest acquired scene image, add the feature point information of that two-dimensional feature point to the VIO state variable, and obtain the initial measurement variance of the three-dimensional feature point that matches the two-dimensional feature point.

[0102] It should be noted here that, because VIO feature points and other feature points may be extracted simultaneously in a scene image, and the feature point information of the extracted VIO feature points may be included in the VIO state variables, and the other extracted feature points may also have matching target 2D feature points in the most recently acquired scene image, for this scene image, both update strategy one and update strategy two can be executed, and update strategy one can be executed first, followed by update strategy two. When update strategy two is executed to filter and update the localization information, the updated localization information is the localization information updated by update strategy one.

[0103] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0104] In the global positioning method provided in this application embodiment, the pose information estimated by VIO and the global feature point map are combined, and the autonomous robot is positioned by filtering and updating. The positioning process does not require GPS participation, which avoids the problem that GPS is affected by interference in scenarios such as tunnels and densely built-up areas when GPS and VIO are combined, resulting in poor global positioning effect.

[0105] This application also provides a method for global positioning, see [link to relevant documentation]. Figure 3 The processing flow of this method may include the following steps:

[0106] Step 301: Obtain scene images and IMU information.

[0107] Step 302: Extract multiple two-dimensional feature points from the scene image.

[0108] Step 303: Determine the estimated pose information based on multiple two-dimensional feature points and IMU information.

[0109] Step 304: Match multiple two-dimensional feature points with the global feature point map to obtain the three-dimensional feature points matched by the multiple two-dimensional feature points in the feature point map.

[0110] Step 305: Determine whether the two-dimensional feature point is a VIO feature point.

[0111] Step 306: If the two-dimensional feature point is a VIO feature point, then continue to determine whether the VIO state variables include the feature point information of the two-dimensional feature point.

[0112] Step 307: If the two-dimensional feature point is a VIO feature point and the VIO state variables include the feature point information of the two-dimensional feature point, then input the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the matched three-dimensional feature point, and the estimated pose information corresponding to the scene image to which it belongs into the EKF to update the positioning information of this device.

[0113] Step 308: If the two-dimensional feature point is a VIO feature point and the feature point information of the two-dimensional feature point is not included in the state variables of the VIO, then add the feature point information of the two-dimensional feature point to the state variables of the VIO and obtain the initial measurement variance of the three-dimensional feature point matched by the two-dimensional feature point.

[0114] Step 309: If the two-dimensional feature point is not a VIO feature point, perform feature tracking on the two-dimensional feature point in the latest acquired scene image to determine whether there is a matching target two-dimensional feature point in the latest acquired scene image.

[0115] Step 310: If the two-dimensional feature point is not a VIO feature point, and a matching target two-dimensional feature point exists in the latest acquired scene image, then calculate the measurement noise based on the coordinates of the matching second three-dimensional feature point. Input the coordinates of the two-dimensional feature point in the camera coordinate system, the coordinates of the second three-dimensional feature point, the initial measurement variance of the second three-dimensional feature point, the estimated pose information corresponding to the scene image, and the measurement noise into the EKF to update the positioning information of this device.

[0116] Step 311: If the two-dimensional feature point is not a VIO feature point and there is no matching target two-dimensional feature point in the latest acquired scene image, add the feature point information of the two-dimensional feature point to the state variable and obtain the initial measurement variance of the three-dimensional feature point matching the two-dimensional feature point.

[0117] It should be noted that the specific processing of step 301 is the same as that of step 101, the specific processing of step 302 is the same as that of step 202, the specific processing of step 303 is the same as that of step 302, the specific processing of step 304 is the same as that of step 104, and the specific processing of steps 305 to 311 is the same as that of step 205, and will not be repeated here.

[0118] In the global positioning method provided in this application embodiment, the pose information estimated by VIO and the global feature point map are combined, and the autonomous robot is positioned by filtering and updating. The positioning process does not require GPS participation, which avoids the problem that GPS is affected by interference in scenarios such as tunnels and densely built-up areas when GPS and VIO are combined, resulting in poor global positioning effect.

[0119] Based on the same technical concept, embodiments of this application also provide a global positioning device, which can be applied to autonomous robots. See also Figure 4 The device includes an acquisition module 410, an extraction module 420, a determination module 430, a matching module 440, and an update module 450. Wherein:

[0120] The acquisition module 410 is used to acquire positioning information, scene images, and IMU information;

[0121] Extraction module 420 is used to extract multiple two-dimensional feature points of the scene image;

[0122] The determination module 430 is used to determine the estimated pose information based on the plurality of two-dimensional feature points and the IMU information;

[0123] The matching module 440 is used to match the plurality of two-dimensional feature points with a global feature point map to obtain three-dimensional feature points that are matched by the plurality of two-dimensional feature points in the global feature point map.

[0124] The update module 450 is used to filter and update the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information.

[0125] In one possible implementation, the update module 450 is configured to:

[0126] If the two-dimensional feature point is a VIO feature point, and the state variables of the VIO include the feature point information of the two-dimensional feature point, then the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point, and the estimated pose information are input into the EKF to filter and update the positioning information of this device. The feature point information of the two-dimensional feature point includes the coordinates of the two-dimensional feature point, the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, and the identifier of the two-dimensional feature point.

[0127] In one possible implementation, the update module 450 is further configured to:

[0128] If the two-dimensional feature point is a VIO feature point, and the state variable of the VIO does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

[0129] In one possible implementation, the update module 450 is configured to:

[0130] If the two-dimensional feature point is not a VIO feature point, feature tracking is performed on the two-dimensional feature point in the latest acquired scene image. If there is a target two-dimensional feature point in the latest acquired scene image that matches the two-dimensional feature point, the measurement noise is calculated based on the coordinates of the second three-dimensional feature point that matches the two-dimensional feature point.

[0131] The coordinates of the two-dimensional feature points, the coordinates of the three-dimensional feature points matched by the two-dimensional feature points, the initial measurement variance corresponding to the three-dimensional feature points matched by the two-dimensional feature points, the estimated pose information, and the measurement noise are input into the EKF to update the positioning information of this device.

[0132] In one possible implementation, the update module 450 is configured to:

[0133] If the two-dimensional feature point is not a VIO feature point, and the target two-dimensional feature point does not exist in the latest acquired scene image, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained. In one possible implementation, the extraction module 420 is used to:

[0134] Extract a first number of two-dimensional feature points from the scene image, and add VIO feature identifiers to the first number of two-dimensional feature points as the VIO feature points;

[0135] Extract a second number of two-dimensional feature points from the scene image.

[0136] In the global positioning method provided in this application embodiment, the pose information estimated by VIO and the global feature point map are combined, and the autonomous robot is positioned by filtering and updating. The positioning process does not require GPS participation, which avoids the problem that GPS is affected by interference in scenarios such as tunnels and densely built-up areas when GPS and VIO are combined, resulting in poor global positioning effect.

[0137] It should be noted that the global positioning device provided in the above embodiments is only illustrated by the division of the functional modules described above. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the autonomous robot can be divided into different functional modules to complete all or part of the functions described above. In addition, the global positioning device and the global positioning method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0138] Figure 5 A structural block diagram of an autonomous robot 500 provided in an exemplary embodiment of this application is shown. The autonomous robot 500 may be a drone, an unmanned delivery vehicle, or the like.

[0139] Typically, an autonomous robot 500 includes a processor 501 and a memory 502.

[0140] Processor 501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 501 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0141] Memory 502 may include one or more computer-readable storage media, which may be non-transitory. Memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 502 is used to store at least one instruction, which is executed by processor 501 to implement the global positioning method provided in the method embodiments of this application.

[0142] In some embodiments, the autonomous robot 500 may also optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, memory 502, and peripheral device interface 503 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 503 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, a positioning assembly 508, and a power supply 509.

[0143] Peripheral device interface 503 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 501 and memory 502. In some embodiments, processor 501, memory 502 and peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 501, memory 502 and peripheral device interface 503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0144] The radio frequency (RF) circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 504 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0145] Display screen 505 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 505 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 501 for processing. In this case, display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 505, disposed on the front panel of the autonomous robot 500; in other embodiments, there may be at least two display screens, disposed on different surfaces of the autonomous robot 500 or in a folded design; in other embodiments, display screen 505 may be a flexible display screen disposed on the outer surface of the autonomous robot 500. Furthermore, display screen 505 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 505 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0146] The camera assembly 506 is used to acquire images or videos. Optionally, the camera assembly 506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0147] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 501 for processing, or to the radio frequency circuit 504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations on the autonomous robot 500. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0148] The positioning component 508 is used to determine the current geographical location of the autonomous robot 500 in order to enable navigation or LBS (Location Based Service). The positioning component 508 can be a positioning component based on GPS (Global Positioning System), BeiDou system, or Galileo system.

[0149] Power supply 509 is used to power the various components in autonomous robot 500. Power supply 509 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0150] In some embodiments, the autonomous robot 500 further includes one or more sensors 510. The one or more sensors 510 include, but are not limited to, an accelerometer 511, a gyroscope 512, and a proximity sensor 513.

[0151] Accelerometer 511 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by the autonomous robot 500. For example, accelerometer 511 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 501 can control display screen 505 to display the user interface in either a horizontal or vertical view based on the gravitational acceleration signal acquired by accelerometer 511. Accelerometer 511 can also be used for games or for acquiring user motion data.

[0152] The gyroscope sensor 512 can detect the body orientation and rotation angle of the autonomous robot 500. The gyroscope sensor 512 can work in conjunction with the accelerometer sensor 511 to collect the user's 3D movements of the autonomous robot 500. Based on the data collected by the gyroscope sensor 512, the processor 501 can perform the following functions: motion sensing, image stabilization during shooting, and inertial navigation.

[0153] The proximity sensor 513, also known as a distance sensor, is typically mounted on the front panel of the autonomous robot 500. The proximity sensor 513 is used to detect the distance between external objects and the front of the autonomous robot 500.

[0154] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the autonomous robot 500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0155] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in an autonomous robot to complete the global localization method described above. This computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.

[0156] It should be noted that all information (including but not limited to user device information, user personal information, and collected scene images), data (including but not limited to data used for analysis, stored data, and displayed data), and signals (including but not limited to signals transmitted between the user terminal and other devices) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the scene images involved in this application were all obtained under full authorization.

[0157] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0158] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for global positioning, characterized in that, The method includes: Obtain location information; Acquire scene images and inertial measurement unit (IMU) information; Extract multiple two-dimensional feature points from the scene image; Based on the multiple two-dimensional feature points and the IMU information, the estimated pose information is determined; The multiple two-dimensional feature points are matched with the global feature point map to obtain the three-dimensional feature points matched by the multiple two-dimensional feature points in the global feature point map; The positioning information is filtered and updated based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information. The step of filtering and updating the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched with each two-dimensional feature point, and the estimated pose information includes: If the two-dimensional feature point is not a VIO feature point, feature tracking is performed on the two-dimensional feature point in the latest acquired scene image. If there is a target two-dimensional feature point in the latest acquired scene image that matches the two-dimensional feature point, the measurement noise is calculated based on the coordinates of the three-dimensional feature point that matches the two-dimensional feature point. The coordinates of the two-dimensional feature points, the coordinates of the three-dimensional feature points matched by the two-dimensional feature points, the initial measurement variance corresponding to the three-dimensional feature points matched by the two-dimensional feature points, the estimated pose information, and the measurement noise are input into the EKF to update the positioning information.

2. The method according to claim 1, characterized in that, The step of filtering and updating the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information includes: If the two-dimensional feature point is a visual inertial odometry (VIO) feature point, and the VIO state variables include the feature point information of the two-dimensional feature point, then the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point, and the estimated pose information are input into an extended Kalman filter (EKF) to filter and update the positioning information. The feature point information of the two-dimensional feature point includes the coordinates of the two-dimensional feature point, the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, and the identifier of the two-dimensional feature point.

3. The method according to claim 2, characterized in that, The method further includes: If the two-dimensional feature point is a VIO feature point, and the state variable of the VIO does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

4. The method according to claim 1, characterized in that, The method further includes: If the two-dimensional feature point is not a VIO feature point, and the target two-dimensional feature point does not exist in the latest acquired scene image, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

5. The method according to any one of claims 2-4, characterized in that, The extraction of multiple two-dimensional feature points from the scene image includes: Extract a first number of two-dimensional feature points from the scene image, and add VIO feature identifiers to the first number of two-dimensional feature points as the VIO feature points; Extract a second number of two-dimensional feature points from the scene image.

6. A global positioning device, characterized in that, The device includes: The acquisition module is used to acquire location information, scene images, and IMU information; An extraction module is used to extract multiple two-dimensional feature points from the scene image; The determination module is used to determine the estimated pose information based on the plurality of two-dimensional feature points and the IMU information; The matching module is used to match the plurality of two-dimensional feature points with a global feature point map to obtain three-dimensional feature points that are matched by the plurality of two-dimensional feature points in the global feature point map. The update module is used to filter and update the positioning information based on the type of each two-dimensional feature point, the coordinates of the three-dimensional feature points matched by each two-dimensional feature point, and the estimated pose information. The update module is used for: If the two-dimensional feature point is not a VIO feature point, feature tracking is performed on the two-dimensional feature point in the latest acquired scene image. If there is a target two-dimensional feature point in the latest acquired scene image that matches the two-dimensional feature point, the measurement noise is calculated based on the coordinates of the second three-dimensional feature point that matches the two-dimensional feature point. The coordinates of the two-dimensional feature points, the coordinates of the three-dimensional feature points matched by the two-dimensional feature points, the initial measurement variance corresponding to the three-dimensional feature points matched by the two-dimensional feature points, the estimated pose information, and the measurement noise are input into the EKF to update the positioning information of this device.

7. The apparatus according to claim 6, characterized in that, The update module is used for: If the two-dimensional feature point is a VIO feature point, and the state variables of the VIO include the feature point information of the two-dimensional feature point, then the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point, and the estimated pose information are input into the EKF to filter and update the positioning information of this device. The feature point information of the two-dimensional feature point includes the coordinates of the two-dimensional feature point, the coordinates of the three-dimensional feature point matched by the two-dimensional feature point, and the identifier of the two-dimensional feature point.

8. The apparatus according to claim 7, characterized in that, The update module is also used for: If the two-dimensional feature point is a VIO feature point, and the state variable of the VIO does not include the feature point information of the two-dimensional feature point, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

9. The apparatus according to claim 6, characterized in that, The update module is used for: If the two-dimensional feature point is not a VIO feature point, and the target two-dimensional feature point does not exist in the latest acquired scene image, then the feature point information of the two-dimensional feature point is added to the state variable of the VIO, and the initial measurement variance corresponding to the three-dimensional feature point matched by the two-dimensional feature point is obtained.

10. The apparatus according to any one of claims 7-9, characterized in that, The extraction module is used for: Extract a first number of two-dimensional feature points from the scene image, and add VIO feature identifiers to the first number of two-dimensional feature points as the VIO feature points; Extract a second number of two-dimensional feature points from the scene image.

11. An autonomous robot, characterized in that, The autonomous robot includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to perform the operation of the global positioning method as described in any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to perform the operation of the global positioning method as described in any one of claims 1 to 5.