Positioning method and device, electronic equipment and storage medium

By generating temporary maps in visual positioning technology, the problem of reducing positioning accuracy caused by scene changes is solved, and the positioning accuracy of the target equipment is improved.

CN120219481APending Publication Date: 2025-06-27HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311800435.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In visual positioning technology, scene changes cause the correlation state of the image collected by the target device to be reduced with the global map, thereby reducing the accuracy of positioning.

Method used

When the latest positioning state is the first preset state, a temporary map of the target scene is generated, and the target position of the target image is determined based on the first pose corresponding to the first image of the previous frame of the target image, the global map, the temporary map and the target image.

Benefits of technology

By generating a temporary map, the correlation status between the target image and the temporary map is improved, and the accuracy of positioning the target device is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219481A_ABST
    Figure CN120219481A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a positioning method and device, electronic equipment and a storage medium, and relates to the technical field of visual localization, and the method comprises the steps: obtaining a global map of a target scene where target equipment is located, a target image of the target scene collected by a camera, and a latest positioning state for positioning the target equipment; when the latest positioning state is a first preset state, determining the target image and an adjacent image collected before the target image is collected as a second image, and generating a temporary map of the target scene based on the second image; the latest positioning state represents a target association state of an image acquired by the target equipment and a global map of the target scene; determining a target pose corresponding to the target image based on the target image, the global map, the temporary map and a first pose corresponding to a previous frame of first image of the target image; the pose corresponding to one image is the pose of the target equipment when the image is acquired. Therefore, the accuracy of positioning the target equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual positioning technology, and in particular, to a positioning method, device, electronic device, and storage medium. Background Art

[0002] In visual positioning technology, an initial image of a target scene is pre-collected by a camera of a collection device, and a global map of the target scene is constructed in combination with other sensor information of the collection device. The global map includes: the pose corresponding to the key-frame image in the initial image, the positions of map points in the target scene, and the feature description information of each feature point in the initial image. The pose corresponding to the key-frame image is the pose of the collection device when the key-frame image is collected.

[0003] When positioning a target device moving in a target scene, a target image of the target scene collected by the target device in real time is matched with the global map, and the 6DOF (Degree Of Freedom) pose of the target device is determined based on the matching result.

[0004] However, for some situations where the scene changes, for example, the illumination and scene structure of the target scene when constructing the global map are different from those when positioning the target device, it will lead to a decrease in the association state between the image collected by the target device and the global map, and further lead to a decrease in the accuracy of matching the target image with the global map, thereby reducing the accuracy of positioning the target device. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a positioning method, device, electronic device, and storage medium to improve the accuracy of positioning a target device. The specific technical solutions are as follows:

[0006] In a first aspect, to achieve the above object, the embodiments of this application provide a positioning method, which is applied to a target device equipped with a camera, and the method includes:

[0007] Obtain the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning status of the target device; when the latest positioning status is the first preset status, determine the target image and the adjacent image collected before the target image as the second image, and generate a temporary map of the target scene based on the second image; wherein, the latest positioning status represents the target association status between the image collected by the target device and the global map; determine the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image; wherein, the pose corresponding to an image is: the pose of the target device when the image is collected.

[0008] Optionally, the determining the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image includes:

[0009] Determine the prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image and the first relative pose of the target device from the time when the first image is collected to the time when the target image is collected; adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when the first pose is determined; the first associated map point is the map point in the target scene associated with the feature point in the first image; determine the first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine the map point associated with the feature point in the first key frame image that observes the first reference map point in the global map from the global map points in the global map, and determine the map point associated with the feature point in the second key frame image that observes the first reference map point in the temporary map from the temporary map points in the temporary map as the first local map points; adjust the candidate pose corresponding to the target image based on the association between the first local map points and the feature points in the target image to obtain the target pose corresponding to the target image.

[0010] Optionally, the adjusting the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image to obtain the candidate pose corresponding to the target image includes:

[0011] Based on the first association relationship between the first associated map point in the target scene and the feature point in the first image, and the prior pose corresponding to the target image, determine the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image; adjust the prior pose corresponding to the target image, and determine the candidate pose corresponding to the target image when the first reprojection error between the first projection position and the first candidate matching feature point satisfies a preset convergence condition; wherein, the first reprojection error is the distance between the first projection position and the first candidate matching feature point.

[0012] Optionally, the step of determining the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image based on the first association relationship between the first associated map point in the target scene and the feature point in the first image, and the prior pose corresponding to the target image includes:

[0013] Based on the first association relationship between the first associated map point in the target scene and the feature point in the first image, and the prior pose corresponding to the target image, project the first associated map point to the pixel position in the target image to obtain the first projection position of the first associated map point in the target image; based on the feature description information of the first projection position and the feature description information of each feature point within the preset neighborhood range of the first projection position in the target image, determine the first candidate matching feature point of the target map point.

[0014] Optionally, the step of adjusting the candidate pose corresponding to the target image based on the association between the first local map point and the feature point in the target image to obtain the target pose corresponding to the target image includes:

[0015] Based on the first association relationship and the candidate pose corresponding to the target image, determine the second projection position of the first local map point in the target image and the second candidate matching feature point of the first local map point in the target image; adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when the second reprojection error between the second projection position and the second candidate matching feature point satisfies a preset convergence condition.

[0016] Optionally, the step of determining the first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image includes:

[0017] Based on the first association relationship and the candidate pose corresponding to the target image, determine the third projection position of the first associated map point in the target image, and the third candidate matching feature point of the first associated map point in the target image; from the first associated map points, determine the map points for which the third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold, to obtain first reference map points.

[0018] Optionally, after adjusting the candidate pose corresponding to the target image based on associating the first local map points with the feature points in the target image to obtain the target pose corresponding to the target image, the method further includes:

[0019] Based on the first association relationship and the target pose corresponding to the target image, determine second reference map points from the global map points among the first associated map points; if the number of the second reference map points is less than a first preset number, update the latest positioning state to a first preset state; if the number of the second reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state.

[0020] After determining, from the first associated map points, the map points for which the third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold to obtain first reference map points, the method further includes:

[0021] If the number of the first reference map points is less than a second preset number, update the latest positioning state to a third preset state; wherein, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0022] Optionally, the first relative pose is determined based on a target sensor in the target device; the first relative pose includes: the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis of the world coordinate system when the target device moves from acquiring the first image to acquiring the target image.

[0023] Optionally, determining that the target image and an adjacent image acquired before acquiring the target image is a second image, and generating a temporary map of the target scene based on the second image, includes:

[0024] Obtain a third preset number of adjacent images from the images acquired by the camera before acquiring the target image, and determine the target image and the adjacent images as second images; for each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determine the position of the temporary map point corresponding to the feature points in this second image in the target scene, and the second pose corresponding to this second image; determine the second key frame images in each second image, and obtain the second pose corresponding to the second key frame image, the feature description information of the feature points in the second key frame image, and the position of the temporary map point, to obtain a temporary map of the target scene.

[0025] Optionally, the step of, for each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determining the position of the temporary map point corresponding to the feature points in this second image in the target scene, and the second pose corresponding to this second image, includes:

[0026] Based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, calculate the first matching relationship between the feature points in this second image and the previous frame of the second image; based on the pixel positions of two matching feature points in this second image and the previous frame of the second image, calculate the distance between the same temporary map point corresponding to the two matching feature points in the target scene and the target device, and determine the position of the temporary map point based on this distance; based on the first matching relationship, the second pose corresponding to the previous frame of the second image, and the second relative pose of the target device from the time of acquiring the previous frame of the second image to the time of acquiring this second image, obtain the prior pose corresponding to this second image; based on the second association relationship between the temporary map point and the feature points in the previous frame of the second image, and the prior pose corresponding to this second image, determine the fourth projection position of the temporary map point in this second image, and the fourth candidate matching feature point of the temporary map point in this second image; respectively adjust the position of the temporary map point and the prior pose corresponding to this second image, and determine the position of the temporary map point and the second pose corresponding to this second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point satisfies a preset convergence condition.

[0027] Optionally, after obtaining the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning state of positioning the target device, the method further includes:

[0028] When the latest positioning state is the second preset state, determine the target pose corresponding to the target image based on the first pose corresponding to the previous frame of the first image of the target image, the target image, and the global map.

[0029] Optionally, the determining the target pose corresponding to the target image based on the first pose corresponding to the previous frame of the first image of the target image, the target image, and the global map includes:

[0030] Determine the prior pose corresponding to the target image based on the first pose corresponding to the previous frame of the first image of the target image and the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image; adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map points of the target scene and the feature points in the first image to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map points are the map points in the target scene associated with the feature points in the first image; determine the third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine the map points associated with the feature points in the first key frame image that observes the third reference map point in the global map from the global map points in the global map as the second local map points; adjust the candidate pose corresponding to the target image based on the association between the second local map points and the feature points in the target image to obtain the target pose corresponding to the target image.

[0031] Optionally, after adjusting the candidate pose corresponding to the target image based on the association between the second local map points and the feature points in the target image to obtain the target pose corresponding to the target image, the method further includes:

[0032] Determine the fourth reference map point from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; if the number of the fourth reference map points is less than the first preset number, update the latest positioning state to the first preset state; if the number of the fourth reference map points is greater than the first preset number, update the latest positioning state to the second preset state; wherein, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state;

[0033] After determining the third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image, the method further includes:

[0034] If the number of the third reference map points is less than a second preset number, update the latest positioning state to a third preset state; wherein, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0035] Optionally, the first preset number when updating the latest positioning state to the second preset state when the latest positioning state is the first preset state is greater than the first preset number when updating the latest positioning state to the first preset state when the latest positioning state is the second preset state.

[0036] Optionally, after obtaining the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning state of positioning the target device, the method further includes:

[0037] When the latest positioning state is the third preset state, determine the target pose corresponding to the target image based on the target image and the global map.

[0038] Optionally, the determining the target pose corresponding to the target image based on the target image and the global map includes:

[0039] Based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map, determine a candidate recall image of the target image from the first key frame image; based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recall image, determine a second matching relationship between the feature points in the target image and the feature points in the candidate recall image; based on the second matching relationship and the third association relationship between the feature points in the candidate recall image and the global map points in the global map, obtain a fourth association relationship between the feature points in the target image and the global map points in the global map; perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image.

[0040] Optionally, the performing pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image includes:

[0041] Perform pose calculation based on the fourth association relationship to obtain the candidate pose corresponding to the target image; determine the prior pose corresponding to the target image based on the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image; perform consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image; when the candidate pose corresponding to the target image passes the verification, determine the target pose as the candidate pose corresponding to the target image.

[0042] Optionally, after performing consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image, the method further includes:

[0043] When the candidate pose corresponding to the target image passes the verification, update the latest positioning state to a second preset state; when the candidate pose corresponding to the target image fails to pass the verification, update the latest positioning state to a third preset state; wherein, the target association state corresponding to the third preset state is lower than the target association state corresponding to the second preset state.

[0044] Optionally, after generating the temporary map of the target scene based on the second image, the method further includes:

[0045] Delete the temporary map after a preset duration from generating the temporary map.

[0046] In a second aspect, to achieve the above object, an embodiment of the present application provides a positioning device, which is applied to a target device, and a camera is installed on the target device. The device includes:

[0047] A data acquisition module, configured to acquire the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning state for positioning the target device;

[0048] A temporary map generation module, configured to determine the target image and an adjacent image acquired before collecting the target image as the second image when the latest positioning state is a first preset state, and generate a temporary map of the target scene based on the second image; wherein, the latest positioning state represents the target association state between the image acquired by the target device and the global map;

[0049] A first positioning module, configured to determine the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the previous frame of the first image of the target image; wherein, the pose corresponding to an image is: the pose of the target device when collecting the image.

[0050] Optionally, the first positioning module is specifically configured to determine a prior pose corresponding to the target image based on the first pose corresponding to the previous frame of the first image of the target image and the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image; adjust the prior pose corresponding to the target image based on a first association relationship between a first associated map point of the target scene and a feature point in the first image to obtain a candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map point is a map point in the target scene associated with the feature point in the first image; determine a first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine, from the global map points in the global map, map points associated with the feature points in the first key frame image that observes the first reference map point in the global map, and determine, from the temporary map points in the temporary map, map points associated with the feature points in the second key frame image that observes the first reference map point in the temporary map as first local map points; and adjust the candidate pose corresponding to the target image based on associating the first local map points with the feature points in the target image to obtain the target pose corresponding to the target image.

[0051] Optionally, the first positioning module is specifically configured to determine a first projection position of the first associated map point in the target image and a first candidate matching feature point of the first associated map point in the target image based on a first association relationship between a first associated map point of the target scene and a feature point in the first image and the prior pose corresponding to the target image; adjust the prior pose corresponding to the target image to determine a candidate pose corresponding to the target image when a first reprojection error preset convergence condition between the first projection position and the first candidate matching feature point is satisfied; wherein, the first reprojection error is the distance between the first projection position and the first candidate matching feature point.

[0052] Optionally, the first positioning module is specifically configured to project the first associated map point to a pixel position in the target image based on a first association relationship between a first associated map point of the target scene and a feature point in the first image and the prior pose corresponding to the target image to obtain a first projection position of the first associated map point in the target image; and determine a first candidate matching feature point of the target map point based on the feature description information of the first projection position and the feature description information of each feature point within a preset neighborhood range of the first projection position in the target image.

[0053] Optionally, the first positioning module is specifically configured to determine a second projection position of the first local map point in the target image and a second candidate matching feature point of the first local map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when the second reprojection error between the second projection position and the second candidate matching feature point satisfies a preset convergence condition.

[0054] Optionally, the first positioning module is specifically configured to determine a third projection position of the first associated map point in the target image and a third candidate matching feature point of the first associated map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; determine, from the first associated map points, map points for which the third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold, to obtain first reference map points.

[0055] Optionally, the apparatus further comprises:

[0056] A first positioning state update module, configured to, after the first positioning module performs an adjustment on the candidate pose corresponding to the target image based on an association between the first local map point and the feature points in the target image to obtain the target pose corresponding to the target image, determine second reference map points from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; if the number of the second reference map points is less than a first preset number, update the latest positioning state to a first preset state; if the number of the second reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state.

[0057] The apparatus further comprises:

[0058] A second positioning state update module, configured to, after the first positioning module determines, from the first associated map points, map points for which the third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold to obtain first reference map points, perform an update of the latest positioning state to a third preset state if the number of the first reference map points is less than a second preset number; wherein the second preset number is less than the first preset number; and the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0059] Optionally, the first relative pose is determined based on a target sensor in the target device; the first relative pose includes: when the target device moves from capturing the first image to capturing the target image, the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis in the world coordinate system.

[0060] Optionally, the temporary map generation module is specifically configured to: obtain a third preset number of adjacent images from the images captured by the camera before capturing the target image, and determine the target image and the adjacent images as second images; for each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determine the positions of the feature points in this second image corresponding to the temporary map points in the target scene, and the second pose corresponding to this second image; determine the second key frame images in each second image, and obtain the second poses corresponding to the second key frame images, the feature description information of the feature points in the second key frame images, and the positions of the temporary map points, to obtain the temporary map of the target scene.

[0061] Optionally, the temporary map generation module is specifically configured to: calculate a first matching relationship between the feature points in this second image and the feature points in the previous frame of the second image based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image; calculate the distance between the two matched feature points in the target scene corresponding to the same temporary map point and the target device based on the pixel positions of the two matched feature points in this second image and the previous frame of the second image, and determine the position of the temporary map point based on this distance; obtain the prior pose corresponding to this second image based on the first matching relationship, the second pose corresponding to the previous frame of the second image, and the second relative pose of the target device from capturing the previous frame of the second image to capturing this second image; determine the fourth projection position of the temporary map point in this second image and the fourth candidate matching feature point of the temporary map point in this second image based on the second association relationship between the temporary map point and the feature points in the previous frame of the second image and the prior pose corresponding to this second image; respectively adjust the position of the temporary map point and the prior pose corresponding to this second image, and determine the position of the temporary map point and the second pose corresponding to this second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point satisfies a preset convergence condition.

[0062] Optionally, the apparatus further includes:

[0063] A second positioning module, configured to, after the data acquisition module executes to acquire a global map of the target scenario where the target device is located, a target image of the target scenario acquired by the camera, and the latest positioning status of positioning the target device, execute to, when the latest positioning status is a second preset status, determine a target pose corresponding to the target image based on a first pose corresponding to a first image of the previous frame of the target image, the target image, and the global map.

[0064] Optionally, the second positioning module is specifically configured to determine a prior pose corresponding to the target image based on a first pose corresponding to a first image of the previous frame of the target image and a first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image; adjust the prior pose corresponding to the target image based on a first association relationship between a first associated map point of the target scenario and a feature point in the first image to obtain a candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map point is a map point in the target scenario associated with the feature point in the first image; determine a third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine, from global map points in the global map, a map point associated with a feature point in a first key frame image that observes the third reference map point in the global map as a second local map point; and adjust the candidate pose corresponding to the target image based on associating the second local map point with the feature point in the target image to obtain the target pose corresponding to the target image.

[0065] Optionally, the device further includes:

[0066] A third positioning status update module, configured to, after the second positioning module executes to adjust the candidate pose corresponding to the target image based on associating the second local map point with the feature point in the target image to obtain the target pose corresponding to the target image, execute to determine a fourth reference map point from global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; if the number of the fourth reference map points is less than a first preset number, update the latest positioning status to a first preset status; if the number of the fourth reference map points is greater than the first preset number, update the latest positioning status to a second preset status; wherein, the target association status corresponding to the first preset status is lower than the target association status corresponding to the second preset status;

[0067] The device further includes:

[0068] A fourth positioning status update module, configured to, after the second positioning module determines a third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image, update the latest positioning status to a third preset status if the number of the third reference map points is less than a second preset number; wherein, the second preset number is less than the first preset number; and the target association status corresponding to the third preset status is lower than the target association status corresponding to the first preset status.

[0069] Optionally, when updating the latest positioning status to the second preset status in the case where the latest positioning status is the first preset status, the first preset number is greater than the first preset number when updating the latest positioning status to the first preset status in the case where the latest positioning status is the second preset status.

[0070] Optionally, the apparatus further includes:

[0071] A third positioning module, configured to, after the data acquisition module acquires the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning status of positioning the target device, determine the target pose corresponding to the target image based on the target image and the global map when the latest positioning status is the third preset status.

[0072] Optionally, the third positioning module is specifically configured to determine a candidate recall image of the target image from the first key frame image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map; determine a second matching relationship between the feature points in the target image and the feature points in the candidate recall image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recall image; obtain a fourth association relationship between the feature points in the target image and the global map points in the global map based on the second matching relationship and a third association relationship between the feature points in the candidate recall image and the global map points in the global map; and perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image.

[0073] Optionally, the third positioning module is specifically configured to perform pose calculation based on the fourth association relationship to obtain a candidate pose corresponding to the target image; determine a prior pose corresponding to the target image based on a first relative pose of the target device from the time of collecting the first image to the time of collecting the target image; perform consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image; and when the candidate pose corresponding to the target image passes the verification, determine the target pose as the candidate pose corresponding to the target image.

[0074] Optionally, the apparatus further includes:

[0075] A fifth positioning status update module, configured to, after the third positioning module performs consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image, update the latest positioning status to a second preset status when the candidate pose corresponding to the target image passes the verification; and update the latest positioning status to a third preset status when the candidate pose corresponding to the target image fails to pass the verification; where a target association status corresponding to the third preset status is lower than a target association status corresponding to the second preset status.

[0076] Optionally, the apparatus further includes:

[0077] A temporary map deletion module, configured to, after the temporary map generation module generates a temporary map of the target scene based on the second image, delete the temporary map after a preset duration from the generation of the temporary map.

[0078] An embodiment of the present application further provides an electronic device, including:

[0079] A memory for storing a computer program;

[0080] A processor, configured to implement the positioning method described in any one of the above when executing the program stored in the memory.

[0081] An embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program, when executed by a processor, implements the positioning method described in any one of the above.

[0082] An embodiment of the present application further provides a computer program product including instructions, which, when running on a computer, cause the computer to execute the positioning method described in any one of the above.

[0083] Advantages of the embodiments of the present application:

[0084] A positioning method, device, electronic device, and storage medium provided by an embodiment of the present application are applied to a target device, and a camera is installed on the target device. The method includes: obtaining a global map of a target scene where the target device is located, a target image of the target scene collected by the camera, and the latest positioning status of positioning the target device; when the latest positioning status is a first preset status, determining the target image and an adjacent image collected before collecting the target image as second images, and generating a temporary map of the target scene based on the second images; the latest positioning status represents the target association status between the image collected by the target device and the global map of the target scene; determining the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image.

[0085] Based on the above processing, when the latest positioning status is the first preset status, it indicates that the target association status between the image collected by the target device and the global map of the target scene is low, that is, the accuracy of positioning the target device is reduced. Then, the target image and the adjacent image collected before collecting the target image are determined as second images, and a temporary map of the target scene is generated based on the second images. Since the target scene changes little when collecting the target image and the adjacent image, when positioning the target device based on the temporary map, the association status between the target image and the temporary map is high, that is, the accuracy of matching the target image and the temporary map is high. Furthermore, the accuracy of positioning the target device can be improved.

[0086] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. Description of the Drawings

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0088] Figure 1 It is a flowchart of the first positioning method provided by an embodiment of the present application;

[0089] Figure 2 It is a flowchart of the second positioning method provided by an embodiment of the present application;

[0090] Figure 3 It is a schematic diagram of the principle of generating a temporary map provided by an embodiment of the present application;

[0091] Figure 4 It is a flowchart of generating a temporary map provided by an embodiment of the present application;

[0092] Figure 5 Flow chart of the third positioning method provided by the embodiment of the present application;

[0093] Figure 6 Schematic diagram of the distribution of temporary map points and global map points provided by the embodiment of the present application;

[0094] Figure 7 Flow chart of positioning a target device based on a global map provided by the embodiment of the present application;

[0095] Figure 8 Comparison chart of thresholds for switching different positioning states provided by the embodiment of the present application;

[0096] Figure 9 Flow chart of switching different positioning states provided by the embodiment of the present application;

[0097] Figure 10 Comparison chart of different positioning modes provided by the embodiment of the present application;

[0098] Figure 11 Flow chart of the fourth positioning method provided by the embodiment of the present application;

[0099] Figure 12 Structural diagram of a positioning device provided by the embodiment of the present application;

[0100] Figure 13 Structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0101] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the protection scope of the present application.

[0102] In the related art, in the case where some scenarios change, for example, when constructing a global map, the lighting and scene structure of the target scene are different from those when positioning the target device, which will reduce the correlation state between the images collected by the target device and the global map, and further reduce the accuracy of matching the target image with the global map, thereby reducing the accuracy of positioning the target device.

[0103] To solve the above problems, an embodiment of the present application provides a positioning method, which is applied to a target device equipped with a camera. The target device can be a vehicle, a mobile robot, or the like. When the target device moves in a target scene, it acquires the global map of the target scene where it is located, the target image of the target scene collected by the camera, and the latest positioning state of the target device; when the latest positioning state is a first preset state, it generates a temporary map of the target scene, and determines the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image, which can improve the accuracy of positioning the target device.

[0104] See Figure 1 , Figure 1 which is a flowchart of a positioning method provided by an embodiment of the present application. The method is applied to a target device equipped with a camera, and the method includes the following steps:

[0105] S101: Acquire the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning state for positioning the target device.

[0106] S102: When the latest positioning state is a first preset state, determine the target image and the adjacent image collected before collecting the target image as the second image, and generate a temporary map of the target scene based on the second image.

[0107] Among them, the latest positioning state represents the target association state between the image collected by the target device and the global map of the target scene.

[0108] S103: Determine the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image.

[0109] Among them, the pose corresponding to an image is: the pose of the target device when the image is collected.

[0110] Based on the positioning method provided by the embodiment of the present application, when the latest positioning state is a first preset state, it indicates that the target association state between the image collected by the target device and the global map of the target scene is low, that is, the accuracy of positioning the target device is reduced. Then, determine the target image and the adjacent image collected before collecting the target image as the second image, and generate a temporary map of the target scene based on the second image. Since the target scene changes little when collecting the target image and the adjacent image, when positioning the target device based on the temporary map, the association state between the target image and the temporary map is high, that is, the accuracy of matching the target image and the temporary map is high. Furthermore, the accuracy of positioning the target device can be improved.

[0111] For step S101, the target scene is the scene where the target device is located. The global map is pre-generated by using an online SLAM (Simultaneous Localization And Mapping) method or an offline SFM (Structure From Motion) method based on the initial image of the target scene collected by an image acquisition device. For example, methods such as ORB-SLAM3 (Oriented FAST and Rotated BRIEF-SLAM3), COLMAP (Map), etc. Figure 3 ) and COLMAP (Map), etc.

[0112] When the target device is a vehicle, the target scene is the street, road, etc. where the vehicle is located; the global map is generated from the initial images of the street and road collected by an image acquisition device.

[0113] When the target device is a mobile robot, for example, when the target device is a floor cleaning robot, the target scene is the room cleaned by the floor cleaning robot; the global map is generated based on the initial images collected when the floor cleaning robot first cleaned the room.

[0114] When the target device is an AGV (Automated Guided Vehicle), the target scene is the area where the AGV transports items; the global map is generated based on the initial images collected when the AGV first transported items in this area.

[0115] A camera is installed on the target device, and the camera is used to collect images in real time. The target image is the image of the target scene collected by the target device at the current moment.

[0116] During the movement of the target device, the target device is located based on the images collected by the camera in real time. Before locating the target device based on the target image, the latest positioning status of the target device can be obtained.

[0117] When the target image is not the first image collected by the camera, the latest positioning status is the positioning status when the target device was located when the target device collected the previous frame of the first image. When the target image is the first image collected by the camera, the latest positioning status can be defaulted to the third preset status. Furthermore, based on the latest positioning status, the accuracy of locating the target device is determined, and then the corresponding positioning method is selected to continue to locate the target device.

[0118] Regarding step S102, the latest positioning state represents the target association state between the image collected by the target device and the global map of the target scene; the higher the association state, the higher the accuracy of positioning the target device, and the lower the association state, the lower the accuracy of positioning the target device.

[0119] The latest positioning state includes multiple states. In some embodiments, the latest positioning state includes three states: a first preset state, a second preset state, and a third preset state. Among them, the first preset state indicates a weak target association state between the image collected by the target device and the global map of the target scene; the second preset state indicates a good target association state between the image collected by the target device and the global map of the target scene; the third preset state indicates that the positioning of the target device fails.

[0120] Correspondingly, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0121] When the latest positioning state is the first preset state, it indicates that the target association state between the image collected by the target device and the map of the target scene is weak. It may be caused by changes in the lighting and scene structure of the target scene during the construction of the global map and when positioning the target device. The accuracy of positioning the target device is relatively low, and a temporary map of the target scene can be generated.

[0122] When the first pose corresponding to the previous frame (the first image) of the target image is determined based on the global map and the first image, it indicates that the accuracy of positioning the target device based on the global map is relatively low. A temporary map is generated based on the newly collected target image of the target device and the adjacent image of the target image.

[0123] When the first pose is determined based on the global map, the previously generated temporary map, and the first image, it indicates that the accuracy of positioning the target device based on the global map and the previously generated temporary map is relatively low. Therefore, a new temporary map is generated based on the newly collected target image of the target device and the adjacent image of the target image.

[0124] In some embodiments, on the basis of Figure 1 , referring to Figure 2 , step S102 may include the following steps:

[0125] S1021: Obtain a third preset number of adjacent images from the images collected by the camera before collecting the target image, and determine the target image and the adjacent images as the second images.

[0126] S1022: For each second image, based on the feature description information of the feature points in the previous frame of the second image and the feature description information of the feature points in the second image, determine the positions of the temporary map points corresponding to the feature points in the second image in the target scene, and the second pose corresponding to the second image.

[0127] S1023: Determine the second key frame images in each second image, and obtain the second poses corresponding to the second key frame images, the feature description information of the feature points in the second key frame images, and the positions of the temporary map points, to obtain a temporary map of the target scene.

[0128] The target device obtains, from the images captured by the camera before capturing the target image, the adjacent third preset number of images (i.e., adjacent images) according to the size of the preset sliding window, and determines the target image and the adjacent images as second images. The third preset number is the size of the sliding window. The third preset number is set by a technician according to requirements.

[0129] The adjacent images may be the images adjacent to the target image in time sequence according to the time sequence of the images captured by the camera.

[0130] Exemplarily, refer to Figure 3 , Figure 3 which is a schematic diagram of the principle of generating a temporary map provided by an embodiment of the present application. Tc represents the current frame as the target image. According to the size of the preset sliding window, the adjacent images are determined to include: the images represented by frames T0, T1, T2, T3, T4,... before the current frame, Tc-1.

[0131] For example, the third preset number is 5. When the target image is the 10th frame image captured by the camera, the second images include the 5th to 10th frame images captured by the camera; when the target image is the 12th frame image captured by the camera, the second images include the 7th to 12th frame images captured by the camera.

[0132] Alternatively, the adjacent images may also be the images with a relatively large co-visible area with the target image in the target scene in space. For example, when the movement route of the target device is circular, when the target device moves to the end point to capture the target image, since the end point of the circle coincides with the starting point, the co-visible area between the target image and the image captured by the target device at the starting point is the largest, and the adjacent image of the target image may be the image captured at the starting point.

[0133] For each image, the feature points in the image include: corner points in the target scene, pixel points where the edge points of objects in the target scene are imaged in the image, etc.

[0134] For each feature point, the feature description information of the feature point can be a feature descriptor. For example, the feature description information of the feature point can be a binary feature descriptor, such as ORB, etc. Or, it can also be a floating-point feature descriptor, such as SuperPoint, etc.

[0135] For each feature point, when the feature description information of the feature point is ORB, within the preset neighborhood range of the feature point in the image to which it belongs, multiple feature points are randomly selected. Based on the pixel values of the selected multiple feature points, a feature descriptor of the feature point is generated. For example, for every two selected feature points, binary encoding is performed in such a way that the feature point with a larger pixel value is 1 and the feature point with a smaller pixel value is 0, to obtain the feature descriptor of the second feature point.

[0136] The preset neighborhood range can be: a circular region with the feature point as the center and a specified radius; or, the neighborhood range of the feature point can also be: a rectangular region with the feature point as the center point, a specified width, and a specified length.

[0137] In some embodiments, step S1022 may include the following steps:

[0138] Step 1, based on the feature description information of the feature points in the previous frame of the second image of the second image, and the feature description information of the feature points in the second image, calculate the first matching relationship between the feature points in the second image and the previous frame of the second image.

[0139] Step 2, based on the pixel positions of two matching feature points in the second image and the previous frame of the second image, calculate the distance between the same temporary map point corresponding to the two matching feature points in the target scene and the target device, and determine the position of the temporary map point based on this distance.

[0140] Step 3, based on the first matching relationship and the second pose corresponding to the previous frame of the second image, and the second relative pose of the target device from the time of acquiring the previous frame of the second image to the time of acquiring the second image, obtain the prior pose corresponding to the second image.

[0141] Step 4, based on the second association relationship between the temporary map point and the feature points in the previous frame of the second image, and the prior pose corresponding to the second image, determine the fourth projection position of the temporary map point in the second image, and the fourth candidate matching feature point of the temporary map point in the second image.

[0142] Step 5, adjust the position of the temporary map point and the prior pose corresponding to the second image respectively, and determine the position of the temporary map point and the second pose corresponding to the second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point satisfies the preset convergence condition.

[0143] For each feature point in the second image, the target device selects a feature point from the feature points in the previous-frame second image as the current feature point to be matched, and calculates the descriptor distance between the feature descriptor of this feature point and the feature descriptor of the current feature point to be matched. For example, when the feature descriptor is a binary feature descriptor such as ORB, the Hamming distance of the feature descriptor is calculated. When the feature descriptor is a floating-point feature descriptor such as SuperPoint, the Euclidean distance of the feature descriptor is calculated.

[0144] If the calculated descriptor distance is not less than the second distance threshold, it is determined that this feature point matches the current feature point to be matched. If the calculated descriptor distance is less than the second distance threshold, it is determined that this feature point does not match the current feature point to be matched. The target device selects an unmatched feature point from the feature points in the previous-frame second image as the current feature point to be matched, and calculates the descriptor distance between the feature descriptor of this feature point and the feature descriptor of the current feature point to be matched, and so on. In this way, the feature point in the previous-frame second image that matches each feature point in the second image can be determined, and thus the first matching relationship between the feature points in the second image and the previous-frame second image is obtained.

[0145] For each map point in the target map, the feature points at which this map point is imaged in different second images are different. The matching relationship between the feature points in two adjacent second images includes: the two feature points at which the same map point in the target scene is imaged in two adjacent second images match each other.

[0146] Correspondingly, after the first matching relationship between the feature points in the second image and the previous-frame second image is determined, for the two matching feature points in the second image and the previous-frame second image, these two feature points are the feature points at which the same map point in the target scene is imaged in the second image and the previous-frame second image.

[0147] That is, the target device can observe this map point when collecting the second image and the previous-frame second image. Furthermore, based on the geometric triangulation relationship of multiple views (i.e., the second image and the previous-frame second image), depth estimation can be performed to obtain the distance between the same map point corresponding to these two matching feature points in the target scene and the target device. Based on this distance, the position of this map point is determined to obtain the position of the temporary map point.

[0148] In some embodiments, the second relative pose represents the transformation relationship of moving from the second pose corresponding to the previous-frame second image to the second pose corresponding to this second image. The second relative pose includes: the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis of the world coordinate system when the target device moves from collecting the previous-frame second image to collecting this second target image.

[0149] The second relative pose can be determined based on a target sensor of the target device; the target sensor is an IMU (Inertial Measurement Unit), a wheel speed sensor, etc. The target sensor outputs the rotation angles and translation distances of the target device on the X-axis, Y-axis, and Z-axis of the world coordinate system based on the distance and direction of the movement of the target device.

[0150] Or, it is determined based on a uniform motion model in the case where no target sensor is installed in the target device. The uniform motion model can output the rotation angles and translation distances of the target device on the X-axis, Y-axis, and Z-axis of the world coordinate system after the target device has moved for a certain period of time in the case of the uniform motion of the target device.

[0151] Furthermore, in the case where the second pose corresponding to the previous-frame second image is known, based on the second pose corresponding to the previous-frame second image, recursive calculation is performed according to the second relative pose to obtain the prior pose corresponding to this second image.

[0152] In one implementation manner, the prior pose corresponding to this second image is directly used as the second pose corresponding to this second image. The second key-frame images in each second image are determined, and the second pose corresponding to the second key-frame image, the feature description information of the feature points in the second key-frame image, and the positions of the temporary map points are obtained to generate a temporary map of the target scene.

[0153] The second key-frame images can include the second images in the second image that contain objects, doors, windows, etc. in the target scene, and the second images when the target device moves to a turning point, etc.

[0154] In another implementation manner, in order to improve the accuracy of the temporary map, the reprojection error of the temporary map points (i.e., the fourth reprojection error) is constructed to optimize the positions of the temporary map points and the prior pose corresponding to the second image.

[0155] For each temporary map point, if the temporary map point is associated with the feature points in the previous-frame second image, based on the second association relationship between the temporary map point and the feature points in the previous-frame second image, and the prior pose corresponding to this second image, the temporary map point is projected to the pixel position in this second image to obtain the fourth projection position of the temporary map point in this second image.

[0156] Determine the feature points within the preset neighborhood range of the fourth projection position in the second image, calculate the distances between the feature description information of the fourth projection position and the feature description information of the determined feature points respectively, and determine the feature point with the smallest corresponding distance to obtain the fourth candidate matching feature point for determining the temporary map point in the second image.

[0157] For each temporary map point, calculate the distance between the fourth projection position of the temporary map point and the fourth candidate matching feature point to obtain the fourth reprojection error corresponding to the temporary map point. Ideally, when there is no error in the position of the determined temporary map point and no error in the prior pose corresponding to the determined second image, the fourth projection position and the fourth candidate matching feature point are the same pixel point, and the fourth reprojection error corresponding to the temporary map point is 0.

[0158] Correspondingly, adjust the positions of the temporary map points and the prior pose corresponding to the second image respectively, and based on the adjusted positions of the temporary map points and the pose corresponding to the second image, determine the fourth reprojection errors corresponding to the temporary map points again until the fourth reprojection errors corresponding to the temporary map points meet the preset convergence condition. At this time, the positions of the temporary map points and the pose corresponding to the second image are relatively accurate, then determine the pose corresponding to the second image at this time as the second pose, and determine the positions of the temporary map points at this time.

[0159] The preset convergence condition is: when adjusting the positions of the temporary map points and the pose corresponding to the second image for a continuous fourth preset number of times, the difference in the loss values of the fourth reprojection errors between adjacent adjustments is less than the first preset value.

[0160] Alternatively, set the convergence condition as: when adjusting the positions of the temporary map points and the pose corresponding to the second image for a continuous fourth preset number of times, the adjustment values between the adjustment states of adjacent adjustments are all less than the second preset value. The adjustment states include: the positions of the temporary map points and the pose corresponding to the second image. The adjustment values between the adjustment states include: the difference in the positions of the temporary map points between adjacent adjustments, and the difference in the poses corresponding to the second image between adjacent adjustments.

[0161] Determine the second key frame images in each second image, and obtain the second pose corresponding to the second key frame images, the feature description information of the feature points in the second key frame images, and the positions of the temporary map points to obtain the temporary map of the target scene.

[0162] Based on the above processing, when the latest positioning state is the first preset state, it indicates that the correlation state between the image collected by the target device and the global map is relatively low, which may be caused by changes in the scene structure or illumination of the target scene. Then, generate a temporary map of the target scene based on the second image. Subsequently, the temporary map and the global mapFigure 1 It is used for the association and positioning pose solution of time series frames (multiple consecutive images collected by the target device), that is, the target device is hybridly positioned, the positioning accuracy of the target device is improved, and the robustness of the positioning method is improved.

[0163] Exemplarily, Figure 3 The temporary map in includes: temporary map - fixed and temporary map - participating in local optimization. Among them, the temporary map - fixed refers to the pose corresponding to the temporary map points and key frame images in the previously generated temporary map, and this part of the temporary map does not need to be optimized. The temporary map - participating in local optimization refers to the position of the temporary map points and the second pose of the second key frame image adjusted according to the above method, and optimization will be performed when generating this part of the temporary map. The solid circle represented by Tn corresponds to the optimization state, and the image of the Tn optimization state is: the nth second image used when adjusting the position of the temporary map points and the second pose of the second image according to the above method. The dashed circle represented by Tn corresponds to the fixed state, and the image of the Tn fixed state is: the nth image used when generating the temporary map - fixed. Among them, the value range of n is 0 - c. And the global map is also in a fixed state. Figure 3 The dashed boxes (dashed ellipses and rectangles) in indicate optimization based on the odometry factor. The solid boxes (solid ellipses and rectangles) indicate optimization based on the visual factor.

[0164] The visual factor is the fourth reprojection error in the foregoing embodiments, that is, for the temporary map points, they are back - projected onto the second image through perspective projection, and the error between the fourth projection position of this reprojection and the fourth candidate matching feature point is the fourth reprojection error. By using non - linear optimization means, by minimizing this fourth reprojection error, the position of the optimized temporary map points and the second pose corresponding to the second image are obtained.

[0165] The odometry factor is: a constraint factor constructed based on the error between the relative pose obtained from the target sensor and the relative pose obtained from the pose for positioning the target device. The odometry factor is jointly optimized in combination with the reprojection error, which can further improve the accuracy of the generated temporary map and the accuracy and robustness of the positioning method.

[0166] Exemplarily, see Figure 4 , Figure 4 which is a flowchart of generating a temporary map provided by an embodiment of the present application.

[0167] When the scene structure of the target scene changes or the illumination changes, the quality of positioning the target device based on the global map deteriorates, that is, the latest positioning state is the first preset state, then the construction of the temporary local map starts.

[0168] Obtain the global map of the target scene, the feature points of the current frame, and the integrated navigation information (e.g., the information output by the IMU and wheel speed sensors). The current frame is the target image, and the feature points of the current frame are the feature points in the target image.

[0169] Based on the obtained information, match the temporary map frame with the feature points of the current frame. The temporary map frame is an adjacent image acquired before the target image is acquired. The target image and the adjacent image are collectively referred to as the second image. That is, match the feature points in two adjacent frames of the second image to obtain the first matching relationship between the feature points in two adjacent frames of the second image.

[0170] Based on the calculated first matching relationship, create temporary map points. For example, determine the feature points in the second image that are not associated with the global map points in the global map, and triangulate based on the feature points in two frames of the second image that are not associated with the global map points to obtain the positions of the temporary map points and the second pose corresponding to the second image.

[0171] Perform local optimization on the positions of the temporary map points and the second pose corresponding to the second image to obtain a temporary map. The temporary map includes: key frame and map point information, that is, the second pose corresponding to the second key frame image and the positions of the temporary map points.

[0172] Based on the above processing, optimize by fixing the global map points and optimizing the newly created temporary map points and the second pose corresponding to the second image, so that the coordinate system of the newly created temporary map is consistent with the global map. When subsequently positioning the target device based on the global map and the temporary map, the global map and the temporary map form a hybrid map for positioning the target device, until the positioning quality for positioning the target device improves, which can improve the positioning accuracy of the target device and the robustness of the positioning method.

[0173] Regarding step S103, the second image includes the adjacent image of the target image. Since the target scene changes little when the target image and the adjacent image are acquired, the association state between the temporary map generated based on the second image and the target image is relatively high, which can improve the accuracy of matching the target image with the temporary map. Furthermore, the target device can be positioned based on the temporary map to obtain the target pose corresponding to the target image. The pose of the target device is a 6DOF pose. The 6DOF pose includes the position and orientation of the target device. The position refers to the three-dimensional coordinates of the target device in the world coordinate system; the orientation refers to the rotation angles and translation distances of the target device relative to the coordinate origin in the three coordinate axis directions in the world coordinate system.

[0174] In some embodiments, on the basis of Figure 1 , referring to Figure 5 , step S103 may include the following steps:

[0175] S1031: Determine a prior pose corresponding to the target image based on the first pose corresponding to the previous frame of the first image of the target image and the first relative pose of the target device from when the first image is acquired to when the target image is acquired.

[0176] S1032: Adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map points of the target scene and the feature points in the first image to obtain a candidate pose corresponding to the target image.

[0177] Wherein, the first association relationship is determined when determining the first pose; the first associated map points are map points in the target scene associated with the feature points in the first image.

[0178] S1033: Determine first reference map points from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image.

[0179] S1034: Determine, from the global map points in the global map, map points associated with the feature points in the first key frame image that observes the first reference map point in the global map, and determine, from the temporary map points in the temporary map, map points associated with the feature points in the second key frame image that observes the first reference map point in the temporary map, as first local map points.

[0180] S1035: Adjust the candidate pose corresponding to the target image based on the association between the first local map points and the feature points in the target image to obtain the target pose corresponding to the target image.

[0181] In some embodiments, the first pose can be obtained by positioning the target device based on the global map; or, the first pose can also be obtained by positioning the target device based on the global map and the previously generated temporary map.

[0182] In some embodiments, the first relative pose represents the conversion relationship from the first pose to the pose corresponding to the target image. That is, the first relative pose includes: the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis of the world coordinate system when the target device moves from acquiring the first image to acquiring the target image.

[0183] The first relative pose is determined based on the target sensor in the target device; or determined based on a uniform motion model. The method for determining the first relative pose is similar to the method for determining the second relative pose, and the relevant introduction in the foregoing embodiments can be referred to.

[0184] Correspondingly, in the case where the first pose corresponding to the first image is known, based on the first pose, a prior pose corresponding to the target image is obtained by recursion according to the first relative pose.

[0185] In one implementation, the prior pose corresponding to the target image is directly used as the target pose corresponding to the target image.

[0186] In another implementation, in order to improve the accuracy of the target pose corresponding to the determined target image, the target device adjusts the prior pose corresponding to the target image based on the first association relationship between the first associated map points in the target scene and the feature points in the first image, obtains the candidate pose corresponding to the target image, and determines the target pose corresponding to the target image based on the candidate pose corresponding to the target image.

[0187] In some embodiments, step S1032 may include the following steps:

[0188] Step 1, based on the first association relationship between the first associated map points in the target scene and the feature points in the first image, and the prior pose corresponding to the target image, determine the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image.

[0189] Step 2, adjust the prior pose corresponding to the target image, and determine the candidate pose corresponding to the target image when the first reprojection error between the first projection position and the first candidate matching feature point satisfies a preset convergence condition.

[0190] Wherein, the first reprojection error is the distance between the first projection position and the first candidate matching feature point.

[0191] When determining the first pose, the first association relationship between the map points in the target scene and the feature points in the first image is also determined. The fact that the first associated map point is associated with the feature point in the first image means that the feature point in the first image is the pixel point where the first associated map point is imaged in the first image.

[0192] If the first pose is determined based on the global map, the first associated map point is the global map point in the global map. If the first pose is determined based on the global map and the temporarily determined map last time, the first associated map point includes the global map point in the global map and the temporarily determined map point in the temporarily determined map last time.

[0193] In some embodiments, step 1 may include the following steps: Based on the first association relationship between the first associated map points in the target scene and the feature points in the first image, and the prior pose corresponding to the target image, project the first associated map point to the pixel position in the target image to obtain the first projection position of the first associated map point in the target image. Based on the feature description information of the first projection position and the feature description information of each feature point within the preset neighborhood range of the first projection position in the target image, determine the first candidate matching feature point of the target map point.

[0194] Based on the first association relationship, the target device can determine which map point (i.e., the first associated map point) in the world coordinate system the feature point in the first image is associated with. Further, based on the position of the first associated map point in the world coordinate system, the pixel position of the feature point associated with the first associated map point in the first image, and the prior pose corresponding to the target image, the pixel position where the first associated map point is projected onto the target image can be calculated to obtain the first projection position of the first associated map point in the target image. Determine the feature points within the preset neighborhood range of the first projection position in the target image, calculate the distances between the feature description information of the first projection position and the feature description information of the determined feature points respectively, and determine the feature point with the smallest corresponding distance to obtain the first candidate matching feature point of the first associated map point in the target image.

[0195] For each first associated map point, calculate the distance between the corresponding first projection position and the first candidate matching feature point of the first associated map point to obtain the first reprojection error corresponding to the first associated map point. Ideally, when there is no error in the position of the determined first associated map point and no error in the prior pose corresponding to the target image, the first projection position and the first candidate matching feature point are the same pixel point, and the first reprojection error is 0.

[0196] Correspondingly, adjust the prior pose corresponding to the target image, and based on the adjusted pose corresponding to the target image, determine the first reprojection error corresponding to each first associated map point again until the first reprojection errors corresponding to each first associated map meet the preset convergence condition. At this time, the accuracy of the pose corresponding to the target image is relatively high, and then determine the pose of the target device at this time as the candidate pose corresponding to the target image. The preset convergence condition can refer to the relevant introduction in the foregoing embodiments.

[0197] In one implementation, directly use the candidate pose corresponding to the target image as the target pose corresponding to the target image.

[0198] In another implementation, in order to improve the accuracy of the target pose corresponding to the determined target image, the target device adjusts the candidate pose corresponding to the target image based on the first association relationship to obtain the target pose corresponding to the target image.

[0199] Correspondingly, based on the first association relationship and the candidate pose corresponding to the target image, determine the first reference map point.

[0200] In some embodiments, step S1033 may include the following steps:

[0201] Step 1: Based on the first association relationship and the candidate pose corresponding to the target image, determine the third projection position of the first associated map point in the target image and the third candidate matching feature point of the first associated map point in the target image.

[0202] Step 2: From the first associated map points, determine the map points for which the third reprojection error between the third projection position and the third candidate matching feature point is less than the first distance threshold, to obtain the first reference map points.

[0203] The target device, based on the first association relationship and the candidate pose corresponding to the target image, reprojects the first associated map points to the pixel positions in the target image, to obtain the new third projection position and the new third candidate matching feature point of the first associated map points in the target image. The method for determining the third projection position and the third candidate matching feature point is similar to the method for determining the first projection position and the first candidate matching feature point, and the relevant introduction in the foregoing embodiments can be referred to.

[0204] Calculate the distance between the third projection position and the third candidate matching feature point to obtain the third reprojection error. Ideally, the third reprojection error is 0. For each map point in the target scene, if the third reprojection error corresponding to this map point is less than the first distance threshold, it indicates that the accuracy of the position of this map point is relatively high, and then determine this map point as the first reference map point.

[0205] In some embodiments, when the first associated map points include global map points in the global map, the first reference map points include global map points in the global map. When the first associated map points include global map points in the global map and temporary map points in the previously generated temporary map, the first reference map points include global map points in the global map and temporary map points in the previously generated temporary map.

[0206] In some embodiments, if the number of the first reference map points is not less than the second preset number, according to the co-visibility relationship recorded in the global map, determine the first key-frame image that observes the first reference map points from the first key-frame images in the global map. And determine the map points associated with the feature points in the first key-frame image that observes the first reference map points from the global map points in the global map, as the first local map points.

[0207] According to the co-visibility relationship recorded in the temporary map, determine the second key-frame image that observes the first reference map points from the second key-frame images in the temporary map. And determine the map points associated with the feature points in the second key-frame image that observes the first reference map points from the temporary map points in the temporary map, as the first local map points.

[0208] The co-visibility relationship includes: the key-frame images in which the map points are observed, and the corresponding relationships between the map points observed by the key-frame images. That a key-frame image observes a map point means that the target device can observe the map point when acquiring the key-frame image, that is, the map point can be imaged in the key-frame image.

[0209] Furthermore, based on the second reprojection error constructed from the first local map points, the candidate poses of the target device when acquiring the target image are optimized to obtain the target pose of the target device when acquiring the target image.

[0210] In some embodiments, step S1035 may include the following steps:

[0211] Step 1, based on the first association relationship and the candidate pose corresponding to the target image, determine the second projection position of the first local map point in the target image and the second candidate matching feature point of the first local map point in the target image.

[0212] Step 2, adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when the second reprojection error between the second projection position and the second candidate matching feature point satisfies a preset convergence condition.

[0213] The target device projects the first local map points to the pixel positions in the target image based on the first association relationship and the candidate pose corresponding to the target image, and obtains the second projection position and the second candidate matching feature point of the first local map point in the target image. The method for determining the second projection position and the second candidate matching feature point is similar to the method for determining the first projection position and the first candidate matching feature point, and the relevant introduction in the foregoing embodiments can be referred to.

[0214] For each first associated map point, calculate the distance between the second projection position of the first associated map point and the second candidate matching feature point to obtain the second reprojection error corresponding to the first associated map point. Ideally, the second reprojection error is 0. Correspondingly, adjust the candidate pose corresponding to the target image, and based on the pose corresponding to the adjusted target image, determine the second reprojection error corresponding to each first associated map point again until the second reprojection error corresponding to each first associated map point is less than the preset convergence condition. At this time, the accuracy of the pose corresponding to the target image is relatively high, and then determine the pose of the target device at this time as the target pose corresponding to the target image. The preset convergence condition can refer to the relevant introduction in the foregoing embodiments.

[0215] Based on the above processing, based on the temporary map and the global map Figure 1 start hybrid positioning of the target device, improve the positioning accuracy of the target device, and improve the robustness of the positioning method.

[0216] Exemplarily, refer to Figure 6 , Figure 6 , which is a schematic diagram of the distribution of temporary map points and global map points provided by an embodiment of the present application. Figure 6 In Figure 6 , the circles represent the global map points in the global map, the dots represent the temporary map points in the temporary map, and the curved lines with arrows represent the positioning trajectories.

[0217] When the target scene changes at time i, the distances between the global map points in the global map and the positioning trajectory are all relatively far, that is, there are few global map points with a relatively short distance from the positioning trajectory, that is, the target association degree between the image collected by the target device and the global map is weak, and the accuracy of positioning the target device based on the global map is relatively low.

[0218] The distances between the temporary map points in the temporary map at time i and the temporary map at time j and the positioning trajectory are all relatively short, that is, there are more temporary map points with a relatively short distance from the positioning trajectory, that is, the target association state between the image collected by the target device and the temporary map is good, and the accuracy of positioning the target device based on the temporary map is relatively high.

[0219] In some embodiments, after step S102, the method may further include the following steps: deleting the temporary map after a preset duration for generating the temporary map.

[0220] During the process of real-time positioning of the target device, when it is detected that the latest positioning state is the first preset state and a temporary map of the target scene is generated, multiple temporary maps will be generated. And the target association state between the temporary maps generated before the current time and the image collected by the target device at the current time is weak, so after the preset duration for generating the temporary map, the temporary map is deleted. That is, the temporary maps generated at times relatively far from the current time can be deleted to save the storage space of the target device.

[0221] In some embodiments, after determining the target pose, based on the first association relationship and the target pose corresponding to the target image, a second reference map point is determined from the first associated map points.

[0222] Specifically, based on the first association relationship and the target pose corresponding to the target image, the new projection position and the new candidate matching feature points of the first associated map points in the target image are determined, and the map points with a reprojection error less than the first distance threshold between the new projection position and the new candidate matching feature points are determined from the global map points among the first associated map points to obtain the second reference map point. Correspondingly, the number of the second reference map points can represent the target association state between the image collected by the target device and the global map.

[0223] If the number of second reference map points is less than the first preset number, it indicates that there are fewer map points with relatively small second reprojection errors, so the accuracy of the determined target pose is relatively low, and the target association state between the image collected by the target device and the global map is weak. Then, the latest positioning state is updated to the first preset state.

[0224] If the number of second reference map points is greater than the first preset number, it indicates that there are more map points with relatively small second reprojection errors, so the accuracy of the determined target pose is relatively high, and the target association state between the image collected by the target device and the global map is good. Then, the latest positioning state is updated to the second preset state. The target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state.

[0225] After determining the first reference map points, the latest positioning state of the target device can also be updated based on the first reference map points.

[0226] If the number of first reference map points is less than the second preset number, it indicates that there are fewer map points with relatively small third reprojection errors, and the pose when the target image collected by the target device cannot be determined. Then, the latest positioning state is updated to the third preset state. Here, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0227] Based on the above processing, the latest positioning state represents the accuracy of positioning the target device, and the latest positioning state of the target device is updated. Correspondingly, different positioning modes can be selected based on the latest positioning state to position the target device, which can improve the accuracy of real-time positioning of the target device.

[0228] In some embodiments, after step S101, the method may further include the following steps: when the latest positioning state is the second preset state, based on the first pose corresponding to the first image of the previous frame of the target image, the target image, and the global map, determine the target pose corresponding to the target image.

[0229] When the latest positioning state is the second preset state, it indicates that the target association state between the image collected by the target device and the global map is good. Then, when constructing the global map, the illumination and scene structure of the target scene change less compared to when positioning the target device. Therefore, a temporary map does not need to be generated, and the target device can be directly positioned based on the first pose, the target image, and the global map.

[0230] In some embodiments, the manner of determining the target pose corresponding to the target image based on the first pose, the target image, and the global map may include the following steps:

[0231] Step 1: Determine the prior pose corresponding to the target image based on the first pose corresponding to the previous frame (the first image) of the target image and the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image.

[0232] Step 2: Adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map points of the target scene and the feature points in the first image to obtain the candidate pose corresponding to the target image.

[0233] Among them, the first association relationship is determined when determining the first pose; the first associated map points are the map points in the target scene associated with the feature points in the first image.

[0234] Step 3: Based on the first association relationship and the candidate pose of the target device when acquiring the target image, determine the third reference map points from the first associated map points.

[0235] Step 4: From the global map points in the global map, determine the map points associated with the feature points in the first key frame image that observes the third reference map points in the global map as the second local map points.

[0236] Step 5: Based on the association between the second local map points and the feature points in the target image, adjust the candidate pose corresponding to the target image to obtain the target pose corresponding to the target image.

[0237] The manner in which the target device determines the candidate pose corresponding to the target image can refer to the relevant introduction in the foregoing embodiments.

[0238] The target device constructs the reprojection error corresponding to the first associated map points based on the first association relationship and the candidate pose corresponding to the target image. The manner of constructing the reprojection error corresponding to the first associated map points can refer to the relevant introduction in the foregoing embodiments. For each of the first associated map points, if the reprojection error corresponding to the first associated map point is less than the first distance threshold, it indicates that the accuracy of the position of the first associated map point is relatively high, and then determine the first associated map point as the third reference map point.

[0239] In some embodiments, when the first associated map points include the global map points in the global map, the third reference map points include the global map points in the global map. When the first associated map points include the global map points in the global map and the temporary map points in the temporarily generated map last time, the third reference map points include the global map points in the global map and the temporary map points in the temporarily generated map last time.

[0240] In some embodiments, if the number of third reference map points is not less than a second preset number, according to the co-visibility relationship recorded in the global map, from the first key-frame images in the global map, determine the first key-frame images that observe the third reference map points. And from the global map points in the global map, determine the map points associated with the feature points in the first key-frame images that observe the third reference map points, as the second local map points.

[0241] Furthermore, based on the second local map points, construct the reprojection error corresponding to the second local map points, and optimize the candidate pose corresponding to the target image to obtain the target pose corresponding to the target image. The manner of optimizing the candidate pose corresponding to the target image based on the second local map points is similar to the manner of optimizing the candidate pose corresponding to the target image based on the first local map points, and the relevant introduction of the foregoing embodiments can be referred to.

[0242] Based on the above processing, when the accuracy of positioning the target device is relatively high, the target device is positioned based on the global map, without generating a temporary map, saving the system resources of the target device and improving the efficiency of positioning the target device.

[0243] Exemplarily, see Figure 7 , Figure 7 which is a flowchart of a method for positioning a target device based on a global map provided by an embodiment of the present application.

[0244] Obtain the previous frame state, current frame feature points, integrated navigation information, and global map. The previous frame is the first image, and the previous frame state includes: the first pose corresponding to the first image, and the first association relationship between the feature points in the first image and the global map points. The current frame is the target image, and the current frame feature points are the feature points in the target image. The integrated navigation information is the first relative pose output based on target sensors (such as IMU and wheel speed sensors).

[0245] Perform projection matching on the previous frame map points. The previous frame map points are the first associated map points. Project the first associated map points onto the target image to obtain the first projection positions of the first associated map points, and search within a preset neighborhood range of the first projection positions according to the feature description information to obtain the first candidate matching feature points.

[0246] Optimize the pose based on the result of the projection matching of the previous frame map points. Based on the first relative pose and the first pose, determine the prior pose corresponding to the target image. Construct the first reprojection error between the first projection position and the first candidate matching feature points. Based on the first reprojection error and the integrated navigation information, jointly optimize the prior pose corresponding to the target image to obtain the candidate pose corresponding to the target image. When performing this optimization, the positions of the first associated map points are fixed.

[0247] Determine whether the number of inliers is greater than a threshold. The inliers are the third reference map points. The threshold is the second preset number. Determine the third reference map points based on the first association relationship and the candidate poses of the target device when acquiring the target image. If the number of the third reference map points is less than the second preset number, it is determined that the positioning fails and the target pose is not output.

[0248] When the number of inliers is greater than the threshold, perform local map point projection matching. The local map points are the second local map points. When the number of inliers is greater than the threshold, determine the first key frame images observing the inliers according to the co-visibility relationship in the global map of the inliers (i.e., the third reference map points) optimized in the current frame, and determine the global map points associated with the feature points in these key frame images as the second local map points.

[0249] Optimize the pose based on the local map points to obtain the positioning pose, the positioning state, and the association between the current frame and the map. The positioning pose is the target pose corresponding to the target image, and the positioning state is the latest positioning state corresponding to the target pose. The association between the current frame and the map is the association relationship between the feature points in the target image and the global map points.

[0250] Project the second local map points onto the target image, construct the reprojection error between the projected positions of the second local map points and the candidate matching feature points, and perform joint optimization based on this reprojection error and the relative pose constraint in the integrated navigation information to obtain the target pose corresponding to the target image, without generating a temporary map, saving the system resources of the target device and improving the efficiency of positioning the target device.

[0251] In some embodiments, when the latest positioning states of continuously positioning the target device for multiple frames are all the first preset state, it indicates that the positioning accuracy of the target device is relatively high and there is no need to use a temporary map, so the previously generated temporary map can be deleted to save the storage space of the target device.

[0252] In some embodiments, based on the first association relationship and the target pose corresponding to the target image, determine the fourth reference map points from the global map points in the first associated map points. The method for determining the fourth reference map points is similar to the method for determining the second reference map points, and the relevant introduction in the foregoing embodiments can be referred to. After determining the fourth reference map points, the latest positioning state of the target device can also be updated based on the number of the fourth reference map points. The latest positioning state indicates the target association state between the images acquired by the target device and the global map when positioning the target device based on the global map.

[0253] If the number of fourth reference map points is less than the first preset number, it indicates that there are fewer map points with relatively small reprojection errors, so the accuracy of the determined target pose is relatively low, and the target association state between the image collected by the target device and the map of the target scene is weak. Then, the latest positioning state is updated to the first preset state.

[0254] If the number of fourth reference map points is greater than the first preset number, it indicates that there are more map points with relatively small reprojection errors, so the accuracy of the determined target pose is relatively high, and the target association state between the image collected by the target device and the map of the target scene is good. Then, the latest positioning state is updated to the second preset state.

[0255] After determining the third reference map points, the latest positioning state of the target device can also be updated based on the third reference map points.

[0256] If the number of third reference map points is less than the second preset number, it indicates that there are fewer map points with relatively small reprojection errors, and the pose corresponding to the target image cannot be determined. Then, the latest positioning state is updated to the third preset state.

[0257] Based on the above processing, the latest positioning state represents the accuracy of positioning the target device. By updating the latest positioning state of the target device, different positioning modes can be selected based on the latest positioning state subsequently to perform positioning on the target device, which can improve the accuracy of real-time positioning of the target device.

[0258] In some embodiments, the first preset number when updating the latest positioning state from the first preset state to the second preset state is greater than the first preset number when updating the latest positioning state from the second preset state to the first preset state.

[0259] When the latest positioning state is the first preset state, if the number of second reference map points is greater than the first preset number, the latest positioning state is updated to the second preset state, that is, the positioning state of the target device switches from the first preset state to the second preset state, that is, from a weak positioning state to a good positioning state.

[0260] When the latest positioning state is the second preset state, if the number of fourth reference map points is less than the first preset number, the latest positioning state is updated to the first preset state, that is, the positioning state of the target device switches from the second preset state to the first preset state, that is, from a good positioning state to a weak positioning state.

[0261] Correspondingly, the first preset number for switching from the first preset state to the second preset state is greater than the first preset number for switching from the second preset state to the first preset state.

[0262] For example, the first preset number for switching from the first preset state to the second preset state is 50, and the first preset number for switching from the second preset state to the first preset state is 40. When the number of second reference map points is greater than 50, the latest positioning state is updated to the second preset state. When the number of fourth reference map points is less than 40, the latest positioning state is updated to the first preset state.

[0263] For example, refer to Figure 8 , the weak positioning state is the first preset state. The good positioning state is the second preset state. Switching from the weak positioning state to the good positioning state requires the number of internal points to be greater than the high threshold. Switching from the good positioning state to the weak positioning state requires the number of internal points to be less than the low threshold. There is a difference between the high threshold and the low threshold (which can be called the scissors difference), that is, the threshold for comparing the number of internal points for switching from the good positioning state to the weak positioning state is lower than the threshold for comparing the number of internal points for switching from the weak positioning state to the good positioning state.

[0264] Based on the above processing, it is possible to avoid the positioning state of the target device from switching too frequently between the first preset state and the second preset state, and improve the stability of positioning the target device.

[0265] In some embodiments, after step S101, the method may further include the following steps: when the latest positioning state is the third preset state, based on the target image and the global map, determine the target pose of the target device corresponding to the target image.

[0266] When the latest positioning state is the third preset state, it indicates that the positioning of the target device fails and the pose of the target device cannot be output. Then, the target device needs to be repositioned. Correspondingly, based on the target image and the global map, determine the target pose of the target device corresponding to the target image.

[0267] In some embodiments, the manner of determining the target pose of the target device corresponding to the target image based on the target image and the global map may include the following steps:

[0268] Step 1, based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map, determine the candidate recall image of the target image from the first key frame image.

[0269] Step 2, based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recall image, determine the second matching relationship between the feature points in the target image and the feature points in the candidate recall image.

[0270] Step 3: Based on the second matching relationship and the third association relationship between the feature points in the candidate recalled image and the global map points in the global map, obtain the fourth association relationship between the feature points in the target image and the global map points in the global map.

[0271] Step 4: Perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image.

[0272] The target device generates global description information of the target image based on the feature description information of each feature point in the target image. For example, calculate statistical values of the feature description information of each feature point in the target image based on the target device, such as mean, variance, maximum value, minimum value, etc. Or, based on a pre-trained network model, fuse the feature description information of each feature point in the target image based on the target device to obtain the global description information of the target image.

[0273] Similarly, the target device generates global description information of the first key frame image based on the feature description information of each feature point in the first key frame image. Then, calculate the distance between the global description information of the target image and the global description information of the first key frame image, and determine, from the first key frame images, the key frame images corresponding to a distance less than the third distance threshold to obtain the candidate recalled images of the target image.

[0274] Based on the feature description information of each feature point in the target image and the feature description information of each feature point in the candidate recalled image, determine the second matching relationship between the feature points in the target image and the feature points in the candidate recalled image. The method for determining the second matching relationship is similar to the method for determining the first matching relationship, and reference can be made to the relevant introduction in the foregoing embodiments.

[0275] Furthermore, for each feature point in the target image, the second matching relationship records which feature point in the candidate recalled image this feature point matches. For each feature point in the candidate recalled image, the third association relationship records which map point in the global map this feature point is associated with. Accordingly, based on the second matching relationship and the third association relationship, it can be determined which global map point in the global map the feature point in the target image is associated with, and thus the fourth association relationship between the feature points in the target image and the global map points in the global map can be obtained.

[0276] Based on the fourth association relationship and a preset pose calculation algorithm, perform pose calculation to obtain the target pose corresponding to the target image. The preset pose calculation algorithm is the PnP (Perspective-n-Point) algorithm.

[0277] In some embodiments, step 4 may include the following steps: performing pose calculation based on the fourth association relationship to obtain the candidate pose corresponding to the target image. Determining the prior pose corresponding to the target image based on the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image. Performing consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image. When the candidate pose corresponding to the target image passes the verification, determining the target pose as the candidate pose corresponding to the target image.

[0278] The method for determining the prior pose corresponding to the target image may refer to the relevant introduction in the foregoing embodiments.

[0279] The prior pose and the candidate pose corresponding to the target image are calculated by different methods. In an ideal situation, the prior pose and the candidate pose corresponding to the target image are the same. The difference between the prior pose and the candidate pose corresponding to the target image can indicate the accuracy of the candidate pose corresponding to the target image. Accordingly, the target device performs consistency verification on the candidate pose based on the prior pose corresponding to the target image.

[0280] For example, calculate the distance between the prior pose and the candidate pose corresponding to the target image. When the calculated distance is less than the fourth distance threshold, determine that the candidate pose corresponding to the target image passes the verification; when the calculated distance is not less than the fourth distance threshold, determine that the candidate pose corresponding to the target image fails the verification.

[0281] Alternatively, the target device performs consistency verification on the candidate pose corresponding to the target image based on a preset verification algorithm and the prior pose corresponding to the target image to obtain a verification result indicating whether the candidate pose corresponding to the target image passes the verification. The preset verification algorithm is RANSAC (Random Sample Consensus).

[0282] When the candidate pose corresponding to the target image passes the verification, it indicates that the accuracy of the candidate pose corresponding to the target image is relatively high, and then determine the target pose as the candidate pose corresponding to the target image.

[0283] In some embodiments, the latest positioning state of the target device may also be updated based on the result of performing consistency verification on the prior pose corresponding to the target image.

[0284] When the candidate pose corresponding to the target image passes the verification, it indicates that the accuracy of the candidate pose corresponding to the target image is relatively high, and update the latest positioning state to the second preset state.

[0285] When the candidate pose corresponding to the target image fails the verification, it indicates that the accuracy of the candidate pose corresponding to the target image is relatively low, and the target pose corresponding to the target image cannot be determined. Update the latest positioning state to the third preset state.

[0286] Among them, the target association state corresponding to the third preset state is lower than the target association state corresponding to the second preset state.

[0287] Based on the above processing, the latest positioning state represents the accuracy of positioning the target device. By updating the latest positioning state of the target device, different positioning modes can be selected based on the latest positioning state for positioning the target device subsequently, which can improve the accuracy of real-time positioning of the target device.

[0288] In some embodiments, refer to Figure 9 , Figure 9 which is a flowchart of a positioning state switching provided by an embodiment of the present application.

[0289] The positioning states include three states: weak positioning state (i.e., the first preset state), good positioning state (i.e., the second preset state), and positioning failure state (i.e., the third preset state).

[0290] Refer to Figure 10 , the first preset state corresponds to the hybrid positioning mode; the hybrid positioning mode refers to positioning the target device based on the first pose of the previous frame of the first image, the global map, the temporary map, and the current frame Tc (i.e., the target image). The temporary map is generated based on the fixed frame Tc-1 (i.e., the adjacent image). In the first preset state, the pose of the target device can be normally output. When the target scene changes, the hybrid positioning mode is used. In the hybrid positioning mode, on the one hand, the global map is used for positioning, and on the other hand, the temporary map is updated based on the newly acquired images. Figure 10 The dotted line connecting adjacent images in

[0291] The second preset state corresponds to the global positioning mode; the global positioning mode refers to positioning the target device based on the first pose of the previous frame of the first image (i.e., the fixed frame Tc-1), the global map, and the current frame. The global map is updated based on the fixed frame. In the second preset state, the pose of the target device can be normally output. In most cases, the global positioning mode is used to position the target device. The global positioning mode continuously associates the current frame with the global map and performs continuous positioning.

[0292] The third preset state corresponds to the relocalization mode. The relocalization mode refers to localizing the target device based on the global map and the current frame when the first pose corresponding to the first image of the previous frame is not obtained. For example, the target image is the first frame image collected by the target device, or the localization of the target device fails when collecting the first image. No pose can be output in the third preset state. For example, the target device is in the third preset state when it is just started or the localization of the target device fails. In the third preset state, the relocalization mode is used to attempt to relocalize the target device until the relocalization is successful.

[0293] The switching between the above three states of the target device is described as follows:

[0294] (1) Maintain the positioning failure state;

[0295] Trigger condition: Global positioning fails.

[0296] Operation: Continuously relocalize.

[0297] In the case where the first pose of the first image of the previous frame is not determined, according to the relocalization mode, the target device is localized based on the target image collected by the target device and the global map. If the target pose corresponding to the target image is not determined, that is, the relocalization of the target device fails in the relocalization mode, the target device maintains the positioning failure state. Subsequently, the target device is localized according to the relocalization mode.

[0298] (2) Switch from the positioning failure state to the well-positioned state;

[0299] Trigger condition: Global positioning is successful.

[0300] Operation: Perform global positioning.

[0301] In the case where the first pose of the first image of the previous frame is not determined, according to the relocalization mode, the target device is localized based on the target image collected by the target device and the global map. If the target pose corresponding to the target image is successfully determined, that is, the relocalization of the target device is successful in the relocalization mode, the target device switches from the positioning failure state to the well-positioned state.

[0302] (3) Maintain the well-positioned state;

[0303] Trigger condition: Well-positioned state.

[0304] Operations: 1. Continuously perform global positioning. 2. Maintain the temporary map sliding window (low-frequency clearing of temporary map frames that are far away).

[0305] When the first pose of the first image in the previous frame is determined and the latest positioning state is the second preset state (good positioning state), the target device is positioned according to the global positioning mode based on the first pose, the target image, and the global map. If the number of second reference map points determined during the positioning process is greater than the first preset number, that is, the inlier number is greater than the threshold, the target device remains in the good positioning state. Subsequently, the target device still performs positioning according to the global positioning mode and clears the temporary map that is far from the current moment.

[0306] (4) Switch from the good positioning state to the weak positioning state;

[0307] Trigger condition: Low global positioning quality.

[0308] Operation: Set the key frame to be added to the temporary map.

[0309] When the first pose of the first image in the previous frame is determined and the latest positioning state is the second preset state (good positioning state), the target device is positioned according to the global positioning mode based on the first pose, the target image, and the global map. If the number of second reference map points determined during the positioning process is less than the first preset number, that is, the inlier number is less than the threshold, the target device switches from the good positioning state to the weak positioning state. Subsequently, the target device is positioned according to the hybrid positioning mode, that is, a temporary map of the target scene is generated, and the target device is positioned by combining the temporary map and the global map.

[0310] (5) Switch from the weak positioning state to the good positioning state;

[0311] Trigger condition: Global positioning is successful and of high quality.

[0312] Operation: Update the sliding window.

[0313] When the first pose of the first image in the previous frame is determined and the latest positioning state is the first preset state (weak positioning state), the target device is positioned according to the hybrid positioning mode based on the first pose, the target image, the global map, and the temporary map. If the number of first reference map points determined during the positioning process is greater than the first preset number, that is, the inlier number is greater than the threshold, the target device switches from the weak positioning state to the good positioning state. Subsequently, the target device is positioned according to the global positioning mode, that is, the target device is positioned based on the global map. And, the sliding window can also be updated, and according to the updated sliding window, the temporary map that is far from the current is deleted.

[0314] (6) Remain in the weak positioning state;

[0315] Trigger condition: Low global positioning quality.

[0316] Operation: 1. If the hybrid positioning quality is good, the positioning pose is normally output; if the hybrid positioning quality deteriorates, key frames are set and local map construction is performed. 2. Sliding window maintenance (if an interpolated frame is added, the farthest frame is cleared).

[0317] When the first pose of the first image of the previous frame is determined and the latest positioning state is the first preset state (weak positioning state), the target device is positioned according to the hybrid positioning mode based on the first pose, the target image, the global map, and the temporary map. If the number of first reference map points determined during the positioning process is less than the first preset number, that is, the inlier number is less than the threshold, the target device remains in the weak positioning state. Subsequently, the target device is positioned according to the hybrid positioning mode, that is, the target device is positioned based on the global map.

[0318] See Figure 11 , Figure 11 which is a flowchart of a positioning method provided by an embodiment of the present application. This method is applied to a positioning system, and the positioning system includes: a relocalization module and a continuous positioning module.

[0319] Obtain image feature point data and a map. The image feature point data includes the feature description information of the feature points in the target image collected at the current moment. The map includes: a global map and the temporary map generated last time.

[0320] Perform positioning based on the obtained image feature point data and the map, and determine whether the positioning state is a preset state.

[0321] When the positioning state is the positioning failure state, call the relocalization module to determine the target pose and the positioning state of the target device.

[0322] The relocalization module includes: a feature extraction sub-module, a candidate recall sub-module, a pose solution operator sub-module, and a consistency verification sub-module.

[0323] The feature extraction sub-module is used to extract the feature description information of the feature points in the target image. The candidate recall sub-module is used to perform candidate recall based on the feature description information of the feature points in the target image and determine the candidate recall image from the global map. The pose solution operator sub-module is used to solve the pose based on the candidate recall image to obtain the candidate pose of the target device. The consistency verification sub-module is used to determine the target pose and the positioning state of the target device based on the result of the consistency verification of the pose output by the above pose solution operator sub-module.

[0324] When the positioning state is the positioning good state or the weak positioning state, call the continuous positioning module to determine the target pose and the positioning state of the target device.

[0325] The continuous positioning module includes: a global positioning sub-module, a temporary map construction sub-module, and a hybrid positioning sub-module.

[0326] The global positioning sub-module is used to locate the target device based on the global map when the positioning status is in a good positioning state, so as to obtain the target pose and positioning status of the target device. The temporary map construction sub-module constructs a temporary map based on the target image when the positioning status is in a weak positioning state. The hybrid positioning sub-module is used to locate the target device based on the hybrid map (global map and temporary map), so as to obtain the target pose and positioning status of the target device.

[0327] Based on the above processing, a hybrid positioning method that constructs a temporary map to assist positioning in the case of scene changes and realizes switching of positioning modes according to positioning quality is provided. The hybrid positioning based on the global map and the temporary map has the advantages of high efficiency and high robustness, and improves the visual positioning efficiency. Moreover, the technical solution provided by the embodiments of the present application is applied to related fields such as automatic driving and parking, such as controllers in the field of intelligent driving, and has a wide range of applications. It can solve the problems that traditional visual map positioning methods (such as visual relocalization) cannot utilize the prior information of time series, resulting in large computational time consumption, low robustness, and large fluctuations in positioning poses. And the method based on global map positioning is difficult to cope with changes in the surrounding scene and lighting changes, resulting in positioning failures.

[0328] See Figure 12 , Figure 12 is a structural diagram of a positioning device provided by an embodiment of the present application. The device is applied to a target device, and a camera is installed on the target device. The device includes:

[0329] A data acquisition module 1201, configured to acquire the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning status of positioning the target device;

[0330] A temporary map generation module 1202, configured to determine the target image and the adjacent image collected before collecting the target image as the second image when the latest positioning status is a first preset status, and generate a temporary map of the target scene based on the second image; wherein, the latest positioning status represents the target association status between the image collected by the target device and the global map;

[0331] A first positioning module 1203, configured to determine the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the previous frame of the first image of the target image; wherein, the pose corresponding to an image is: the pose of the target device when the image is collected.

[0332] Optionally, the first positioning module 1203 is specifically configured to determine a prior pose corresponding to the target image based on the first pose corresponding to the previous-frame first image of the target image and the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image; adjust the prior pose corresponding to the target image based on a first association relationship between a first associated map point of the target scene and a feature point in the first image to obtain a candidate pose corresponding to the target image; wherein the first association relationship is determined when determining the first pose; the first associated map point is a map point in the target scene associated with the feature point in the first image; determine a first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine, from the global map points in the global map, map points associated with the feature points in the first key-frame image in which the first reference map point is observed in the global map, and determine, from the temporary map points in the temporary map, map points associated with the feature points in the second key-frame image in which the first reference map point is observed in the temporary map, as first local map points; and adjust the candidate pose corresponding to the target image based on associating the first local map points with the feature points in the target image to obtain the target pose corresponding to the target image.

[0333] Optionally, the first positioning module 1203 is specifically configured to determine a first projection position of the first associated map point in the target image and a first candidate matching feature point of the first associated map point in the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image and the prior pose corresponding to the target image; adjust the prior pose corresponding to the target image to determine a candidate pose corresponding to the target image when a first reprojection error between the first projection position and the first candidate matching feature point satisfies a preset convergence condition; wherein the first reprojection error is the distance between the first projection position and the first candidate matching feature point.

[0334] Optionally, the first positioning module 1203 is specifically configured to project the first associated map point to a pixel position in the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image and the prior pose corresponding to the target image to obtain a first projection position of the first associated map point in the target image; and determine a first candidate matching feature point of the target map point based on the feature description information of the first projection position and the feature description information of each feature point within a preset neighborhood range of the first projection position in the target image.

[0335] Optionally, the first positioning module 1203 is specifically configured to determine a second projection position of the first local map point in the target image and a second candidate matching feature point of the first local map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when a second reprojection error between the second projection position and the second candidate matching feature point satisfies a preset convergence condition.

[0336] Optionally, the first positioning module 1203 is specifically configured to determine a third projection position of the first associated map point in the target image and a third candidate matching feature point of the first associated map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; determine, from the first associated map points, map points for which a third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold, to obtain first reference map points.

[0337] Optionally, the apparatus further includes:

[0338] A first positioning state update module, configured to, after the first positioning module performs an adjustment on the candidate pose corresponding to the target image based on an association between the first local map point and a feature point in the target image to obtain the target pose corresponding to the target image, determine second reference map points from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; if the number of the second reference map points is less than a first preset number, update the latest positioning state to a first preset state; if the number of the second reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state.

[0339] The apparatus further includes:

[0340] A second positioning state update module, configured to, after the first positioning module determines, from the first associated map points, map points for which a third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold to obtain first reference map points, perform an update to the latest positioning state to a third preset state if the number of the first reference map points is less than a second preset number; wherein the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

[0341] Optionally, the first relative pose is determined based on a target sensor in the target device; the first relative pose includes: when the target device moves from capturing the first image to capturing the target image, the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis in the world coordinate system.

[0342] Optionally, the temporary map generation module 1202 is specifically configured to obtain a third preset number of adjacent images from the images captured by the camera before capturing the target image, and determine the target image and the adjacent images as second images; for each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determine the position of the temporary map points corresponding to the feature points in this second image in the target scene, and the second pose corresponding to this second image; determine the second key frame images in each second image, and obtain the second pose corresponding to the second key frame image, the feature description information of the feature points in the second key frame image, and the position of the temporary map points, to obtain the temporary map of the target scene.

[0343] Optionally, the temporary map generation module 1202 is specifically configured to calculate a first matching relationship between the feature points in this second image and the feature points in the previous frame of the second image based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image; calculate the distance between the two matching feature points in the target scene corresponding to the same temporary map point and the target device based on the pixel positions of the two matching feature points in this second image and the previous frame of the second image, and determine the position of the temporary map point based on this distance; obtain the prior pose corresponding to this second image based on the first matching relationship, the second pose corresponding to the previous frame of the second image, and the second relative pose of the target device from capturing the previous frame of the second image to capturing this second image; determine the fourth projection position of the temporary map point in this second image and the fourth candidate matching feature point of the temporary map point in this second image based on the second association relationship between the temporary map point and the feature points in the previous frame of the second image and the prior pose corresponding to this second image; respectively adjust the position of the temporary map point and the prior pose corresponding to this second image, and determine the position of the temporary map point and the second pose corresponding to this second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point satisfies a preset convergence condition.

[0344] Optionally, the device further includes:

[0345] A second positioning module, configured to, after the data acquisition module 1201 acquires the global map of the target scenario where the target device is located, the target image of the target scenario acquired by the camera, and the latest positioning status of positioning the target device, determine the target pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image, the target image, and the global map when the latest positioning status is a second preset status.

[0346] Optionally, the second positioning module is specifically configured to determine a prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image and the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image; adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map points of the target scenario and the feature points in the first image to obtain a candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map points are the map points in the target scenario associated with the feature points in the first image; determine a third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image; determine, as a second local map point, the map point in the global map associated with the feature points in the first key frame image that observes the third reference map point in the global map; and adjust the candidate pose corresponding to the target image based on the association between the second local map point and the feature points in the target image to obtain the target pose corresponding to the target image.

[0347] Optionally, the device further includes:

[0348] A third positioning status update module, configured to, after the second positioning module adjusts the candidate pose corresponding to the target image based on the association between the second local map point and the feature points in the target image to obtain the target pose corresponding to the target image, determine a fourth reference map point from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; if the number of the fourth reference map points is less than a first preset number, update the latest positioning status to a first preset status; if the number of the fourth reference map points is greater than the first preset number, update the latest positioning status to a second preset status; wherein, the target association status corresponding to the first preset status is lower than the target association status corresponding to the second preset status;

[0349] The device further includes:

[0350] A fourth positioning status update module, configured to, after the second positioning module determines a third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image, execute: if the number of the third reference map points is less than a second preset number, update the latest positioning status to a third preset status; wherein the second preset number is less than the first preset number; and the target association status corresponding to the third preset status is lower than the target association status corresponding to the first preset status.

[0351] Optionally, the first preset number when updating the latest positioning status to the second preset status when the latest positioning status is the first preset status is greater than the first preset number when updating the latest positioning status to the first preset status when the latest positioning status is the second preset status.

[0352] Optionally, the apparatus further includes:

[0353] A third positioning module, configured to, after the data acquisition module 1201 acquires the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning status for positioning the target device, execute: when the latest positioning status is the third preset status, determine the target pose corresponding to the target image based on the target image and the global map.

[0354] Optionally, the third positioning module is specifically configured to: determine a candidate recall image of the target image from the first key frame image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map; determine a second matching relationship between the feature points in the target image and the feature points in the candidate recall image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recall image; obtain a fourth association relationship between the feature points in the target image and the global map points in the global map based on the second matching relationship and a third association relationship between the feature points in the candidate recall image and the global map points in the global map; and perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image.

[0355] Optionally, the third positioning module is specifically configured to perform pose solution based on the fourth association relationship to obtain a candidate pose corresponding to the target image; determine a prior pose corresponding to the target image based on a first relative pose of the target device from when the first image is acquired to when the target image is acquired; perform consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image; and when the candidate pose corresponding to the target image passes the verification, determine that the target pose is the candidate pose corresponding to the target image.

[0356] Optionally, the apparatus further includes:

[0357] A fifth positioning status update module, configured to, after the third positioning module performs consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image, update the latest positioning status to a second preset status when the candidate pose corresponding to the target image passes the verification; and update the latest positioning status to a third preset status when the candidate pose corresponding to the target image fails to pass the verification; wherein a target association status corresponding to the third preset status is lower than a target association status corresponding to the second preset status.

[0358] Optionally, the apparatus further includes:

[0359] A temporary map deletion module, configured to, after the temporary map generation module generates a temporary map of the target scene based on the second image, delete the temporary map after a preset duration from when the temporary map is generated.

[0360] Based on the positioning apparatus provided in the embodiments of the present application, when the latest positioning status is a first preset status, it indicates that the target association status between the image acquired by the target device and the global map of the target scene is low, that is, the accuracy of positioning the target device is reduced. Then, the target image and the adjacent image acquired before the target image are determined as the second image, and a temporary map of the target scene is generated based on the second image. Since the target scene changes little when the target image and the adjacent image are acquired, when the target device is positioned based on the temporary map, the association status between the target image and the temporary map is high, that is, the accuracy of matching the target image and the temporary map is high. Furthermore, the accuracy of positioning the target device can be improved.

[0361] Embodiments of the present application further provide an electronic device, as Figure 13 shown, including:

[0362] A memory 1301 for storing a computer program;

[0363] A processor 1302, configured to implement the steps of any of the above positioning methods when executing the program stored in the memory 1301.

[0364] And the above-mentioned electronic device may further include a communication bus and / or a communication interface, and the processor 1302, the communication interface, and the memory 1301 complete communication with each other through the communication bus.

[0365] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0366] The communication interface is used for communication between the above-mentioned electronic device and other devices.

[0367] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0368] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0369] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above positioning methods are implemented.

[0370] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the positioning methods in the above embodiments.

[0371] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0372] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0373] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0374] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A positioning method, characterized in that, The method is applied to a target device, on which a camera is installed. The method includes: Obtaining a global map of a target scene where the target device is located, a target image of the target scene collected by the camera, and the latest positioning state of positioning the target device. When the latest positioning state is a first preset state, determining the target image and an adjacent image collected before collecting the target image as a second image, and generating a temporary map of the target scene based on the second image; wherein, the latest positioning state represents the target association state between the image collected by the target device and the global map. Determining the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image; wherein, the pose corresponding to an image is: the pose of the target device when collecting the image.

2. The method according to claim 1, characterized in that, The determining the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image includes: Determining the prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image and the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image. Adjusting the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map point is the map point in the target scene associated with the feature point in the first image. Determining a first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image. Determining, from the global map points in the global map, the map points associated with the feature points in the first key frame image that observes the first reference map point in the global map, and determining, from the temporary map points in the temporary map, the map points associated with the feature points in the second key frame image that observes the first reference map point in the temporary map as first local map points. Adjusting the candidate pose corresponding to the target image based on the association between the first local map points and the feature points in the target image to obtain the target pose corresponding to the target image.

3. The method according to claim 2, wherein The adjusting the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image to obtain the candidate pose corresponding to the target image includes: Determining the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image based on the first association relationship between the first associated map point in the target scene and the feature point in the first image and the prior pose corresponding to the target image. Adjust the prior pose corresponding to the target image, and determine the candidate pose corresponding to the target image when the first reprojection error between the first projection position and the first candidate matching feature point satisfies a preset convergence condition; wherein, the first reprojection error is the distance between the first projection position and the first candidate matching feature point.

4. The method according to claim 3, wherein The determining the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image based on the first association relationship between the first associated map point in the target scene and the feature point in the first image, and the prior pose corresponding to the target image, includes: Based on the first association relationship between the first associated map point in the target scene and the feature point in the first image, and the prior pose corresponding to the target image, project the first associated map point to the pixel position in the target image to obtain the first projection position of the first associated map point in the target image; Based on the feature description information of the first projection position and the feature description information of each feature point within a preset neighborhood range of the first projection position in the target image, determine the first candidate matching feature point of the target map point.

5. The method according to claim 2, wherein The adjusting the candidate pose corresponding to the target image based on the association between the first local map point and the feature point in the target image to obtain the target pose corresponding to the target image, includes: Based on the first association relationship and the candidate pose corresponding to the target image, determine the second projection position of the first local map point in the target image and the second candidate matching feature point of the first local map point in the target image; Adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when the second reprojection error between the second projection position and the second candidate matching feature point satisfies a preset convergence condition.

6. The method according to claim 2, characterized in that The determining the first reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image, includes: Based on the first association relationship and the candidate pose corresponding to the target image, determine the third projection position of the first associated map point in the target image and the third candidate matching feature point of the first associated map point in the target image; From the first associated map points, determine the map points whose third reprojection error between the third projection position and the third candidate matching feature point is less than a first distance threshold to obtain the first reference map point.

7. The method according to claim 6, characterized in that, After the adjusting the candidate pose corresponding to the target image based on the association between the first local map point and the feature point in the target image to obtain the target pose corresponding to the target image, the method further includes: Based on the first association relationship and the target pose corresponding to the target image, determine the second reference map point from the global map points among the first associated map points; If the number of the second reference map points is less than a first preset number, update the latest positioning status to a first preset status; If the number of the second reference map points is greater than the first preset number, update the latest positioning status to a second preset status; wherein, the target association status corresponding to the first preset status is lower than the target association status corresponding to the second preset status; After determining, from the first associated map points, the map points for which the third reprojection error between the third projection position and the third candidate matching feature points is less than a first distance threshold to obtain first reference map points, the method further includes: If the number of the first reference map points is less than a second preset number, update the latest positioning status to a third preset status; wherein, the second preset number is less than the first preset number; the target association status corresponding to the third preset status is lower than the target association status corresponding to the first preset status.

8. The method according to claim 2, wherein The first relative pose is determined based on a target sensor in the target device; the first relative pose includes: the rotation angles and translation distances on the X-axis, Y-axis, and Z-axis of the world coordinate system when the target device moves from capturing the first image to capturing the target image.

9. The method according to claim 1, characterized in that The determining the target image and the adjacent image captured before capturing the target image as a second image, and generating a temporary map of the target scene based on the second image, includes: Obtaining a third preset number of adjacent images from the images captured by the camera before capturing the target image, and determining the target image and the adjacent images as the second images; For each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determining the position of the feature points in this second image corresponding to the temporary map points in the target scene, and the second pose corresponding to this second image; Determining the second key frame images in each second image, and obtaining the second poses corresponding to the second key frame images, the feature description information of the feature points in the second key frame images, and the positions of the temporary map points, to obtain the temporary map of the target scene.

10. The method according to claim 9, characterized in that, The for each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determining the position of the feature points in this second image corresponding to the temporary map points in the target scene, and the second pose corresponding to this second image, includes: Calculating a first matching relationship between the feature points in this second image and the feature points in the previous frame of the second image based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image; Calculating the distance between the same temporary map point corresponding to the two matched feature points in the target scene and the target device based on the pixel positions of the two matched feature points in this second image and the previous frame of the second image, and determining the position of the temporary map point based on this distance; Obtain the prior pose corresponding to the second image based on the first matching relationship, the second pose corresponding to the second image of the previous frame, and the second relative pose of the target device from the time of acquiring the second image of the previous frame to the time of acquiring this second image; Based on the second association relationship between the temporary map point and the feature point in the second image of the previous frame, and the prior pose corresponding to this second image, determine the fourth projection position of the temporary map point in this second image, and the fourth candidate matching feature point of the temporary map point in this second image; Adjust the position of the temporary map point and the prior pose corresponding to this second image respectively, and determine the position of the temporary map point and the second pose corresponding to this second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point meets the preset convergence condition.

11. The method according to claim 1, wherein After obtaining the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning state for positioning the target device, the method further includes: When the latest positioning state is the second preset state, determine the target pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image, the target image, and the global map.

12. The method according to claim 11, wherein The determining the target pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image, the target image, and the global map includes: Determine the prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image, and the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image; Adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image, to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map point is the map point in the target scene associated with the feature point in the first image; Based on the first association relationship and the candidate pose corresponding to the target image, determine the third reference map point from the first associated map points; Determine, from the global map points in the global map, the map point associated with the feature point in the first key frame image that observes the third reference map point in the global map as the second local map point; Based on associating the second local map point with the feature point in the target image, adjust the candidate pose corresponding to the target image to obtain the target pose corresponding to the target image.

13. The method according to claim 12, wherein After adjusting the candidate pose corresponding to the target image based on associating the second local map point with the feature point in the target image to obtain the target pose corresponding to the target image, the method further includes: Based on the first association relationship and the target pose corresponding to the target image, determine the fourth reference map point from the global map points in the first associated map points; If the number of the fourth reference map points is less than a first preset number, update the latest positioning state to a first preset state; If the number of the fourth reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state; After determining third reference map points from the first associated map points based on the first association relationship and the candidate poses corresponding to the target image, the method further includes: If the number of the third reference map points is less than a second preset number, update the latest positioning state to a third preset state; wherein, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state.

14. The method according to claim 13, wherein The first preset number when updating the latest positioning state from the first preset state to the second preset state is greater than the first preset number when updating the latest positioning state from the second preset state to the first preset state.

15. The method according to claim 1, characterized in that, After obtaining the global map of the target scene where the target device is located, the target image of the target scene collected by the camera, and the latest positioning state of positioning the target device, the method further includes: When the latest positioning state is the third preset state, determine the target pose corresponding to the target image based on the target image and the global map.

16. The method according to claim 15, characterized in that, Determining the target pose corresponding to the target image based on the target image and the global map includes: Based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map, determine the candidate recalled image of the target image from the first key frame image; Based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recalled image, determine the second matching relationship between the feature points in the target image and the feature points in the candidate recalled image; Based on the second matching relationship and the third association relationship between the feature points in the candidate recalled image and the global map points in the global map, obtain the fourth association relationship between the feature points in the target image and the global map points in the global map; Perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image.

17. The method according to claim 16, characterized in that, Performing pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image includes: Perform pose calculation based on the fourth association relationship to obtain the candidate pose corresponding to the target image; Based on the first relative pose of the target device from the time of collecting the first image to the time of collecting the target image, determine the prior pose corresponding to the target image; Based on the prior pose corresponding to the target image, perform consistency verification on the candidate pose corresponding to the target image. When the candidate pose corresponding to the target image passes the verification, determine the target pose as the candidate pose corresponding to the target image.

18. The method according to claim 17, wherein After performing the consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image, the method further includes: When the candidate pose corresponding to the target image passes the verification, update the latest positioning state to a second preset state; When the candidate pose corresponding to the target image fails to pass the verification, update the latest positioning state to a third preset state; Wherein, the target association state corresponding to the third preset state is lower than the target association state corresponding to the second preset state.

19. The method according to any one of claims 1 to 18, characterized in that, After generating the temporary map of the target scene based on the second image, the method further includes: Delete the temporary map after a preset duration of generating the temporary map.

20. A positioning device, characterized in that, The device is applied to a target device, a camera is installed on the target device, and the device includes: A data acquisition module, configured to acquire a global map of the target scene where the target device is located, a target image of the target scene acquired by the camera, and the latest positioning state for positioning the target device; A temporary map generation module, configured to, when the latest positioning state is a first preset state, determine the target image and an adjacent image acquired before the target image as the second image, and generate a temporary map of the target scene based on the second image; wherein, the latest positioning state represents the target association state between the image acquired by the target device and the global map; A first positioning module, configured to determine the target pose corresponding to the target image based on the target image, the global map, the temporary map, and the first pose corresponding to the first image of the previous frame of the target image; wherein, the pose corresponding to an image is: the pose of the target device when the image is acquired.

21. The device according to claim 20, wherein, Specifically, the first positioning module is configured to determine the prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image and the first relative pose of the target device from the time when the first image is acquired to the time when the target image is acquired; Adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map point of the target scene and the feature point in the first image to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when the first pose is determined; the first associated map point is the map point in the target scene associated with the feature point in the first image; Based on the first association relationship and the candidate pose corresponding to the target image, determine a first reference map point from the first associated map points; Determine, from the global map points in the global map, the map points associated with the feature points in the first key frame image that observes the first reference map point in the global map, and determine, from the temporary map points in the temporary map, the map points associated with the feature points in the second key frame image that observes the first reference map point in the temporary map, as the first local map points; Based on the association between the first local map points and the feature points in the target image, adjust the candidate pose corresponding to the target image to obtain the target pose corresponding to the target image; The first positioning module is specifically configured to determine the first projection position of the first associated map point in the target image and the first candidate matching feature point of the first associated map point in the target image based on the first association relationship between the first associated map point in the target scene and the feature points in the first image, and the prior pose corresponding to the target image; Adjust the prior pose corresponding to the target image, and determine the candidate pose corresponding to the target image when the first reprojection error between the first projection position and the first candidate matching feature point meets a preset convergence condition; wherein, the first reprojection error is the distance between the first projection position and the first candidate matching feature point; The first positioning module is specifically configured to project the first associated map point to the pixel position in the target image based on the first association relationship between the first associated map point in the target scene and the feature points in the first image, and the prior pose corresponding to the target image, to obtain the first projection position of the first associated map point in the target image; Based on the feature description information of the first projection position and the feature description information of each feature point within a preset neighborhood range of the first projection position in the target image, determine the first candidate matching feature point of the target map point; The first positioning module is specifically configured to determine the second projection position of the first local map point in the target image and the second candidate matching feature point of the first local map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; Adjust the candidate pose corresponding to the target image, and determine the target pose corresponding to the target image when the second reprojection error between the second projection position and the second candidate matching feature point meets the preset convergence condition; The first positioning module is specifically configured to determine the third projection position of the first associated map point in the target image and the third candidate matching feature point of the first associated map point in the target image based on the first association relationship and the candidate pose corresponding to the target image; From the first associated map points, determine the map points where the third reprojection error between the third projection position and the third candidate matching feature point is less than the first distance threshold, to obtain the first reference map points; The device further includes: The first positioning state update module is configured to, after the first positioning module performs the operation of adjusting the candidate pose corresponding to the target image based on the association between the first local map points and the feature points in the target image to obtain the target pose corresponding to the target image, determine the second reference map points from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; If the number of the second reference map points is less than a first preset number, update the latest positioning state to a first preset state; If the number of the second reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state; The device further includes: A second positioning state update module, configured to, after the first positioning module executes to determine, from the first associated map points, map points for which a third reprojection error between the third projection position and the third candidate matching feature points is less than a first distance threshold, and obtain first reference map points, execute: if the number of the first reference map points is less than a second preset number, update the latest positioning state to a third preset state; Wherein, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state; The first relative pose is determined based on a target sensor in the target device; the first relative pose includes: rotation angles and translation distances on the X-axis, Y-axis, and Z-axis of the world coordinate system when the target device moves from collecting the first image to collecting the target image; The temporary map generation module is specifically configured to obtain, from images collected by the camera before collecting the target image, a third preset number of adjacent images, and determine the target image and the adjacent images as second images; For each second image, based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image, determine the position of the feature points in this second image corresponding to the temporary map points in the target scene, and the second pose corresponding to this second image; Determine second key frame images in each second image, and obtain the second poses corresponding to the second key frame images, the feature description information of the feature points in the second key frame images, and the positions of the temporary map points, to obtain a temporary map of the target scene; The temporary map generation module is specifically configured to calculate a first matching relationship between the feature points in this second image and the feature points in the previous frame of the second image based on the feature description information of the feature points in the previous frame of the second image of this second image and the feature description information of the feature points in this second image; Based on the pixel positions of two matching feature points in this second image and the previous frame of the second image, calculate the distance between the same temporary map point corresponding to the two matching feature points in the target scene and the target device, and determine the position of the temporary map point based on this distance; Based on the first matching relationship, the second pose corresponding to the previous frame of the second image, and the second relative pose of the target device from collecting the previous frame of the second image to collecting this second image, obtain the prior pose corresponding to this second image; Based on the second association relationship between the temporary map point and the feature points in the second image of the previous frame, and the prior pose corresponding to the second image, determine the fourth projection position of the temporary map point in the second image and the fourth candidate matching feature point of the temporary map point in the second image; Adjust the position of the temporary map point and the prior pose corresponding to the second image respectively, and determine the position of the temporary map point and the second pose corresponding to the second image when the fourth reprojection error between the fourth projection position and the fourth candidate matching feature point meets the preset convergence condition; The device further includes: A second positioning module, configured to, after the data acquisition module executes to acquire the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning state for positioning the target device, execute when the latest positioning state is the second preset state, based on the first pose corresponding to the first image of the previous frame of the target image, the target image, and the global map, determine the target pose corresponding to the target image; Specifically, the second positioning module is configured to determine the prior pose corresponding to the target image based on the first pose corresponding to the first image of the previous frame of the target image and the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image; Adjust the prior pose corresponding to the target image based on the first association relationship between the first associated map points of the target scene and the feature points in the first image to obtain the candidate pose corresponding to the target image; wherein, the first association relationship is determined when determining the first pose; the first associated map points are the map points in the target scene associated with the feature points in the first image; Based on the first association relationship and the candidate pose corresponding to the target image, determine the third reference map point from the first associated map points; Determine, from the global map points in the global map, the map points associated with the feature points in the first key frame image where the third reference map point is observed in the global map as the second local map points; Based on associating the second local map points with the feature points in the target image, adjust the candidate pose corresponding to the target image to obtain the target pose corresponding to the target image; The device further includes: A third positioning state update module, configured to, after the second positioning module executes to adjust the candidate pose corresponding to the target image based on associating the second local map points with the feature points in the target image to obtain the target pose corresponding to the target image, execute to determine the fourth reference map point from the global map points in the first associated map points based on the first association relationship and the target pose corresponding to the target image; If the number of the fourth reference map points is less than the first preset number, update the latest positioning state to the first preset state; If the number of the fourth reference map points is greater than the first preset number, update the latest positioning state to a second preset state; wherein, the target association state corresponding to the first preset state is lower than the target association state corresponding to the second preset state; The device further includes: A fourth positioning state update module, configured to, after the second positioning module determines a third reference map point from the first associated map points based on the first association relationship and the candidate pose corresponding to the target image, execute an operation of updating the latest positioning state to a third preset state if the number of the third reference map points is less than a second preset number; wherein, the second preset number is less than the first preset number; the target association state corresponding to the third preset state is lower than the target association state corresponding to the first preset state; The first preset number when updating the latest positioning state to the second preset state when the latest positioning state is the first preset state is greater than the first preset number when updating the latest positioning state to the first preset state when the latest positioning state is the second preset state; The device further includes: A third positioning module, configured to, after the data acquisition module acquires the global map of the target scene where the target device is located, the target image of the target scene acquired by the camera, and the latest positioning state of positioning the target device, execute an operation of determining the target pose corresponding to the target image based on the target image and the global map when the latest positioning state is the third preset state; The third positioning module is specifically configured to determine a candidate recall image of the target image from the first key frame image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the first key frame image in the global map; Determine a second matching relationship between the feature points in the target image and the feature points in the candidate recall image based on the feature description information of the feature points in the target image and the feature description information of the feature points in the candidate recall image; Based on the second matching relationship and the third association relationship between the feature points in the candidate recall image and the global map points in the global map, obtain a fourth association relationship between the feature points in the target image and the global map points in the global map; Perform pose calculation based on the fourth association relationship to obtain the target pose corresponding to the target image; The third positioning module is specifically configured to perform pose calculation based on the fourth association relationship to obtain a candidate pose corresponding to the target image; Determine the prior pose corresponding to the target image based on the first relative pose of the target device from the time of acquiring the first image to the time of acquiring the target image; Perform consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image; When the candidate pose corresponding to the target image passes the verification, determine the target pose as the candidate pose corresponding to the target image; The device further includes: The fifth positioning status update module is configured to, after the third positioning module performs consistency verification on the candidate pose corresponding to the target image based on the prior pose corresponding to the target image, update the latest positioning status to a second preset status when the candidate pose corresponding to the target image passes the verification; when the candidate pose corresponding to the target image fails to pass the verification, update the latest positioning status to a third preset status; wherein the target association status corresponding to the third preset status is lower than the target association status corresponding to the second preset status; The apparatus further includes: The temporary map deletion module is configured to, after the temporary map generation module generates a temporary map of the target scene based on the second image, delete the temporary map after a preset duration from the generation of the temporary map.

22. An electronic device, characterized in that, including: a memory for storing a computer program; a processor for implementing the method according to any one of claims 1-19 when executing the program stored on the memory.

23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program, when executed by the processor, implements the method according to any one of claims 1-19.