A visual positioning method, device, electronic device, and storage medium

By combining feature information from current and historical frame images and optimizing transformation parameters, the problem of inaccurate detection results from single-frame images in visual positioning is solved, thereby improving the accuracy and stability of robot visual positioning.

CN116563378BActive Publication Date: 2026-04-03INNER MONGOLIA ZHENGXIN CLOUD TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing visual localization methods, the accuracy of single-frame image detection results fluctuates greatly, leading to inaccurate localization results. Furthermore, judging localization results solely based on simple standards can easily cause system anomalies.

Method used

By acquiring the feature information of the current frame image and combining it with the matching relationship between the feature information of historical frame images in the current and offline maps, the transformation parameters between the current coordinate system and the coordinate system of the offline map are determined. The transformation parameters are then optimized using the feature information of multiple frames to improve the positioning accuracy.

Benefits of technology

It improves the accuracy and stability of robot visual positioning, reduces positioning error, and enhances the reliability of positioning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563378B_ABST
    Figure CN116563378B_ABST
Patent Text Reader

Abstract

This application discloses a visual positioning method, apparatus, electronic device, and storage medium. The visual positioning method includes: determining matching feature information of the current frame image in an offline map based on the feature information of the acquired current frame image; determining transformation parameters between the current coordinate system and the coordinate system of the offline map based on the feature information of the current frame image and its corresponding matching feature information, and the feature information of historical frame images preceding the current frame image and their corresponding matching feature information; determining whether to output the detection position information of the current frame image in the offline map based on the transformation parameters; and determining the detection position information of the current frame image through the position information of the matching feature information of the current frame image in the offline map. The visual positioning method of this application provides more accurate and stable positioning results, improving the positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a visual positioning method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Visual localization is a crucial problem in computer vision and robotics. It has important applications in many fields, such as augmented reality (AR), virtual reality (VR), and robot path planning.

[0003] Because visual relocalization often uses only the detection results of a single frame, the accuracy of the localization results can fluctuate significantly. A single incorrect or inaccurate localization result can easily cause the system to malfunction. Furthermore, the visual localization process relies on simple criteria (such as the number of data points that conform to the model or the distribution of the data) to judge the localization results, leading to inaccurate outcomes. Summary of the Invention

[0004] This application provides at least one visual positioning method, device, electronic device, and storage medium.

[0005] The first aspect of this application provides a visual positioning method, which includes:

[0006] Based on the feature information of the current frame image, determine the matching feature information of the current frame image in the offline map;

[0007] Based on the position information of the feature information of the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, as well as the position information of the feature information of the historical frame images before the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, the transformation parameters between the coordinate system of the current coordinate system and the coordinate system of the offline map are determined.

[0008] Based on the transformation parameters, it is determined whether to output the detection location information of the current frame image in the offline map; the detection location information of the current frame image is determined by the location information of the matching feature information of the current frame image in the offline map.

[0009] Therefore, in this embodiment, the correspondence between the current coordinate system and the offline map coordinate system is determined based on the feature information of the current frame image and the position information of the feature information corresponding to each historical frame image in the current coordinate system, and the position information of the matching feature information corresponding to the feature information of the current frame image and the historical frame image in the coordinate system of the offline map. By determining the correspondence between the current coordinate system and the offline map coordinate system through the feature information corresponding to multiple frames, the accuracy of robot visual positioning is improved.

[0010] In some embodiments, the visual positioning method further includes:

[0011] If the current frame image is the first frame image, then the transformation parameters between the current coordinate system and the offline map coordinate system are determined based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the coordinate system of the offline map.

[0012] Therefore, in this embodiment, when the current frame image is the first frame image, the correspondence between the current coordinate system and the coordinate system of the offline map is determined by the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the coordinate system of the offline map.

[0013] In some embodiments, the visual positioning method further includes:

[0014] Based on the transformation parameters, the position information of the matching feature information in the coordinate system of the offline map is transformed to obtain the predicted position information of the matching feature information in the current coordinate system;

[0015] The transformation parameters are adjusted based on the error between the predicted position information of the matched feature information in the current coordinate system and the position information of the corresponding feature information in the current coordinate system.

[0016] Therefore, in this embodiment, the transformation parameters are optimized and their accuracy is improved by using the error between the position information of the feature information corresponding to the current frame image and the historical frame image in the current coordinate system and the predicted position information in the current coordinate system based on the transformation parameters.

[0017] In some embodiments, the feature information includes feature point information;

[0018] Based on the feature information of the acquired current frame image, determine the matching feature information of the current frame image in the offline map, including:

[0019] Feature extraction is performed on the acquired current frame image to obtain the feature point information of the current frame image;

[0020] The feature point information of the current frame image is compared with the preset key point information on the offline map;

[0021] If the similarity between the feature point information of the current frame image and the preset key point information on the offline map exceeds a similarity threshold, then the preset key point information corresponding to the similarity is determined as the matching feature information of the feature point information of the current frame image on the offline map.

[0022] Therefore, in this embodiment, by using the similarity between feature point information and preset key point information, matching feature information corresponding to the feature information of the current frame image is determined, thereby improving the positioning accuracy of the current frame image.

[0023] In some embodiments, based on the feature information of the acquired current frame image, matching feature information of the current frame image in the offline map is determined, and then the method further includes:

[0024] The preset key point information and feature point information corresponding to similarity exceeding the similarity threshold are combined into feature point pairs;

[0025] The feature point pairs corresponding to the current frame image are filtered based on the noise reduction algorithm.

[0026] Therefore, in this embodiment, the feature point pairs are filtered by a noise reduction algorithm to improve the accuracy of the feature point pairs corresponding to the current frame image, thereby improving the positioning accuracy of the current frame image.

[0027] In some embodiments, the visual positioning method further includes:

[0028] If the number of feature point pairs in the current frame exceeds a preset number, the current frame will be archived in the historical trajectory image set.

[0029] Therefore, in this embodiment, determining whether the current frame image can be archived into the historical trajectory image set based on the number of feature point pairs in the current frame image can improve the robot's positioning accuracy.

[0030] In some embodiments, determining whether to output the detection location information of the current frame image in the offline map based on transformation parameters includes:

[0031] Based on the current frame image, historical frame images, and their corresponding feature information and transformation parameters, determine whether to output the detection location information of the current frame image in the offline map.

[0032] In some embodiments, determining whether to output the detection location information of the current frame image in the offline map based on the current frame image, historical frame images and their corresponding feature information and transformation parameters includes:

[0033] Based on the current frame image, historical frame images, and their corresponding feature information and transformation parameters, determine whether to output the detection location information of the current frame image in the offline map, including:

[0034] The covariance matrix of the feature information is determined by the current frame image, historical frame images and their corresponding feature information and transformation parameters.

[0035] Based on the estimated value of the covariance matrix, determine whether to output the detection location information of the current frame image in the offline map.

[0036] In some embodiments, the covariance matrix of the feature information is determined using the current frame image, historical frame images, and their corresponding feature information and transformation parameters, including:

[0037] Based on the transformation parameters, the position information of the current frame image and the historical frame images in the current coordinate system are converted into position information in the coordinate system of the offline map, respectively.

[0038] Based on the location information of the current frame image and historical frame images in the coordinate system of the offline map, the key frame distribution information is determined;

[0039] Based on the spatial distribution of each matching feature information corresponding to the feature point information of the current frame image and the feature point information of the historical frame images in the coordinate system of the offline map, the feature point distribution information is determined.

[0040] Based on the keyframe distribution information corresponding to the current frame image and historical frame images, the feature point information of the current frame image and the feature point information of historical frame images, the covariance matrix of the feature information is determined.

[0041] In some embodiments, based on the estimated value of the covariance matrix, determining whether to output the detection location information of the current frame image in the offline map includes:

[0042] The covariance matrix of the feature information is calculated to obtain an estimated value of the covariance matrix;

[0043] If the estimated value of the covariance matrix is ​​less than the preset estimate, the detection location information of the current frame image in the offline map is output.

[0044] Therefore, in this embodiment, the covariance matrix of the feature information is determined based on the current frame image, the historical frame images and their corresponding feature information, and the quality of the positioning results of the current frame image and the historical frame images is evaluated based on the covariance matrix to improve the reliability of the positioning results.

[0045] A second aspect of this application provides a visual positioning device, which includes:

[0046] The feature matching module is used to determine the matching feature information of the current frame image in the offline map based on the feature information of the current frame image.

[0047] The processing module is used to determine the transformation parameters between the current coordinate system and the coordinate system of the offline map based on the position information of the feature information of the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, as well as the position information of the feature information of the historical frame images before the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map.

[0048] The analysis module is used to determine whether to output the detection location information of the current frame image in the offline map based on the transformation parameters; the detection location information of the current frame image is determined by the location information of the matching feature information of the current frame image in the offline map.

[0049] Therefore, in this embodiment, the correspondence between the current coordinate system and the offline map coordinate system is determined based on the feature information of the current frame image and the position information of the feature information corresponding to each historical frame image in the current coordinate system, and the position information of the matching feature information corresponding to the feature information of the current frame image and the historical frame image in the coordinate system of the offline map. By determining the correspondence between the current coordinate system and the offline map coordinate system through the feature information corresponding to multiple frames, the accuracy of robot visual positioning is improved.

[0050] A third aspect of this application provides an electronic device including a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the visual positioning method of the first aspect described above.

[0051] The fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the visual positioning method of the first aspect described above.

[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0054] Figure 1 This is a flowchart illustrating the visual positioning method provided in this application;

[0055] Figure 2 This is a flowchart illustrating an embodiment of the visual positioning method provided in this application;

[0056] Figure 3 This is a schematic diagram of a specific embodiment of the visual positioning method provided in this application;

[0057] Figure 4 This is a schematic diagram of the framework of an embodiment of the visual positioning device provided in this application;

[0058] Figure 5 This is a schematic diagram of the framework of an embodiment of the electronic device of this application;

[0059] Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0060] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0061] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0062] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0063] Please see Figure 1 , Figure 1 This is a flowchart illustrating the visual positioning method provided in this application. This embodiment provides a visual positioning method, the executing entity of which can be an image processing device. This image processing device can be any terminal device, server, or other processing device capable of executing the method embodiment of this application. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, wearable device, etc. The visual positioning method provided in this embodiment is applicable to robots, drones, etc. Specifically, the visual positioning method may include the following steps:

[0064] Step S11: Based on the feature information of the acquired current frame image, determine the matching feature information of the current frame image in the offline map.

[0065] In one embodiment, the current frame image is acquired through image acquisition settings, and feature points are extracted from the current frame image to obtain the feature point information of the current frame image. In this embodiment, when performing visual localization on the robot, the current frame image is acquired by the robot, and then the position of the current frame image is determined.

[0066] In one embodiment, the feature information includes feature point information. Feature extraction is performed on the acquired current frame image to obtain the feature point information of the current frame image; the feature point information of the current frame image is compared with preset key point information on an offline map; in response to the similarity between the feature point information of the current frame image and the preset key point information on the offline map exceeding a similarity threshold, the preset key point information corresponding to the similarity is determined as the matching feature information of the feature point information of the current frame image on the offline map. The preset key point information and feature point information corresponding to the similarity exceeding the similarity threshold are combined to form feature point pairs; the feature point pairs corresponding to the current frame image are filtered based on a denoising algorithm, for example, the denoising algorithm can be the RANSAC (RANdom SAmple Consensus) algorithm.

[0067] Step S12: Based on the position information of the feature information of the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, and the position information of the feature information of the historical frame images before the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, determine the transformation parameters between the coordinate system of the current coordinate system and the coordinate system of the offline map.

[0068] In one embodiment, based on transformation parameters, the position information of the matching feature information in the coordinate system of the offline map is transformed to obtain the predicted position information of the matching feature information in the current coordinate system; based on the error between the predicted position information of the matching feature information in the current coordinate system and the position information of the corresponding feature information in the current coordinate system, the transformation parameters are adjusted.

[0069] Step S13: Based on the transformation parameters, determine whether to output the detection location information of the current frame image in the offline map; the detection location information of the current frame image is determined by the location information of the matching feature information of the current frame image in the offline map.

[0070] In one embodiment, based on the current frame image, historical frame images and their corresponding feature information and transformation parameters, it is determined whether to output the detection location information of the current frame image in the offline map.

[0071] Specifically, the covariance matrix of the feature information is determined by using the current frame image, historical frame images and their corresponding feature information and transformation parameters; based on the estimated value of the covariance matrix, it is determined whether to output the detection location information of the current frame image in the offline map.

[0072] In one specific embodiment, based on transformation parameters, the position information of the current frame image and historical frame images in the current coordinate system are converted into position information in the coordinate system of the offline map. Keyframe distribution information is determined based on the position information of the current frame image and historical frame images in the coordinate system of the offline map. Feature point distribution information is determined based on the spatial distribution of each matching feature information corresponding to the feature point information of the current frame image and the feature point information of the historical frame images in the coordinate system of the offline map. The covariance matrix of the feature information is determined based on the keyframe distribution information corresponding to the current frame image and the historical frame images, as well as the feature point distribution information corresponding to the feature point information of the current frame image and the feature point information of the historical frame images. The covariance matrix of the feature information is calculated to obtain an estimated value. If the estimated value of the covariance matrix is ​​less than a preset estimated value, the detection position information of the current frame image in the offline map is output.

[0073] In another embodiment, in response to the current frame image being the first frame image, the transformation parameters between the current coordinate system and the coordinate system of the offline map are determined based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the coordinate system of the offline map.

[0074] The visual positioning method provided in this embodiment determines the correspondence between the current coordinate system and the offline map coordinate system based on the feature information of the current frame image and the position information of the feature information corresponding to each historical frame image in the current coordinate system, and the position information of the matching feature information corresponding to the feature information of the current frame image and the historical frame image in the coordinate system of the offline map. By determining the correspondence between the current coordinate system and the offline map coordinate system through the feature information corresponding to multiple frames, the accuracy of robot visual positioning is improved.

[0075] Please see Figure 2 and Figure 3 , Figure 2 This is a flowchart illustrating an embodiment of the visual positioning method provided in this application; Figure 3This is a schematic diagram of a specific embodiment of the visual positioning method provided in this application. This embodiment provides a visual positioning method, the executing entity of which can be an image processing device. This image processing device can be any terminal device, server, or other processing device capable of executing the method embodiment of this application. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, wearable device, etc. Specifically, the visual positioning method may include the following steps:

[0076] Step S201: Extract features from the acquired current frame image to obtain the feature information of the current frame image.

[0077] In one embodiment, the current frame image is acquired through image acquisition settings, and feature points are extracted from the current frame image to obtain its feature point information. In this embodiment, when performing visual localization on the robot, the current frame image is acquired through a camera mounted on the robot, and then the position of the current frame image is detected. In other embodiments, feature map extraction can also be performed on the current frame image to obtain its feature information. In a specific embodiment, feature points of the current frame image are obtained by performing ORB (Oriented Fast and Rotated BRIEF) feature point detection on the acquired current frame image. ORB uses the FAST (features from accelerated segment test) algorithm to detect the feature points of the current frame image.

[0078] Step S202: Compare the feature point information of the current frame image with the preset key point information on the offline map.

[0079] Specifically, to determine the corresponding location of the current frame image on the offline map, it is necessary to compare the feature points of the current frame image with all preset key points marked on the offline map. A similarity comparison is performed between each feature point of the current frame image and the corresponding preset key points on the offline map to determine the preset key points that match the feature points of the current frame image.

[0080] Step S203: In response to the similarity between the feature point information of the current frame image and the preset key point information on the offline map exceeding the similarity threshold, the preset key point information corresponding to the similarity is determined as the matching feature information of the feature point information of the current frame image on the offline map.

[0081] Specifically, if the similarity between a feature point in the current frame image and a preset key point on the offline map exceeds a similarity threshold, then the feature point corresponding to the similarity is determined to match the preset key point, and the preset key point is used as the matching feature point of the current frame image on the offline map. The similarity threshold can be set according to the actual situation. For example, the similarity threshold can be 95% or 99%.

[0082] In one embodiment, a coarse matching is performed between feature points of the current frame image and preset key points on the offline map to obtain candidate feature points. Specifically, for the current frame image and the offline map, the Hamming distance between the binarized gradient feature vectors of the feature points is used as a metric for determining feature point similarity. If the Hamming distance between a feature point of the current frame image and the feature vector of a preset key point in the offline map is less than a preset distance, then the feature point of the current frame image is determined to match the preset key point in the offline map. In other alternative embodiments, the similarity can also be evaluated using methods such as Euclidean distance and cosine similarity between the feature points of the current frame image and the preset key points in the offline map.

[0083] Step S204: Combine the preset key point information and feature point information corresponding to similarity exceeding the similarity threshold into feature point pairs.

[0084] Specifically, to improve the accuracy of robot localization, each feature point in the current frame image is combined with a preset key point on the offline map to obtain candidate feature point pairs corresponding to the current frame image and the offline map. Through step S203, a coarse matching can be performed between the feature points of the current frame image and the preset key points on the offline map to obtain candidate feature points.

[0085] Step S205: Filter the feature point pairs corresponding to the current frame image based on the noise reduction algorithm.

[0086] Specifically, RANSAC is used to remove mismatched feature point pairs from the candidate feature points to obtain precisely matched feature point pairs. In other optional embodiments, feature point pairs corresponding to the current frame image can also be filtered in other ways to obtain precisely matched feature point pairs, thereby improving the localization accuracy of the current frame image.

[0087] Step S206: In response to the number of feature point pairs in the current frame image exceeding a preset number, determine to archive the current frame image into the historical trajectory image set.

[0088] Specifically, the number of precisely matched feature point pairs corresponding to the current frame image is determined, and it is then determined whether the number of precisely matched feature point pairs corresponding to the current frame image exceeds a preset number. If the number of precisely matched feature point pairs corresponding to the current frame image exceeds the preset number, the current frame image is archived in the historical trajectory image set.

[0089] Step S207: Determine whether the current frame image is the first frame image.

[0090] In one embodiment, a current coordinate system is established using any feature point in the first frame of the robot's current trajectory as the origin. The coordinate positions of each feature point in the current frame are then determined within this coordinate system. Alternatively, the current coordinate system can be established using other locations as the origin, and the coordinate positions of each feature point in the current frame can be determined within this coordinate system.

[0091] Offline maps have a corresponding coordinate system, in which the coordinates of each preset key point can be determined.

[0092] Since the current coordinate system differs from the offline map's coordinate system, it is necessary to determine the correspondence between them. When the current frame image has historical frame images, the correspondence between the current coordinate system and the offline map's coordinate system can be determined by combining the feature point information of the current frame image and the corresponding feature point information of the historical frame images, thereby improving the accuracy of the correspondence between the two coordinate systems.

[0093] If the current frame image is the first frame image, then the current frame image does not have historical frame images, and the process jumps directly to step S208; if the current frame image is not the first frame image, then the current frame image has historical frame images, and the process jumps directly to step S209.

[0094] Step S208: Based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the coordinate system of the offline map, determine the transformation parameters between the current coordinate system and the coordinate system of the offline map.

[0095] Specifically, when the current frame image is the first frame image, the coordinate positions of each feature point in the current frame image in the camera coordinate system are P. c The coordinates of each feature point in the offline coordinate system are denoted as P. w1 The coordinates P of the same feature point in the current frame image in the camera coordinate system are used to determine its position. c and its coordinate position P in the offline coordinate system w1 Get the coordinates T of the current frame image in the offline map. cw1 Given that the coordinates of the current frame image in this runtime, in a coordinate system with the origin at the starting point, are T... cw2 .

[0096] It can be determined based on the coordinate position T of the current frame image in the current coordinate system. cw2 And the coordinate position T of the current frame image in the coordinate system of the offline map. cw1Determine the transformation parameter dT between the coordinates of the feature point in the current coordinate system and the coordinates of the matched feature point in the coordinate system of the offline map, where,

[0097] Step S209: Based on the position information of the feature information of the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, and the position information of the feature information of the historical frame images before the current frame image in the current coordinate system and the position information of its corresponding matching feature information in the offline map, determine the transformation parameters between the coordinate system of the current coordinate system and the coordinate system of the offline map.

[0098] Specifically, when the current frame image is not the first frame image, the coordinate positions of each feature point in the current frame image in the camera coordinate system are P. c The coordinates of each feature point in the offline coordinate system are denoted as P. w1 The coordinates P of the same feature point in the current frame image in the camera coordinate system are used to determine its position. c and its coordinate position P in the offline coordinate system w1 Get the coordinates T of the current frame image in the offline map. cw1 Given that the coordinates of the current frame image in this runtime, in a coordinate system with the origin at the starting point, are T... cw2 .

[0099] It can be determined based on the coordinate position T of the current frame image in the current coordinate system. cw2 and its coordinate position T in the coordinate system of the offline map relative to the current frame image. cw1 The coordinate positions T of a preset number of historical image frames preceding the current frame in the current coordinate system. cw2 And their respective coordinate positions T in the coordinate system of the offline map. cw1 Determine the transformation parameter dT between the coordinates of the feature point in the current coordinate system and the coordinates of the matched feature point in the coordinate system of the offline map, where,

[0100] In this embodiment, combining multiple frames of images to determine the transformation parameters between the current coordinate system and the offline map coordinate system can improve the accuracy of the transformation parameters and facilitate the improvement of the accuracy of the positioning results.

[0101] In this embodiment, the current frame image and the historical frame images preceding the current frame image share a single transformation parameter.

[0102] Step S210: Adjust the transformation parameters based on the error between the predicted position information of the matching feature information in the current coordinate system and the position information of the corresponding feature information in the current coordinate system.

[0103] Specifically, to further improve the accuracy of the transformation parameters between the current coordinate system and the offline map's coordinate system, the coordinate positions of the matching feature information corresponding to the feature points in the offline map's coordinate system are transformed using the transformation parameters to obtain the predicted coordinate positions of the matching feature information in the current coordinate system. Based on the difference between the coordinate positions of each feature point in the current coordinate system and the predicted coordinate positions, the transformation parameters are continuously optimized until the difference between the coordinate positions of each feature point in the current coordinate system and the predicted coordinate positions is less than a preset value. For example, the preset value can be set to 3 pixels, which can be customized according to actual needs.

[0104] Based on the transformation parameters, the position coordinates of the current frame image in the current coordinate system are converted to coordinates in the coordinate system of the offline map, thereby determining the localization result of the current frame image. The localization result of the current frame image is then archived into the historical trajectory image set and associated with the current frame image.

[0105] Step S211: Determine the covariance matrix of the feature information using the current frame image, historical frame images and their corresponding feature information and transformation parameters.

[0106] In one specific embodiment, based on transformation parameters, the coordinate positions of the current frame image and the historical frame images in the current coordinate system are converted into coordinate positions in the coordinate system of the offline map, respectively; based on the coordinate positions of the current frame image and the historical frame images in the coordinate system of the offline map, the distribution information of the current image frame and all historical image frames before the current image frame is determined.

[0107] Based on the spatial distribution of each matching feature information corresponding to the feature point information of the current frame image and the feature point information of the historical frame image in the coordinate system of the offline map, the feature point distribution information is determined; based on the distribution information of the current frame image and the historical frame image, and the feature point distribution information corresponding to the feature point information of the current frame image and the historical frame image, the covariance matrix of the feature information is determined.

[0108] Step S212: Based on the estimated value of the covariance matrix, determine whether to output the detection location information of the current frame image in the offline map.

[0109] Specifically, the covariance matrix of the feature information is calculated to obtain an estimated value of the covariance matrix. In one embodiment, the magnitude of the covariance matrix in the principal direction is used as the criterion for evaluating the uncertainty of the localization result of the current frame image.

[0110] If the estimated value of the covariance matrix is ​​less than a preset estimate, the detection location information of the current frame image in the offline map is output. In other words, if the magnitude of the covariance matrix in the principal direction does not exceed the set value, the localization result of the current frame image is determined to be reliable, and the localization result of the current frame image can be output. If the magnitude of the covariance matrix in the principal direction exceeds the set value, the localization result of the current frame image is determined to be unreliable, and the localization result of the current frame image can be modified until the localization result of the current frame image is determined to be reliable.

[0111] In this embodiment, by using the estimated value of the covariance matrix to evaluate the positioning results, erroneous and inaccurate positioning results can be filtered out.

[0112] The visual positioning method provided in this embodiment determines the correspondence between the current coordinate system and the offline map coordinate system based on the position information of the feature information of the current frame image and the feature information corresponding to each historical frame image in the current coordinate system, and the position information of the matching feature information corresponding to the feature information of the current frame image and the historical frame images in the coordinate system of the offline map. By determining the correspondence between the current coordinate system and the offline map coordinate system through the feature information corresponding to multiple frames, the positioning result is more accurate and stable, less affected by mismatches and mapping accuracy, and improves the accuracy of robot visual positioning. This method enables the robot to retrieve the map of the explored environment and unify the multiple established coordinate systems into the coordinate system of the offline map, thereby supporting the realization of multi-person AR.

[0113] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0114] Please see Figure 4 , Figure 4This is a schematic diagram of the framework of an embodiment of the visual positioning device of this application. The visual positioning device 60 includes a feature matching module 61, a processing module 62, and an analysis module 63. The feature matching module 61 is used to determine the matching feature information of the current frame image in the offline map based on the feature information of the acquired current frame image; the processing module 62 is used to determine the transformation parameters between the coordinate system of the current coordinate system and the coordinate system of the offline map based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the offline map, and the position information of the feature information of the historical frame images before the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the offline map; the analysis module 63 is used to determine whether to output the detection position information of the current frame image in the offline map based on the transformation parameters; the detection position information of the current frame image is determined by the position information of the matching feature information of the current frame image in the offline map.

[0115] In the visual positioning device provided in this embodiment, the correspondence between the current coordinate system and the offline map coordinate system is determined based on the feature information of the current frame image and the position information of the feature information corresponding to each historical frame image in the current coordinate system, and the position information of the matching feature information corresponding to the feature information of the current frame image and the historical frame image in the coordinate system of the offline map. By determining the correspondence between the current coordinate system and the offline map coordinate system through the feature information corresponding to multiple frames, the accuracy of robot visual positioning is improved.

[0116] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the electronic device of this application. The electronic device 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described visual positioning method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to, a microcomputer, a server, etc. In addition, the electronic device 80 may also include mobile devices such as laptops and tablets, which are not limited here.

[0117] Specifically, processor 82 controls itself and memory 81 to implement the steps in any of the above-described visual positioning method embodiments. Processor 82 can also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 82 can be implemented using integrated circuit chips.

[0118] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 90 stores program instructions 901 that can be executed by a processor. The program instructions 901 are used to implement the steps of any of the above-described visual positioning method embodiments.

[0119] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0120] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0121] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0122] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0123] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0124] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A visual positioning method, characterized in that, include: Based on the feature information of the acquired current frame image, determine the matching feature information of the current frame image in the offline map; Based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the offline map, and the position information of the feature information of the historical frame images preceding the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the offline map, a transformation parameter between the coordinate system of the current coordinate system and the coordinate system of the offline map is determined; based on the transformation parameter, it is determined whether to output the detection position information of the current frame image in the offline map; the detection position information of the current frame image is determined by the position information of the matching feature information of the feature information of the current frame image in the offline map. The visual positioning method further includes: Based on the transformation parameters, the position information of the matching feature information in the coordinate system of the offline map is transformed to obtain the predicted position information of the matching feature information in the current coordinate system; based on the error between the predicted position information of the matching feature information in the current coordinate system and the position information of the corresponding feature information in the current coordinate system, the transformation parameters are adjusted.

2. The visual positioning method according to claim 1, characterized in that, Also includes: In response to the fact that the current frame image is the first frame image, the transformation parameters between the current coordinate system and the coordinate system of the offline map are determined based on the position information of the feature information of the current frame image in the current coordinate system and the position information of the corresponding matching feature information in the coordinate system of the offline map.

3. The visual positioning method according to claim 1, characterized in that, The feature information includes feature point information; determining the matching feature information of the current frame image in the offline map based on the acquired feature information of the current frame image includes: Feature extraction is performed on the acquired current frame image to obtain feature point information of the current frame image; the feature point information of the current frame image is compared with preset key point information on the offline map; in response to the similarity between the feature point information of the current frame image and the preset key point information on the offline map exceeding a similarity threshold, the preset key point information corresponding to the similarity is determined to be the matching feature information of the feature point information of the current frame image on the offline map.

4. The visual positioning method according to claim 3, characterized in that, The step of determining the matching feature information of the current frame image in the offline map based on the acquired feature information of the current frame image further includes: The preset key point information and feature point information corresponding to the similarity exceeding the similarity threshold are combined to form feature point pairs; the feature point pairs corresponding to the current frame image are filtered based on the noise reduction algorithm.

5. The visual positioning method according to claim 3 or 4, characterized in that, The visual positioning method further includes: If the number of feature point pairs in the current frame image exceeds a preset number, then the current frame image is determined to be archived in the historical trajectory image set.

6. The visual positioning method according to claim 3, characterized in that, The step of determining whether to output the detection location information of the current frame image in the offline map based on the transformation parameters includes: determining whether to output the detection location information of the current frame image in the offline map based on the current frame image, the historical frame images and their corresponding feature information and the transformation parameters.

7. The visual positioning method according to claim 6, characterized in that, The step of determining whether to output the detection location information of the current frame image in the offline map based on the current frame image, the historical frame images and their corresponding feature information and transformation parameters includes: The covariance matrix of the feature information is determined by the current frame image, the historical frame images and their corresponding feature information and transformation parameters; based on the estimated value of the covariance matrix, it is determined whether to output the detection location information of the current frame image in the offline map.

8. The visual positioning method according to claim 7, characterized in that, The step of determining the covariance matrix of the feature information using the current frame image, the historical frame images and their corresponding feature information, and the transformation parameters includes: converting the position information of the current frame image and the historical frame image in the current coordinate system to position information in the coordinate system of the offline map based on the transformation parameters; determining keyframe distribution information based on the position information of the current frame image and the historical frame image in the coordinate system of the offline map; determining feature point distribution information based on the spatial distribution of each matching feature information corresponding to the feature point information of the current frame image and the feature point information of the historical frame image in the coordinate system of the offline map; and determining the covariance matrix of the feature information based on the keyframe distribution information corresponding to the current frame image and the historical frame image, and the feature point distribution information corresponding to the feature point information of the current frame image and the feature point information of the historical frame image.

9. The visual positioning method according to claim 7, characterized in that, The step of determining whether to output the detection location information of the current frame image in the offline map based on the estimated value of the covariance matrix includes: calculating the covariance matrix of the feature information to obtain an estimated value of the covariance matrix; and outputting the detection location information of the current frame image in the offline map in response to the estimated value of the covariance matrix being less than a preset estimated value.

10. A visual positioning device, characterized in that, The visual positioning device includes: a feature matching module, used to determine matching feature information of the current frame image in an offline map based on the feature information of the acquired current frame image; a processing module, used to determine transformation parameters between the current coordinate system and the coordinate system of the offline map based on the position information of the feature information of the current frame image in the current coordinate system and the corresponding position information of the matching feature information in the offline map, and the position information of the feature information of historical frame images preceding the current frame image in the current coordinate system and the corresponding position information of the matching feature information in the offline map; and an analysis module, used to determine whether to output the detection position information of the current frame image in the offline map based on the transformation parameters; the detection position information of the current frame image is determined by the position information of the matching feature information of the current frame image in the offline map. The visual positioning device further includes: Based on the transformation parameters, the position information of the matching feature information in the coordinate system of the offline map is transformed to obtain the predicted position information of the matching feature information in the current coordinate system; based on the error between the predicted position information of the matching feature information in the current coordinate system and the position information of the corresponding feature information in the current coordinate system, the transformation parameters are adjusted.

11. An electronic device, characterized in that, It includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the visual positioning method according to any one of claims 1 to 9.

12. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the visual positioning method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Visual positioning method and device based on visual map

    CN111780764A

  • Method for detecting obstacle, electronic equipment, roadside equipment and cloud control platform

    CN112560769A