Visual positioning method and device, electronic equipment and medium
By using extended reality equipment and preset image sets in visual positioning technology for feature point matching and transformation matrix fine-tuning, the problems of environmental interference and low positioning accuracy in the prior art are solved, higher robustness and positioning success rate are achieved, and an effective evaluation method for fixed position reliability is provided.
Patent Information
- Application Number
- CN202510205098.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-24
AI Technical Summary
Existing visual positioning technologies show poor robustness, low positioning accuracy and low success rate in areas with environmental interference, poor texture, or areas with insufficient mapping coverage, and lack effective method for evaluating fixed position reliability.
By extending the real-life equipment, the feature point matching is performed by combining the scene images in the preset image set, the two-dimensional and three-dimensional coordinates of the target feature point are determined, and the conversion matrix is fine-tuned by the optimization algorithm to evaluate the availability of the conversion matrix.
It improves the robustness of visual positioning, reduces the impact of environmental interference, significantly improves the positioning success rate of areas with poor texture or areas with insufficient mapping coverage, and provides an effective method of fixed position reliability evaluation.
Smart Images

Figure CN120198501A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision technology, and in particular to a visual positioning method, device, electronic equipment and medium. Background Art
[0002] High-precision positioning is to use advanced technical means to accurately determine the position of the target object in space. In the field of machine vision, high-precision positioning is one of the core functions. By using computer vision, image processing and sensor technology, the real-time tracking and positioning of the target object can be achieved by analyzing the image data captured by the camera. This technology can be applied to a variety of scenarios, such as robot grasping in industrial automation, road recognition for driverless cars, and surgical assistance in the medical field.
[0003] High-precision positioning is a key technology for multi-person large spaces. Existing solutions include laser positioning, UWB positioning, multi-sensor fusion positioning, and visual positioning. Laser positioning has high positioning accuracy, but it requires additional base stations to be set up and terminal locators to be installed, and the terminal locators cannot be blocked. The cost is relatively high and it is easily affected by environmental interference. The positioning accuracy of UWB positioning in large space scenes cannot reach the centimeter level, and it also requires the establishment of base stations and the installation of terminal locators. The cost is high, the positioning accuracy is insufficient, and it is easily affected by environmental interference. Multi-sensor fusion solutions often combine multi-source data such as images, IMUs, or lasers to achieve high-precision positioning. It has the advantages of good robustness and high accuracy, but it requires the configuration of multiple sensors, which has the disadvantages of high cost and difficulty in algorithm implementation. Visual positioning uses the camera that comes with XR and combines it with a pre-reconstructed positioning map to achieve high-precision positioning. It has the advantages of low cost and high accuracy, but there are the following problems: 1) Poor robustness, and positioning accuracy is easily affected by the environment. 1) Traditional visual positioning methods use artificially designed feature extraction methods such as ORB and SIFT, which are easily affected by environmental changes such as ambient light brightness and human occlusion; 2) In areas with poor texture or insufficient map coverage, insufficient feature matching is prone to occur, resulting in a low positioning success rate; 3) There is a lack of automatic evaluation methods for positioning reliability, which makes it difficult to determine whether the positioning results are usable. Summary of the invention
[0004] The embodiments of the present application provide a visual positioning method, device, electronic device and medium to reduce the degree of environmental interference and improve the accuracy and success rate of visual positioning.
[0005] According to one aspect of the present application, a visual positioning method is provided, the method comprising:
[0006] Capturing an image of a target scene through an extended reality device to obtain a positioning image, matching feature points of the positioning image with scene images in a preset image set, and determining target feature points;
[0007] Determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point according to the point cloud data in the target scene;
[0008] Perform positioning calculation based on the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determine the transformation matrix corresponding to the positioning image, and perform fine-tuning calculation on the transformation matrix based on an optimization algorithm;
[0009] Determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine-tuning parameters in the fine-tuning calculation process, and determine the available evaluation result of the transformation matrix according to the evaluation score.
[0010] According to one aspect of the present application, there is provided a visual positioning device, the device includes:
[0011] A target feature point determination module, configured to collect an image of a target scene through an augmented reality device to obtain a positioning image, perform feature point matching on the positioning image and a scene image in a preset image set, and determine a target feature point;
[0012] A coordinate determination module, configured to determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point according to the point cloud data in the target scene;
[0013] A transformation matrix determination module, configured to perform positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determine the transformation matrix corresponding to the positioning image, and perform fine-tuning calculation on the transformation matrix based on an optimization algorithm;
[0014] A transformation matrix evaluation module, configured to determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine-tuning parameters in the fine-tuning calculation process, and determine the available evaluation result of the transformation matrix according to the evaluation score.
[0015] According to another aspect of the present application, there is provided an electronic device, the electronic device includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to at least one processor; wherein,
[0018] The memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor so that at least one processor can execute the visual positioning method of any embodiment of the present application.
[0019] According to another aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the visual positioning method of any embodiment of the present application when executed.
[0020] In the technical solution of the embodiment of the present application, an extended reality device is used to collect an image of a target scene to obtain a positioning image, and feature point matching is performed between the positioning image and the scene images in a preset image set to determine target feature points, which can improve robustness, is not easily affected by the environment, and can greatly increase the number of image feature matches, and can significantly improve the positioning success rate in areas with insufficient texture or insufficient mapping coverage. Determine the two-dimensional coordinates of the target feature points in the positioning image, and determine the three-dimensional coordinates of the target feature points according to the point cloud data in the target scene; perform positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature points to determine the transformation matrix corresponding to the positioning image, and perform fine adjustment calculation on the transformation matrix based on an optimization algorithm; determine an evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine adjustment parameters in the fine adjustment calculation process, and determine an available evaluation result of the transformation matrix according to the evaluation score, so as to effectively evaluate the transformation matrix and select an ideal and effective transformation matrix for use.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a visual positioning method provided by an embodiment of the present application;
[0024] Figure 2 It is a flowchart of a visual positioning method provided by another embodiment of the present application;
[0025] Figure 3 It is a flowchart of a visual positioning method provided by yet another embodiment of the present application;
[0026] Figure 4Flowchart of the specific implementation provided by the embodiments of the present application;
[0027] Figure 5 Schematic structural diagram of a visual positioning device provided by the embodiments of the present application;
[0028] Figure 6 Schematic structural diagram of an electronic device provided by the embodiments of the present application. Specific implementation manners
[0029] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Figure 1 Flowchart of a visual positioning method provided by the embodiments of the present application. The embodiments of the present application are applicable to the case of visual positioning based on the positioning images collected by an extended reality device. This method can be executed by a visual positioning device, which can be implemented in the form of hardware and / or software, and the visual positioning device can be configured in an electronic device. As Figure 1 shown, this method includes:
[0032] S110. Collect a positioning image of a target scene through an extended reality device, match feature points of the positioning image with scene images in a preset image set, and determine target feature points.
[0033] Among them, the extended reality device, i.e., the XR device, is a hardware device integrating virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies. In the embodiments of the present application, there may be at least one extended reality device, and the solutions in the embodiments of the present application are executed for each extended reality device. The extended reality device can be driven to move by being worn on the head, held in the hand, or carried in other ways by the user, or carried by a robot to move, and perform scanning image acquisition on the target scene. The target scene can be a pre-set scene, and visual positioning needs to be performed on the target scene. The positioning map corresponding to the target scene is known in advance, which includes point cloud data and a preset image set. The point cloud data is the three-dimensional coordinates of the preset feature points in the target scene, and the preset image set is a set of images obtained by pre-performing image acquisition on the target scene. Each image in the preset image set contains at least one feature point in the target scene, and the set of feature points included in each image in the preset image set contains all the feature points in the target scene, that is, the feature points in the point cloud of the target scene can all find corresponding feature points in the preset image set. The scene image can be selected from the preset images in the preset image set, and can be specifically selected according to the image quality of the preset image, the image acquisition perspective, the characteristics of the feature points included in the image, etc.
[0034] Exemplarily, in the target scene, the extended reality device can be used to perform image acquisition on the target scene to obtain a positioning image. The positioning image is an image obtained by the extended reality device performing image acquisition on the target scene at a position and a posture, which contains the characteristics of the target scene. The feature points of the positioning image can be matched with the scene images in the preset image set to determine the target feature points, that is, to determine the target feature points that appear in both the positioning image and the scene image and belong to the same point in the target scene.
[0035] Specifically, the feature points of the positioning image and the scene image can be matched based on an image matching algorithm. The preset images in the preset image set can be used as the scene images respectively, and the feature points of the positioning image and the scene image can be matched. Or some preset images can be selected from the preset image set as the scene images, and the feature points of the positioning image and the scene image can be matched.
[0036] S120. Determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point according to the point cloud data in the target scene.
[0037] Exemplarily, the two-dimensional coordinates of the target feature point in the positioning image can be determined by recognizing the positioning image, that is, the pixel coordinates in the positioning image. In the process of matching the positioning image with the scene image mentioned above, the target feature point that matches the scene image in the positioning image is determined. The target feature point in the scene image can correspond to the point cloud data. Therefore, the point cloud data of the target feature point, that is, the three-dimensional coordinates of the target feature point, can be determined. For a target feature point, its two-dimensional coordinates and three-dimensional coordinates can be determined.
[0038] S130. Perform positioning calculation based on the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determine the transformation matrix corresponding to the positioning image, and perform fine adjustment on the transformation matrix based on an optimization algorithm.
[0039] Exemplarily, based on the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, positioning calculation can be performed, that is, determine the transformation matrix between the two-dimensional coordinates and the three-dimensional coordinates, which reflects the position and posture of the extended reality device during image acquisition, as the transformation matrix corresponding to the positioning image. The two-dimensional coordinates of each point in the positioning image can be transformed into three-dimensional coordinates based on this transformation matrix.
[0040] After determining the transformation matrix, perform fine adjustment on the transformation matrix based on an optimization algorithm to optimize the transformation matrix, so that the transformation matrix can more accurately reflect the transformation relationship between the two-dimensional coordinates and the three-dimensional coordinates.
[0041] In the embodiment of the present application, the optimization algorithm can adopt the Bundle Adjustment Algorithm (BA algorithm), and optimize the transformation matrix and the three-dimensional coordinates by minimizing the reprojection error.
[0042] S140. Determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine adjustment parameters in the fine adjustment process, and determine the available evaluation result of the transformation matrix according to the evaluation score.
[0043] Among them, the feature point matching process refers to the process of matching the positioning image with the scene image in S110. The matching parameters are the parameters that can reflect the feature point matching effect obtained during the feature point matching process, such as the number, ratio, confidence level, etc. of the feature points that can be successfully matched. The positioning calculation process refers to the process of performing positioning calculation according to the two-dimensional coordinates and three-dimensional coordinates of the target feature points in S130. The calculation parameters can be the parameters that reflect the positioning calculation effect during the positioning calculation process. For example, to reflect the accuracy of the positioning calculation, it can be the reprojection error, confidence level of the transformation matrix obtained during the positioning calculation process, the number and proportion of coordinate pairs of the two-dimensional coordinates and three-dimensional coordinates corresponding to the transformation matrix with the reprojection error less than the preset reprojection error, and the number and proportion of coordinate pairs of the two-dimensional coordinates and three-dimensional coordinates corresponding to the transformation matrix with the confidence level greater than the preset confidence threshold. The calculation fine-tuning process refers to the process of performing fine-tuning on the transformation matrix in S130. The fine-tuning parameters can be the parameters that reflect the effect of the optimized transformation matrix during the calculation fine-tuning. For example, the reprojection error and covariance estimator of the optimized transformation matrix output corresponding during the calculation fine-tuning process.
[0044] In the embodiments of the present application, at least one of the matching parameter, the calculation parameter, and the fine-tuning parameter can reflect the effect of the transformation matrix. The evaluation score of the transformation matrix can be determined according to at least one of the matching parameter, the calculation parameter, and the fine-tuning parameter. The transformation matrix can be evaluated through the evaluation score to determine the available evaluation result of the transformation matrix and determine whether the transformation matrix is available. Specifically, the matching parameter, the calculation parameter, and the fine-tuning parameter are all quantified parameters, and at least one of the matching parameter, the calculation parameter, and the fine-tuning parameter can be operated to obtain the final quantified evaluation score. For example, in the case where there are at least two of the matching parameter, the calculation parameter, and the fine-tuning parameter, the evaluation score can be obtained by performing weighted summation on at least two of them.
[0045] In the technical solution of the embodiment of the present application, a positioning image is obtained by collecting an image of a target scene through an extended reality device, and feature point matching is performed between the positioning image and the scene images in a preset image set to determine target feature points, which can improve robustness, is not easily affected by the environment, and can greatly increase the number of image feature matches, and can significantly improve the positioning success rate in areas with insufficient texture or insufficient mapping coverage. Determine the two-dimensional coordinates of the target feature points in the positioning image, and determine the three-dimensional coordinates of the target feature points according to the point cloud data in the target scene; perform positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature points to determine the transformation matrix corresponding to the positioning image, and perform fine-tuning calculation on the transformation matrix based on an optimization algorithm; determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine-tuning parameters in the fine-tuning calculation process, and determine the available evaluation result of the transformation matrix according to the evaluation score, so as to effectively evaluate the transformation matrix and select an ideal and effective transformation matrix for use.
[0046] Figure 2 The flowchart of a visual positioning method provided in another embodiment of the present application. The embodiment of the present application is optimized based on the above embodiment, and the solutions not described in detail in the embodiment of the present application can be seen in the above embodiment. As Figure 2 shown, the method of the embodiment of the present application specifically includes the following steps:
[0047] S210. Collect an image of a target scene through an extended reality device to obtain a positioning image.
[0048] S220. Extract feature points from the positioning image to determine the feature points in the positioning image, and extract vectors from the positioning image to determine the image vector of the positioning image.
[0049] Exemplarily, a deep learning model can be used to extract feature points from the positioning image to obtain the feature points included in the positioning image. The deep learning model for feature point extraction can be an image feature point extraction model based on deep learning such as SurperPoint, R2D2, D2Net, SoSNet, Disk, etc. A deep learning model can be used to extract vectors from the positioning image to determine the image vector of the positioning image, which reflects the content of the positioning image. The deep learning model for vector extraction can be an image feature vector extraction model based on deep learning such as NetVLad, CosPlace, etc.
[0050] S230. Determine a scene image from the preset image set according to the image vector of the positioning image and the image vectors of the preset images in the preset image set.
[0051] Exemplarily, a scene image can be selected from the preset images in the preset image set according to the image vector. Since the image vector can reflect the content of the image, it is possible to determine whether two images are the same or similar images based on the comparison of the image vectors. The image vector of the positioning image can be compared with the image vectors of the preset images in the preset image set, so as to select a preset image with content similar to that of the positioning image from the preset image set as the scene image.
[0052] S240. Match the feature points in the positioning image with the feature points in the scene image to determine the target feature points.
[0053] Exemplarily, the scene image is an image with a relatively high content similarity to the positioning image. Therefore, the probability that there are matching feature points between the scene image and the positioning image is the highest. The feature points in the positioning image can be matched with the feature points in the scene image to query the target feature points that match the positioning image in the scene image. There may be at least two pairs of matching feature points in the positioning image and the scene image, and both are used as target feature points.
[0054] In the embodiment of the present application, determining the scene image from the preset image set according to the image vector of the positioning image and the image vectors of the preset images in the preset image set includes:
[0055] Compare the image vector of the positioning image with the image vectors of the preset images to compare the similarity, and use the preset number of preset images with the highest similarity as the scene images;
[0056] Correspondingly, matching the feature points in the positioning image with the feature points in the scene image to determine the target feature points includes:
[0057] Traverse each scene image, respectively match the feature points in the scene image with the feature points in the positioning image to match the similarity, and use the feature point pairs with similarity exceeding the preset similarity threshold as candidate feature point pairs;
[0058] Verify the candidate feature point pairs based on the verification algorithm, use the candidate feature point pairs with the first confidence level greater than the first preset confidence level threshold as the target feature points, and record the first number of inliers and the first inlier rate.
[0059] Exemplarily, in the process of determining the scene image, the image vector of the positioning image can be compared with the image vector of the preset image, which is equivalent to comparing the similarity of the image content between the positioning image and the preset image. The preset number of preset images with the highest similarity is used as the scene image, and the preset number can be determined according to the actual situation, such as 5. The similarity can be cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc. For each scene image, traverse the scene image, and respectively match the feature points in the scene image with the feature points in the positioning image to search for feature point pairs with higher similarity. The feature point pairs with similarity exceeding the preset similarity threshold are used as candidate feature point pairs. To further verify the matching of the candidate feature point pairs and more accurately find the matching feature points, the candidate feature point pairs can be verified based on a verification algorithm, and the candidate feature point pairs with the first confidence level greater than the first preset confidence threshold are used as the target feature points, that is, the target feature points that finally determine the same point in the same target scene that matches between the positioning image and the scene image. The first confidence level can be the authenticity that can be correspondingly output based on the verification algorithm to reflect that the candidate feature point pair belongs to the same feature point. The first inlier number can be the number of finally determined target feature points, and the first inlier ratio can be the ratio of the number of finally determined target feature points to the number of candidate feature points. The first inlier ratio and the first inlier number can be reference indicators for evaluating the determined transformation matrix in the subsequent process.
[0060] S250. Determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point according to the point cloud data in the target scene.
[0061] S260. Perform positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point to determine the transformation matrix corresponding to the positioning image, and perform fine adjustment on the transformation matrix based on an optimization algorithm.
[0062] S270. Determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the calculation parameters in the positioning calculation process, and the fine adjustment parameters in the fine adjustment process, and determine the available evaluation result of the transformation matrix according to the evaluation score.
[0063] An embodiment of the present application provides a visual positioning method, which extracts feature points from the positioning image to determine the feature points in the positioning image, and extracts vectors from the positioning image to determine the image vector of the positioning image; according to the image vector of the positioning image and the image vectors of the preset images in the preset image set, a scene image is determined from the preset image set; the feature points in the positioning image are matched with the feature points in the scene image to determine the target feature points. The above solution can effectively reduce the influence of ambient light brightness and task occlusion on the positioning performance through the comparison of image vectors and the similarity comparison of feature points, objectively determine the target feature points through the matching of features, and determine the transformation matrix based on the two-dimensional coordinates and three-dimensional coordinates of the target features.
[0064] Figure 3 The flowchart of a visual positioning method provided by another embodiment of the present application. The embodiment of the present application is optimized based on the above embodiment, and the solutions not described in detail in the embodiment of the present application can be seen in the above embodiment. As Figure 3 shown, the method of the embodiment of the present application specifically includes the following steps:
[0065] S310. Collect an image of the target scene through an extended reality device to obtain a positioning image, and match the feature points of the positioning image with the scene images in the preset image set to determine the target feature points.
[0066] S320. Determine the two-dimensional coordinates of the target feature points in the positioning image, and determine the three-dimensional coordinates of the target feature points according to the point cloud data in the target scene.
[0067] S330. Call a positioning calculation function to perform positioning calculation on the two-dimensional coordinates and the three-dimensional coordinates to determine the transformation matrix corresponding to the positioning image, and correspondingly output a second confidence level.
[0068] Exemplarily, a positioning calculation function can be called to perform positioning calculation on the two-dimensional coordinates and the three-dimensional coordinates. For example, the positioning calculation function can be the cv2.solvePnPRansac function provided by OpenCV. After the positioning calculation, a transformation matrix that can realize the transformation between the two-dimensional coordinates and the three-dimensional coordinates of the target feature points can be obtained, which corresponds to the positioning image, that is, the three-dimensional coordinates of all the feature points in the positioning image can be obtained through the transformation matrix. During the process of calling the positioning calculation matrix to perform positioning calculation on the two-dimensional coordinates and the three-dimensional coordinates, a second confidence level is correspondingly output, which reflects the authenticity of the matching between the two-dimensional coordinates and the three-dimensional coordinates.
[0069] S340. Use the two-dimensional coordinates and the three-dimensional coordinates whose second confidence level is greater than the second preset confidence threshold as target coordinate pairs, and record the second inlier number and the second inlier rate.
[0070] The second confidence level reflects the authenticity of the matching between the two-dimensional coordinates and the three-dimensional coordinates of the target feature points. The higher the second confidence level, the higher the matching degree of the target coordinate pair of the two-dimensional coordinates and the three-dimensional coordinates, and the more accurate the transformation matrix calculated therefrom. The second number of inliers may be the number of target coordinate pairs, and the second inlier rate may be the ratio of the number of target coordinate pairs to the number of coordinate pairs of the two-dimensional coordinates and the three-dimensional coordinates.
[0071] S350. Based on an optimization algorithm, perform resolution fine-tuning on the transformation matrix to obtain an optimized transformation matrix, and correspondingly output the translation covariance and the rotation covariance.
[0072] Exemplarily, the optimization algorithm may be a Bundle Adjustment (BA) algorithm. The transformation matrix may be resolved and fine-tuned based on the BA optimization algorithm to obtain an optimized transformation matrix. This algorithm correspondingly outputs the translation covariance and the transformation covariance under optimized conditions.
[0073] S360. Determine an evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the resolution parameters in the positioning resolution process, and the fine-tuning parameters in the resolution fine-tuning process, and determine an available evaluation result of the transformation matrix according to the evaluation score.
[0074] Exemplarily, determining an evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the resolution parameters in the positioning resolution process, and the fine-tuning parameters in the resolution fine-tuning process may be based on at least one of the first number of inliers, the first inlier rate, the second number of inliers, the second inlier rate, the translation covariance, and the rotation covariance output from the above process to determine the evaluation score of the transformation matrix.
[0075] The embodiment of the present application provides a visual positioning method. By calling a positioning resolution function to perform positioning resolution on two-dimensional coordinates and three-dimensional coordinates, target coordinate pairs are screened according to the second confidence level, the second number of inliers and the second inlier rate are determined, and the translation covariance and the rotation covariance are determined during the resolution fine-tuning process. Based on the various parameters calculated in the above process, the transformation matrix is evaluated, and the evaluation score of the transformation matrix is quantitatively calculated to accurately evaluate the transformation matrix, thereby reflecting the available degree of the transformation matrix and providing a reliability reference for the use of the transformation matrix.
[0076] In the embodiments of the present application, the matching parameters in the feature point matching process include the first inlier number and the first inlier ratio of the target feature point pairs with the first confidence greater than the first confidence threshold selected during the verification of the candidate feature point pairs. The candidate feature point pairs are the feature point pairs with a similarity exceeding a preset similarity threshold among the feature points in the scene image and the feature points in the positioning image; the solution parameters in the positioning solution process include the second inlier number and the second inlier ratio of the two-dimensional coordinates and the three-dimensional coordinates corresponding to the output with the second confidence greater than the second confidence threshold during the positioning solution of the two-dimensional coordinates and the three-dimensional coordinates; the fine-tuning parameters in the solution fine-tuning process include the translation covariance and the rotation covariance corresponding to the output during the solution fine-tuning process of the transformation matrix.
[0077] Correspondingly, according to at least two of the matching parameters in the feature point matching process, the solution parameters in the positioning solution process, and the fine-tuning parameters in the solution fine-tuning process, the evaluation score of the transformation matrix is determined, including:
[0078] Normalize at least two of the first inlier number, the first inlier ratio, the second inlier number, the second inlier ratio, the translation covariance, and the rotation covariance respectively, and perform weighted summation on the normalized values to obtain the evaluation score of the transformation matrix; wherein, the evaluation score of the transformation matrix is positively correlated with the first inlier number, the first inlier ratio, the second inlier number, and the second inlier ratio respectively, and negatively correlated with the translation covariance and the rotation covariance respectively.
[0079] Exemplarily, in the above process, the first inlier number, the first inlier ratio, the second inlier number, the second inlier ratio, the translation covariance, and the rotation covariance are determined, and the evaluation score of the transformation matrix can be calculated according to at least one of the above. Specifically, in the case of determining the evaluation score of the transformation matrix according to at least two of the above, specifically, at least two of the above parameters can be normalized, and the normalized values are weighted and summed to obtain the evaluation score of the transformation matrix.
[0080] Specifically, assume that the first inlier ratio is ratio1, the first inlier number is num1, the second inlier ratio is ratio2, the second inlier number is num2, the translation covariance is sigma_position, and the rotation covariance is sigma_rotation. Among them, if the values of ratio1 and ratio2 are between 0 and 1, no normalization is required. If it needs to be a percentage, it is normalized to between 0 and 1. Truncated normalization operations are performed on num1 and num2. Assume that the upper limit maximum value is MaxNum. The specific formula is as follows:
[0081] norm_num1 = num1 / max(num1, MaxNum);
[0082] norm_num2 = num2 / max(num2, MaxNum);
[0083] The BA solution covariance estimators sigma_position and sigma_rotation can effectively measure the uncertainties of the translational part and rotational part in the positioning result. Therefore, it is a reverse index. The smaller the value, the smaller the uncertainty and the more accurate the positioning result. Therefore, a reverse-phase normalization operation is adopted. Assuming the upper limit maximum values are MaxSigmaPosition and MaxSigmaRotation, the specific formula is as follows:
[0084] norm_sigma_position = 1 - sigma_position / max(sigma_position, MaxSigmaPosition);
[0085] norm_sigma_position = 1 - sigma_rotation / max(sigma_rotation, MaxSigmaRotation);
[0086] The parameters MaxNum, MaxSigmaPosition, and MaxSigmaRotation in the normalization can be selected according to the actual situation. In the embodiment, for example, they are respectively valued 100, 0.001, and 0.0001.
[0087] The evaluation score conf of the transformation matrix = weight11 * norm_num1 + weight12 * norm_num2 + weight21 * ratio1 + weight22 * ratio2 + weight31 * norm_sigma_position + weight32 * norm_sigma_position.
[0088] The parameters in the formula are assigned according to the contribution degree. In the embodiment, weight11 and weight12 can be respectively taken as 0.15; weight21 and weight22 can be respectively taken as 0.15; weight31 and weight32 can be respectively taken as 0.4.
[0089] In the embodiment of the present application, determining the available evaluation result of the transformation matrix according to the evaluation score includes:
[0090] If the evaluation score is greater than the preset evaluation score, it is determined that the transformation matrix is available for positioning based on the rotation matrix;
[0091] If the evaluation score is less than or equal to a preset evaluation score, it is determined that the transformation matrix is unavailable, and the process of re-determining the transformation matrix starts from performing feature point matching between the positioning image and the scene images in the preset image set.
[0092] Exemplarily, a preset evaluation score can be determined in advance. If the evaluation score of the calculated transformation matrix is greater than the preset evaluation score, it is determined that the transformation matrix is available, and positioning is performed based on the transformation matrix. If the evaluation score is less than or equal to the preset evaluation score, it reflects that the transformation matrix is not accurate enough, and it is determined that the transformation matrix is unavailable. Then, the transformation matrix is re-determined based on the above process to obtain a more accurate transformation matrix.
[0093] Figure 4 It is a flowchart of a specific implementation manner provided by the embodiments of the present application. As Figure 5 shown, a specific implementation process includes:
[0094] 1. Camera data acquisition module
[0095] The data includes positioning image data I, the internal parameters K of the camera corresponding to the image, and the distortion parameters D of the camera corresponding to the image.
[0096] 2. Image processing module
[0097] (1) Feature points of the positioning image I are extracted based on the deep learning model A, and the extracted image features are denoted as Ki. The model A uses a deep learning feature point extraction model, which can be a deep learning-based image feature point extraction model such as SurperPoint, R2D2, D2Net, SoSNet, Disk, etc.
[0098] (2) Image vectors of the positioning image I are extracted based on the deep learning model B, and the extracted image vectors are denoted as Vi. The image vector extraction model used by the model B can be a deep learning-based image feature point extraction model such as NetVLad, CosPlace, etc.
[0099] 3. Image query module
[0100] Similarity query is performed on the image vector set Vmap of the image library Q using the image vector Vi. The similarity calculation formula uses cosine similarity, and the TopN image set Iq with a similarity greater than Threshold1 is taken as the final result. Among them, Threshold1 is taken as 0.8, and TopN is taken as 5; if the result greater than Threshold1 is greater than TopN, then take TopN results, and if it is less than TopN, then take the result of Threshold1.
[0101] 4. Multi-image feature matching module
[0102] Implement the feature point matching between the positioning map I and the TopN image set Iq, and perform geometric verification on the feature matching results to obtain a sufficient number and quality of 2D-3D matching degrees.
[0103] (1) Use the model C to perform one-by-one feature point matching on the positioning image I and the TopN query map to obtain the candidate feature point pairs Match1.
[0104] (2) Use the cv2.findEssentialMat or cv2.findHomography function provided by Opencv to perform one-by-one geometric verification on the candidate feature point pairs Match1, retain the candidate feature point pairs with the first confidence level greater than 0.9 to form the target feature points Match2, and record the inlier rate ratio1 and the number of inliers num1 at the same time.
[0105] (3) Use the target feature points Match2 to obtain the 2D coordinates of the target feature points of the positioning image and the 3D point cloud coordinates of the corresponding positioning map, and form the 2D-3D matching pairs Match3.
[0106] The model C adopts a deep learning feature point matching model, which can be a deep learning-based feature point matching model such as LightGlue or SurperGlue.
[0107] 5. Positioning solution module
[0108] (1) Use the cv2.solvePnPRansac function provided by OpenCV to perform positioning solution on the 2D-3D matching pairs Match3, record the transformation matrix Pose1, retain the 2D-3D matching pairs with the second confidence level greater than 0.6 to form the 2D-3D inlier matching pairs Match4, and record the inlier rate ratio2 and the number of inliers num2 at the same time.
[0109] (2) Use the BA algorithm to solve and fine-tune the 2D-3D inlier matching pairs Match4, use Pose1 as the initial value, and record the fine-tuned optimized transformation matrix Pose2, as well as the corresponding covariance estimators sigma_position and sigma_rotation.
[0110] 6. Evaluation score determination module
[0111] Use the inlier rates ratio1, ratio2, the number of inliers num1, num2, and the covariance estimators sigma_position and sigma_rotation obtained by the BA solution to evaluate the transformation matrix, and distinguish the quality of the positioning. The evaluation score calculation process is as follows:
[0112] (1) Data normalization
[0113] a. Among them, when ratio1 and ratio2 take values from 0 to 1, no normalization is required.
[0114] b. Truncated normalization operations are performed on num1 and num2. Assuming the upper limit maximum value is MaxNum, the specific formula is as follows:
[0115] norm_num1 = num1 / max(num1, MaxNum)
[0116] norm_num2 = num2 / max(num2, MaxNum)
[0117] c. The BA solution covariance estimators sigma_position and sigma_rotation can well measure the uncertainties in the translational part and rotational part of the positioning result. Therefore, it is a reverse index. The smaller the value, the smaller the uncertainty and the more accurate the positioning result. Therefore, a reverse-phase normalization operation is adopted. Assuming the upper limit maximum values are MaxSigmaPosition and MaxSigmaRotation, the specific formula is as follows:
[0118] norm_sigma_position = 1 - sigma_position / max(sigma_position, MaxSigmaPosition)
[0119] norm_sigma_position = 1 - sigma_rotation / max(sigma_rotation, MaxSigmaRotation)
[0120] In the embodiment, the parameters MaxNum, MaxSigmaPosition, and MaxSigmaRotation in the normalization take values 100, 0.001, and 0.0001 respectively.
[0121] (2) Calculation of weighted evaluation score
[0122] The evaluation score conf = weight11 * norm_num1 + weight12 * norm_num2 + weight21 * ratio1 + weight22 * ratio2 + weight31 * norm_sigma_position + weight32 * norm_sigma_position.
[0123] The parameters in the formula are assigned according to the contribution degree. In the embodiment, weight11 and weight12 are respectively taken as 0.15; weight21 and weight22 are respectively taken as 0.15; weight31 and weight32 are respectively taken as 0.4.
[0124] When using the positioning result, the accuracy of the positioning result can be judged according to the evaluation score. Assume that the evaluation score threshold is conf_threshold. When the evaluation score is greater than conf_threshold, the transformation matrix is considered available, otherwise the transformation matrix is unavailable.
[0125] Figure 5 The figure is a schematic structural diagram of a visual positioning device provided by an embodiment of the present application. The device can execute the visual positioning method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. As Figure 5 shown, the device includes:
[0126] A target feature point determination module 410, configured to collect an image of a target scene through an augmented reality device to obtain a positioning image, perform feature point matching on the positioning image and a scene image in a preset image set, and determine target feature points;
[0127] A coordinate determination module 420, configured to determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point according to the point cloud data in the target scene;
[0128] A transformation matrix determination module 430, configured to perform positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determine a transformation matrix corresponding to the positioning image, and perform fine adjustment calculation on the transformation matrix based on an optimization algorithm;
[0129] A transformation matrix evaluation module 440, configured to determine an evaluation score of the transformation matrix according to at least one of matching parameters in the feature point matching process, calculation parameters in the positioning calculation process, and fine adjustment parameters in the fine adjustment calculation process, and determine an available evaluation result of the transformation matrix according to the evaluation score.
[0130] In the embodiment of the present application, the target feature point determination module 410 performs feature point matching on the positioning image and a scene image in a preset image set to determine target feature points, including:
[0131] Extract feature points from the positioning image, determine the feature points in the positioning image, and extract image vectors of the positioning image to determine the image vectors of the positioning image;
[0132] Determine a scene image from the preset image set according to the image vector of the positioning image and the image vectors of the preset images in the preset image set;
[0133] Match the feature points in the positioning image with the feature points in the scene image to determine target feature points.
[0134] In the embodiment of the present application, the target feature point determination module 410 determines a scene image from the preset image set according to the image vector of the positioning image and the image vectors of the preset images in the preset image set, including:
[0135] Perform a similarity comparison on the image vector of the positioning image and the image vectors of the preset images, and use the preset number of preset images with the highest similarity as the scene image;
[0136] Correspondingly, matching the feature points in the positioning image with the feature points in the scene image to determine target feature points includes:
[0137] Traverse each scene image, respectively perform similarity matching on the feature points in the scene image and the feature points in the positioning image, and use the feature point pairs with similarity exceeding the preset similarity threshold as candidate feature point pairs;
[0138] Verify the candidate feature point pairs based on a verification algorithm, use the candidate feature point pairs with the first confidence level greater than the first preset confidence threshold as the target feature points, and record the first inlier number and the first inlier rate.
[0139] In the embodiment of the present application, the transformation matrix determination module 430 performs positioning calculation based on the two-dimensional coordinates and the three-dimensional coordinates of the target feature points to determine the transformation matrix corresponding to the positioning image, including:
[0140] Call a positioning calculation function to perform positioning calculation on the two-dimensional coordinates and the three-dimensional coordinates to determine the transformation matrix corresponding to the positioning image, and correspondingly output a second confidence level;
[0141] Use the two-dimensional coordinates and the three-dimensional coordinates with the second confidence level greater than the second preset confidence threshold as target coordinate pairs, and record the second inlier number and the second inlier rate.
[0142] In the embodiment of the present application, the transformation matrix determination module 430 performs fine calculation and adjustment on the transformation matrix based on an optimization algorithm, including:
[0143] Perform fine calculation and adjustment on the transformation matrix based on an optimization algorithm to obtain an optimized transformation matrix, and correspondingly output a translation covariance and a rotation covariance.
[0144] In the embodiments of the present application, the matching parameters in the feature point matching process include the first inlier number and the first inlier rate of the target feature point pairs whose first confidence is greater than the first confidence threshold selected during the verification of candidate feature point pairs. The candidate feature point pairs are the feature point pairs with a similarity exceeding a preset similarity threshold among the feature points in the scene image and the feature points in the positioning image; the solution parameters in the positioning solution process include the second inlier number and the second inlier rate of the two-dimensional coordinates and the three-dimensional coordinates corresponding to the output when performing positioning solution on the two-dimensional coordinates and the three-dimensional coordinates, where the second confidence is greater than the second confidence threshold; the fine-tuning parameters in the fine-tuning process of the solution include the translation covariance and the rotation covariance corresponding to the output during the fine-tuning process of the solution of the transformation matrix;
[0145] Accordingly, the transformation matrix determination module 430 determines the evaluation score of the transformation matrix based on at least two of the matching parameters in the feature point matching process, the solution parameters in the positioning solution process, and the fine-tuning parameters in the fine-tuning process of the solution, including:
[0146] Normalize at least two of the first inlier number, the first inlier rate, the second inlier number, the second inlier rate, the translation covariance, and the rotation covariance respectively, and perform weighted summation on the normalized values to obtain the evaluation score of the transformation matrix; where the evaluation score of the transformation matrix is positively correlated with the first inlier number, the first inlier rate, the second inlier number, and the second inlier rate respectively, and negatively correlated with the translation covariance and the rotation covariance respectively.
[0147] In the embodiments of the present application, the transformation matrix evaluation module 440 determines the available evaluation result of the transformation matrix based on the evaluation score, including:
[0148] If the evaluation score is greater than the preset evaluation score, it is determined that the transformation matrix is available for positioning based on the rotation matrix;
[0149] If the evaluation score is less than or equal to the preset evaluation score, it is determined that the transformation matrix is unavailable, and the process of re-determining the transformation matrix starts from performing feature point matching between the positioning image and the scene images in the preset image set.
[0150] A visual positioning device provided in the embodiments of the present application can execute a visual positioning method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the method.
[0151] Figure 6The structural schematic diagram of the electronic device 10 that can be used to implement the embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0152] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0153] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless vision positioning transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0154] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the vision positioning method.
[0155] In some embodiments, the visual positioning method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the visual positioning method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the visual positioning method by any other suitable means (e.g., by means of firmware).
[0156] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0157] The computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable visual positioning device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0158] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0160] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0161] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0162] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the information expected by the technical solution of this application can be achieved, and no limitation is imposed herein.
[0163] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A visual positioning method, characterized in that: The method comprises: Capturing an image of a target scene through an extended reality device to obtain a positioning image, matching feature points of the positioning image with scene images in a preset image set, and determining target feature points; Determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point based on the point cloud data in the target scene; Performing positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determining a transformation matrix corresponding to the positioning image, and calculating and fine-tuning the transformation matrix based on an optimization algorithm; The evaluation score of the transformation matrix is determined according to at least one of the matching parameters in the feature point matching process, the solution parameters in the positioning solution process, and the fine-tuning parameters in the solution fine-tuning process, and the available evaluation result of the transformation matrix is determined according to the evaluation score.
2. The method according to claim 1, characterized in that: Matching feature points of the positioning image with scene images in a preset image set to determine target feature points includes: Extracting feature points from the positioning image to determine feature points in the positioning image, and extracting vectors from the positioning image to determine image vectors of the positioning image; Determine a scene image from the preset image set according to an image vector of the positioning image and image vectors of preset images in the preset image set; The feature points in the positioning image are matched with the feature points in the scene image to determine the target feature points.
3. The method according to claim 2, characterized in that Determining a scene image from the preset image set according to an image vector of the positioning image and an image vector of a preset image in the preset image set includes: Comparing the image vector of the positioning image with the image vector of the preset image for similarity, and taking a preset number of preset images with the highest similarity as scene images; Accordingly, matching the feature points in the positioning image with the feature points in the scene image to determine the target feature points includes: Traversing each scene image, respectively matching the feature points in the scene image with the feature points in the positioning image by similarity, and taking the feature point pairs whose similarity exceeds a preset similarity threshold as candidate feature point pairs; The candidate feature point pairs are verified based on the verification algorithm, the candidate feature point pairs whose first confidence is greater than the first preset confidence threshold are taken as target feature points, and the first number of inliers and the first inlier rate are recorded.
4. The method according to claim 1, characterized in that: Performing positioning calculation according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point to determine a transformation matrix corresponding to the positioning image includes: Calling a positioning solution function to perform positioning solution on the two-dimensional coordinates and the three-dimensional coordinates, determining a transformation matrix corresponding to the positioning image, and outputting a second confidence level accordingly; The two-dimensional coordinates and the three-dimensional coordinates whose second confidence is greater than a second preset confidence threshold are taken as a target coordinate pair, and a second number of inliers and a second inlier rate are recorded.
5. The method according to claim 4, characterized in that The conversion matrix is solved and fine-tuned based on an optimization algorithm, including: The transformation matrix is solved and fine-tuned based on an optimization algorithm to obtain an optimized transformation matrix, and a translation covariance and a rotation covariance are output accordingly.
6. The method according to any one of claims 1 to 5, characterized in that The matching parameters in the feature point matching process include the first inliers number and the first inliers rate of the target feature point pairs whose first confidence is greater than the first confidence threshold value selected when checking the candidate feature point pairs, and the candidate feature point pairs are feature point pairs whose similarity between the feature points in the scene image and the feature points in the positioning image exceeds the preset similarity threshold value; the solution parameters in the positioning solution process include the second inliers number and the second inliers rate of the two-dimensional coordinates and the three-dimensional coordinates whose corresponding output second confidence is greater than the second confidence threshold value when positioning solution is performed on the two-dimensional coordinates and the three-dimensional coordinates; the fine-tuning parameters in the solution fine-tuning process include the translation covariance and the rotation covariance correspondingly output in the solution fine-tuning process of the transformation matrix; Accordingly, the evaluation score of the transformation matrix is determined according to at least two of the matching parameters in the feature point matching process, the solution parameters in the positioning solution process, and the fine-tuning parameters in the solution fine-tuning process, including: At least two of the first number of inliers, the first inlier rate, the second number of inliers, the second inlier rate, the translation covariance and the rotation covariance are normalized respectively, and the normalized values are weighted summed to obtain an evaluation score of the transformation matrix; wherein the evaluation score of the transformation matrix is positively correlated with the first number of inliers, the first inlier rate, the second number of inliers and the second inlier rate, respectively, and negatively correlated with the translation covariance and the rotation covariance, respectively.
7. The method according to claim 1, characterized in that Determining an available evaluation result of the conversion matrix according to the evaluation score includes: If the evaluation score is greater than a preset evaluation score, determining that the transformation matrix is available to perform positioning based on the rotation matrix; If the evaluation score is less than or equal to a preset evaluation score, it is determined that the conversion matrix is unavailable, and the conversion matrix is re-determined starting from matching feature points between the positioning image and a scene image in a preset image set.
8. A visual positioning device, characterized in that: The device comprises: A target feature point determination module is used to acquire an image of a target scene through an extended reality device to obtain a positioning image, match feature points of the positioning image with a scene image in a preset image set, and determine target feature points; A coordinate determination module, used to determine the two-dimensional coordinates of the target feature point in the positioning image, and determine the three-dimensional coordinates of the target feature point based on the point cloud data in the target scene; A conversion matrix determination module is used to perform positioning and solving according to the two-dimensional coordinates and the three-dimensional coordinates of the target feature point, determine the conversion matrix corresponding to the positioning image, and solve and fine-tune the conversion matrix based on an optimization algorithm; The transformation matrix evaluation module is used to determine the evaluation score of the transformation matrix according to at least one of the matching parameters in the feature point matching process, the solution parameters in the positioning solution process, and the fine-tuning parameters in the solution fine-tuning process, and determine the available evaluation result of the transformation matrix according to the evaluation score.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the visual positioning method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the visual positioning method according to any one of claims 1 to 7 when executed.
Citation Information
Cited By
Ion implanter camera offset detection method and system based on feature matching
CN121095240A