Continuous Visual Localization Method in Large-Scale Point Cloud Maps Based on Cross-Domain Joint Retrieval
By constructing a visual positioning method for cross-domain joint retrieval in large-scale point cloud maps, using cross-domain matching of images and point cloud features, the problems of inaccurate positioning and high computational complexity in the existing technology are solved, and real-time precise positioning is achieved in complex environments.
Patent Information
- Application Number
- CN202311060237.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-08-22
AI Technical Summary
The existing visual positioning methods are difficult to find the correct matching relationship when dealing with large point clouds, resulting in inaccurate positioning results; while the image retrieval method has limited positioning capabilities in real-time, complex and dynamic environments, high computational complexity and sensitive to environmental changes.
A continuous visual positioning method in large-scale point cloud maps based on cross-domain joint retrieval is proposed. By building an image search module and a point cloud search module, image feature extraction and point cloud encoding are used using residual network and local sensitive hash algorithm, and cross-domain matching is performed through attention fusion model and multi-layer perceptron, and 3D-2D matching is performed by combining RANSAC and EPnP methods.
It improves the speed and accuracy of visual positioning, enhances the robustness of the system, and can achieve real-time precise positioning in large-scale point clouds, suitable for complex and dynamic environments.
Smart Images

Figure CN117235294B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual positioning, and relates to a continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval. Background Art
[0002] Visual positioning refers to accurately positioning and tracking a target in three-dimensional space through computer vision algorithms and sensor information. Visual positioning technology plays an important role in fields such as unmanned driving, augmented reality, and intelligent navigation. Traditional visual positioning technologies mainly rely on feature extraction and matching algorithms. By extracting feature points in images or videos and matching them with feature points in maps or reference images, positioning is achieved. However, these traditional methods often face many challenges in complex environments, such as problems like light changes, occlusion, and image blurring, resulting in low accuracy and stability of positioning.
[0003] Currently, visual positioning methods in point cloud maps mainly rely on point cloud and image matching. For example, a DeepI2P method disclosed by J. Li et al. in the literature "J. Li and G. Hee Lee, 'DeepI2P: Image-to-Point Cloud Registration via Deep Classification,' in Proc. Conference on Computer Vision and Pattern Recognition, 2021, pp. 15955–15964." proposed a visual positioning method based on 2D-3D matching. This method mainly proposed a camera field-of-view classification algorithm based on an attention mechanism to reduce outliers in the matching process and improve the accuracy of 2D-3D matching. However, this method will be unable to find the matching relationship between the point cloud and the image or find a large number of incorrect matching relationships when dealing with large point clouds, which will interfere with the estimation of RANSAC and thus unable to give the correct positioning result.
[0004] There are also some methods that use image retrieval to complete visual localization in large-scale point clouds. For example, a cross-season six-degree-of-freedom visual localization method SensLoc guided by mobile sensors, disclosed by S. Yan et al. in the literature "S.Yan et al., 'Long-term Visual Localization with Mobile Sensors,' in Proc. Conference on Computer Vision and Pattern Recognition, 2023, pp. 17245–17255." This method uses the sensor data built into the mobile device to provide effective initial poses and constraint conditions for visual localization, thereby reducing the search space for image retrieval and pose estimation. However, the image retrieval process of this method still relies on the database and is not optimized for the matching characteristics of images and point clouds. When the database is large, the image retrieval process will become very time-consuming and unable to meet the requirements of real-time localization. At the same time, image retrieval methods are relatively sensitive to environmental changes such as lighting, viewing angles, and occlusions. And due to the limited reference images in the database, when encountering new scenes or unseen situations, image retrieval methods may not be able to provide accurate localization results. Summary of the Invention
[0005] Technical Problems to be Solved
[0006] To avoid the deficiencies of the prior art, the present invention proposes a continuous visual localization method in a large-scale point cloud map based on cross-domain joint retrieval, which mainly solves the limitations of existing visual localization methods and further improves the speed of visual localization. Specifically, the object of the present invention is to improve the following aspects.
[0007] 1. Existing visual localization methods based on point-to-image matching are applicable to a certain extent for localization in small-scale point clouds, but it is difficult to find the correct matching relationship when dealing with large-scale point cloud data and cannot obtain the correct localization result.
[0008] 2. Existing visual localization methods based on image retrieval are applicable to a certain extent for localization in large-scale point clouds, but the disadvantages of this method, such as relying on the database, high computational complexity, sensitivity to environmental changes, and limited generalization ability, limit its localization ability in real-time, complex, and dynamic environments.
[0009] Technical Solution
[0010] A continuous visual localization method in a large-scale point cloud map based on cross-domain joint retrieval, characterized by the following steps:
[0011] Construct an image retrieval module:
[0012] Augment the image set X = {x1, x n ,..., x n} collected during the process of generating the point cloud map to obtain the augmented image set X' = {S1, S2,..., S n}, where S i = {x i , x' i1 , x' i2 ,..., x' ik} is the augmented image set of any image x i ∈ X;
[0013] Supervisedly train an image feature extractor F img based on the residual network using the augmented image set X';
[0014] Extract the features of all images in the image set X using the trained F img and construct a retrieval function Q for these features using the locality-sensitive hashing algorithm;
[0015] For any input image x q , the image retrieval module uses the retrieval function Q to complete the image retrieval process, i.e., X result = Q(F img (x q ))), where X result = {x1, x2,..., x m} is all the images in the image set X that are similar to x q ;
[0016] Construct a point cloud retrieval module:
[0017] Use the KD-Tree algorithm to construct a point cloud lookup function F cut such that it can find the point cloud within a certain range around the input position according to the input position;
[0018] For the point cloud p at the shooting position of any image i = F cut (t i , r), p j = F cut (t j , r), where t i and t j are the shooting positions of the images x i and the image x j respectively, and r is the radius of the point clouds p i and p j . Construct a point cloud encoder F pcd to encode the point cloud to obtain the point cloud feature f i = Fpcd (p i ), f j = F pcd (p j );
[0019] Construct the attention fusion model F a Extract the fusion features from both the point cloud and image modalities, and construct a cross-domain matching model M1 based on a multi-layer perceptron such that it can determine whether an image was taken in the point cloud according to the cross-domain matching relationship between the image features and the point cloud features;
[0020] For any input image x q , the point cloud retrieval module first uses the image retrieval module to obtain all the images X in the image set X that are similar to x q , and uses the point cloud search function F resul to obtain the point cloud P around the shooting positions of these images cut = {p1, p2,..., p result}, then the point cloud is the point cloud near the input image x n : q
[0021]
[0022] Image-point cloud registration module
[0023] For any image x i ∈ X, its shooting position The point cloud p around the shooting position i = F cut (t i , r), where r is the radius of the point cloud p i ;
[0024] Divide p i into the points pi q in x i and the points po q not in x i , the indices of the points in p i are respectively denoted as Ii i and Io i , project p i onto the camera plane to obtain the pixel coordinates q of the point cloud i , where the camera plane is a known parameter during the shooting process of the image x i ;
[0025] Use the point cloud encoder F in the point cloud retrieval module pcd and the attention fusion network F a to extract the fusion features from both the point cloud and image modalities
[0026] Construct and train a field-of-view classification model M2 based on a multi-layer perceptron such that M2(c a )[Ii i = {1} and M2(c a )[Io i = {0}, where the operation [·] represents obtaining the corresponding data according to the index;
[0027] At the same time, construct and train a feature regression model M3 based on a multi-layer perceptron such that M3(c a )[Ii i = q i [Ii i ;
[0028] For any input image x q , the image point cloud registration module first uses the point cloud retrieval module to obtain the point cloud p q around x q , calculates the point pi q of the point cloud p i in the image x q using M2, and then calculates the pixel coordinates of pi q using M3, that is, obtains the 3D-2D matching relationship between the point cloud p q and the image x i ;
[0029] Use the RANSAC and EPnP methods to determine the shooting pose of the query captured image x q from the 3D-2D matching relationship, and use the prior knowledge of the gravity direction to suppress the multiple solutions of the PnP method, that is, complete the visual positioning of the query image x q .
[0030] The acquisition of the augmented image set S i = {x i , x' i1 , x' i2 ,..., x' ik}: For any image x i ∈ X, perform random rotation, translation, scaling, mirroring, cropping, and color jittering to obtain the transformed image x' i , repeat this process k times to obtain the augmented image set S i of the image x i = {x i , x' i1 , x' i2 ,..., x' ik}.
[0031] The training of the image feature extractor F img : For any Randomly select image pairs Image pairs such that and is as small as possible and and is as large as possible.
[0032] The method of constructing the point cloud search function F using the KD-Tree algorithm cut : P cut = F cut (x, y), where x is the central position of the point cloud to be searched and y is the radius of the point cloud to be searched, is the found point cloud.
[0033] The method of constructing the attention fusion model F a Extracts the fusion features from two modalities of point cloud and image, and constructs a cross-domain matching model M1 based on a multi-layer perceptron so that it can judge whether an image is taken in a point cloud according to the cross-domain matching relationship between the image features and the point cloud features: Train F pcd , F a and M1 such that and equals 1 and and equals 0, where g i = F img (x i ), g j = F img (x j ), represents the concatenation operation.
[0034] An application of the continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval, characterized in that: it is used for scenarios with real-time positioning requirements to perform real-time and accurate positioning of targets in a large-scale point cloud.
[0035] An electronic device, characterized in that it includes a processor and a memory, and the processor is used to implement the steps of the method when executing the computer program stored in the memory.
[0036] A readable storage medium, characterized in that a computer program is stored on the readable storage medium, and the computer program implements the steps of the method when executed by a processor.
[0037] Beneficial effects
[0038] A continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval proposed by the present invention constructs an image retrieval module. For any input image x q, the image retrieval module uses the retrieval function Q to complete the image retrieval process; a point cloud retrieval module is constructed. For any input image x q , the point cloud retrieval module first uses the image retrieval module to obtain all the images X in the image set X that are similar to x q similar to the image X result , and uses the point cloud search function F cut to obtain the point cloud around the shooting positions of these images; an image-point cloud registration module is constructed. For any input image x q , the image-point cloud registration module first uses the point cloud retrieval module to obtain the point cloud p q around x q , and uses M2 to calculate the points pi q of the point cloud p i in the image x q , and then uses M3 to calculate the pixel coordinates of pi q , and thus the 3D-2D matching relationship between the point cloud p q and the image x i can be obtained. The RANSAC and EPnP methods can be used to determine the shooting pose of the query image x q from the 3D-2D matching relationship. By using the prior knowledge of the gravity direction to suppress the multiple solutions of the PnP method, the visual positioning of the query image x q can be completed.
[0039] The beneficial effects of the present invention are mainly as follows:
[0040] 1. The system has stronger robustness. The point cloud retrieval module eliminates the ambiguous solutions of the image retrieval module to ensure the effectiveness and robustness of the image retrieval algorithm in visual positioning, thereby realizing visual positioning in a large-scale point cloud.
[0041] 2. The continuous visual positioning speed is fast. The LS-Hash method is used to accelerate the image retrieval speed, and a certain strategy is used to short-circuit the image retrieval module, thereby meeting the speed requirements of real-time positioning.
[0042] 3. The training speed is fast. The image retrieval module, the point cloud retrieval, and the image-point cloud registration module use the image feature extractor and the point cloud feature extractor with shared parameters, thereby ensuring the feature semantic consistency among the three modules and accelerating the model training speed and inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is the flowchart in the present invention:
[0044] a: Flowchart of the image retrieval module, b: Flowchart of the point cloud retrieval module, c: Flowchart of the image-point cloud registration module;
[0045] Flowchart of the image retrieval module
[0046] Figure 2 : The positioning results of continuous positioning requests are obtained using a moving vehicle in the experiment
[0047] Figure 3 : The positioning error of the present invention is mainly distributed in the range of 1m - 3m Specific implementation manner
[0048] The present invention will be further described below in conjunction with embodiments and the accompanying drawings
[0049] In order to overcome the defects that existing visual positioning methods are difficult to be used in large-scale point clouds or cannot perform real-time positioning in large-scale point clouds, the present invention proposes a continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval. Its technology is divided into three major modules: an image retrieval module, a point cloud retrieval module, and an image-point cloud registration module
[0050] Step 1. In order to more effectively meet the requirements of image retrieval in a complex and dynamic environment, the present invention first performs data augmentation on the image set X = {x1, x2,..., x n} collected during the generation of the point cloud map. The main purpose is to improve the accuracy and robustness of image retrieval by increasing the diversity of image data. Specifically, for any image x i ∈ X, a series of random transformations are first performed on it, including rotation, translation, scaling, mirroring, cropping, and color jittering, etc. These transformations can generate a series of new images x' i that are visually different from the original image but consistent in content. This process will be repeated k times to generate the augmented image set S i = {x i , x' i , x' i1 ,..., x' i2} of image x ik . Applying this method to all images in the image set X and converting each image into the corresponding augmented image set, the augmented image set X' = {S1, S2,..., S n} of the image set X can be obtained
[0051] Step 2. Randomly select from the augmented image set X' of X and use ResNet50 as the backbone network to train an image feature extractor F img such that the feature distances of similar images are as small as possible and the feature distances of different images are as large as possible, that is:
[0052]
[0053] Step 3, use F img Extract the features of all images in the image set X and construct a feature set C. Based on the features in C, construct a feature retrieval set using the LS-Hash algorithm. Specifically, we use n random hyperplanes to partition the feature space and map each feature vector to an n-bit hash code. Each hyperplane defines a hash function. If the feature vector is on one side of the hyperplane, the value of the hash function is 1; otherwise, it is 0. In this way, we can map each feature vector to a binary string (i.e., hash code), where each bit corresponds to a hyperplane. In this way, images with similar features will be mapped to the same or similar hash codes, thus achieving fast and efficient feature retrieval. Subsequently, we store these hash codes in the LS-Hash feature retrieval set for subsequent feature matching and image retrieval.
[0054] Step 4, use the KD Tree algorithm to construct a point cloud lookup function F cut such that it can give the point cloud within a certain range around the input position within the complexity of O(logn), that is, P cut = F cut (x, y), where x is the central position of the point cloud to be searched, and y is the radius of the point cloud to be searched, is the retrieved point cloud, where n is the number of points in the point cloud map. Use F cut to search for the point cloud P corresponding to all images in the image set X all = {F cut (t1, r), F cut (t2, r),..., F cut (t n , r)}, where t i is the shooting position of the image x i ∈ X, and r is the preset point cloud radius.
[0055] Step 5, for any point cloud P ∈ P found all , the method proposed by the present invention refers to SO-Net to encode it. Specifically, use farthest point sampling to generate a node set from P Use PointNet to process all the nodes in to obtain where M1 is the number of points of each node, and C1 is the length of the feature vector. Repeat the steps of farthest point sampling and feature extraction to obtain Finally, use PointNet to process P (2) to obtain the point cloud feature
[0056] Step 6, construct an attention fusion model Fa Extract the fused features from two modalities of point cloud and image. Specifically, connect the features of each layer of the point cloud and the image, and use PointNet and Softmax to process them to obtain the attention weights. Weight the point cloud features and image features respectively according to the weights and input them into PointNet. Finally, obtain the feature representations of the point cloud and the image with cross-domain attention fusion between the point cloud and the image. Randomly select from the augmented image set X' of X Construct a cross-domain matching model M1 based on a multi-layer perceptron so that it can judge whether an image is taken in the point cloud according to the cross-domain matching relationship between the image features and the point cloud features, that is, optimize the following expression:
[0057]
[0058] Step 7, for any image x i ∈X, its corresponding point cloud p i =P all , divide p i into the points pi q in x i and the points po q not in x i . Construct and train a field-of-view classification model M2 based on a multi-layer perceptron and a feature regression model M3 based on a multi-layer perceptron, so that M2(c a )[Ii i ={1}, M2(c a )[Io i ={0} and M3(c a )[Ii i =q i [Ii i , where where Ii i and Io i are the indices of pi i and po i in the point cloud p i , q i is the projection of p i on the camera plane, and the operation [·] represents obtaining the corresponding data according to the index.
[0059] Step 8, after completing the training work of the present invention, we use the following steps for continuous visual positioning:
[0060] (1) Use the image retrieval module to screen out all similar images from the feature retrieval set with the image provided by the photographer as the input image. The shooting positions of these images are the possible shooting positions of the input image;
[0061] (2) Use the point cloud search function to obtain the point cloud within a certain range near all possible shooting positions, and use the point cloud retrieval model to select the best point cloud from all possible point clouds as the retrieval result of the point cloud retrieval module. The center of the point cloud is the rough positioning result;
[0062] (3) Use the field of view classification model M2 in the image-point cloud registration module to filter out all the points within the camera's field of view from the point cloud corresponding to the rough positioning result, and use the feature regression model M3 to obtain the pixel coordinates of these points. The RANSAC and EPnP methods can be used to obtain multiple groups of possible camera poses from the matching relationship between the three-dimensional coordinates and pixel coordinates of the points. Based on the prior knowledge of the gravity direction, the most suitable one is selected from multiple groups of possible camera poses as the final positioning result;
[0063] (4) When the shooter moves and requests positioning again, the present invention determines whether it is necessary to re-perform image retrieval according to the request time interval and moving speed of the shooter. If not, we use the result of the previous positioning as the rough positioning result in (2) and start executing the algorithm process from (3). If necessary, we start re-positioning from (1) to complete the continuous positioning loop.
[0064] The effects of the present invention can be further illustrated by the following simulation experiments.
[0065] 1. Simulation conditions
[0066] The present invention conducts simulation experiments on a central processing unit Intel(R) Xeon(R) Silver 4210 CPU, 64G of memory, a graphics processing unit NVIDIA GeForce RTX 2080Ti GPU, and a Windows 10 operating system, using Python software and the PyTorch deep learning framework. The dataset used in the simulation is a self-collected dataset.
[0067] 2. Simulation content
[0068] The experiment uses a moving vehicle to make continuous positioning requests, and the positioning results are as Figure 2 shown. The coordinates in the middle represent the coordinate values of the vehicle after subtracting the center point (460000, 4400000) in the EGSP:4548 projection coordinate system. The point in the upper right corner represents the true position of the vehicle, the point on the left represents the positioning result of the vehicle by the present method, and the point in the lower right corner represents the result of vehicle positioning using the image retrieval method. We conducted an error analysis on the positioning results, and the analysis results are as Figure 3 shown, where the abscissa represents the L2 distance between the vehicle positioning result and the true position of the vehicle, and the ordinate represents the frequency of occurrence of this error.
[0069] From Figure 2It can be seen that the results obtained by continuous visual positioning basically coincide with the actual driving trajectory of the vehicle, which proves that the present invention can effectively perform accurate positioning of targets in large-scale point clouds. As Figure 2 shown in the lower right corner, the number of points for which the present invention calls image retrieval for positioning is very small, which effectively proves the effectiveness of the short-circuit strategy of the present invention. During the test, the positioning request of the vehicle is 30 times per second, which proves that the present invention can be applied to scenarios with real-time positioning requirements.
[0070] From Figure 3 it can be seen that the positioning error of the present invention is mainly distributed in the range of 1m - 3m, and the error is in the same error gradient as the positioning method in small-scale point clouds of the same type, which effectively proves that the present invention can perform real-time accurate positioning of targets in large-scale point clouds.
Claims
1. A continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval, characterized in that The steps are as follows: Construct an image retrieval module: Augment the image set X = {x1, x2,..., x n} collected during the process of generating the point cloud map to obtain the augmented image set X' = {S1, S2,..., S n}, where S i = {x i , x' i1 , x' i2 ,..., x' ik} is the augmented image set of any image x i ∈ X; Supervisedly train a residual network-based image feature extractor F with the augmented image set X'. img ; Use the trained F img Extract the features of all images in the image set X, and use the locality-sensitive hashing algorithm to construct these features into the retrieval function Q; For any input image x q , the image retrieval module uses the retrieval function Q to complete the image retrieval process, i.e., X result = Q(F img (x q ))), where X result = {x1, x2,..., x m} is all the images in the image set X that are similar to x q ; Construct a point cloud retrieval module: Construct a point cloud search function F using the KD-Tree algorithm cut so that it can find the point cloud within a certain range around the input position according to the input position; For any image the point cloud p at its shooting position i = F cut (t i , r), p j = F cut (t j , r), where t i and t j are the shooting positions of image x i and image x j respectively, and r is the radius of point cloud p i and p j . Construct a point cloud encoder F pcd to encode the point cloud to obtain the point cloud feature f i = F pcd (p i ), f j = F pcd (p j ); Construct the attention fusion model F a Extract the fusion features from two modalities of point cloud and image, and construct a cross-domain matching model M1 based on a multi-layer perceptron so that it can judge whether the image is taken in the point cloud according to the cross-domain matching relationship between the image features and the point cloud features; For any input image x q The point cloud retrieval module first uses the image retrieval module to obtain all the points in the image set X that are related to x. q Similar ImagesX result , use the point cloud to find the function F cut Get the point cloud P around the location where these images were taken result ={p1,p2,...,p n }, then the point cloud is the input image x q Nearby point clouds: Image point cloud registration module For any image x i ∈ X, its shooting position The point cloud p around the shooting position i = F cut (t i , r), where r is the radius of the point cloud p i ; Divide p i into points pi q in x i and points po q not in x i , and the indices of the points in p i are respectively represented as Ii i and Io i . Project p i onto the camera plane to obtain the pixel coordinates q i of the point cloud, where the camera plane is a known parameter during the image x i shooting process; Using the point cloud encoder F in the point cloud retrieval module pcd and the attention fusion network F a Extract the fused features from both the point cloud and image modalities Construct and train a vision classification model M2 based on a multi-layer perceptron such that M2(c a )[Ii i = {1} and M2(c a )[Io i = {0}, where the operation [·] represents obtaining the corresponding data according to the index; Construct and train a feature regression model M3 based on a multi-layer perceptron simultaneously such that M3(c a )[Ii i = q i [Ii i ; For any input image x q , the image point cloud registration module first uses the point cloud retrieval module to obtain the point cloud p q around x q , and uses M2 to calculate the points pi q of the point cloud p i in the image x q , and then uses M3 to calculate the pixel coordinates of pi q , that is, the 3D-2D matching relationship between the point cloud p q and the image x i is obtained; Determine the pose of the query captured image x from the 3D-2D matching relationship using the RANSAC and EPnP methods, and utilize the prior knowledge of the gravity direction to suppress the multiple solution cases of the PnP method, thus completing the visual localization of the query image x q q 2. The continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval according to claim 1, wherein: The acquisition of the augmented image set S i ={x i , x' i1 , x' i2 ,..., x' ik}: For any image x i ∈X, perform random rotation, translation, scaling, mirroring, cropping, and color jittering to obtain the transformed image x' i . Repeat this process k times to obtain the augmented image set S i of the image x i ={x i , x' i1 , x' i2 ,..., x' ik}.
3. The continuous visual positioning method in the large-scale point cloud map based on cross-domain joint retrieval according to claim 1, characterized in that: The image feature extractor F img Training: For any Randomly select image pairs Image pair such that and is as small as possible and and is as large as possible.
4. The continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval according to claim 1, characterized in that: The point cloud search function F constructed using the KD-Tree algorithm cut : P cut = F cut (x, y), where x is the central position of the point cloud to be searched, and y is the radius of the point cloud to be searched, is the searched point cloud.
5. The continuous visual positioning method in a large-scale point cloud map based on cross-domain joint retrieval according to claim 1, characterized in that: The constructed attention fusion model F a extracts the fusion features from two modalities of point cloud and image, and constructs a cross-domain matching model M1 based on a multi-layer perceptron so that it can judge whether an image is taken in a point cloud according to the cross-domain matching relationship between the image features and the point cloud features as: Train F pcd , F a and M1 such that and equals 1 and and equals 0, where g i = F img (x i ), g j = F img (x j ), represents the concatenation operation.
6. Application of the continuous visual positioning method in the large-scale point cloud map based on cross-domain joint retrieval according to any one of claims 1 to 5, characterized in that: For scenarios with real-time positioning requirements, perform real-time and precise positioning of the target in a large-scale point cloud.
7. An electronic device, characterized in that, It includes a processor and a memory. When the processor executes the computer program stored in the memory, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium. When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Visual positioning method based on point cloud map
CN114723920A
Visual position identification method for cross-modal retrieval, storage medium and electronic equipment
CN115457125A