XR equipment repositioning method without map prior
By building an image dictionary of offline query library and online data, combined with PnP and PGO algorithms, XR device relocation without map priors is achieved, solving the problems of high cost and unstable accuracy in the existing technology, and improving positioning accuracy and scalability.
Patent Information
- Application Number
- CN202510510550.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing XR device relocation technology relies on pre-built 3D point cloud maps, resulting in high costs, large storage space occupancy, and unstable positioning accuracy due to ambient light and texture.
Using a map-free prior method, the image dictionary of offline query library and online data is constructed, keyframes are selected for reconstruction of 3D point cloud maps, and the pose is calculated using PnP and PGO algorithms to realize the repositioning of XR devices.
It improves positioning accuracy, reduces storage and deployment costs, reduces redundant calculations, improves the accuracy of three-dimensional reconstruction, and has a wider range of applications.
Smart Images

Figure CN120070818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual positioning, and particularly to a method for relocating an XR device without prior map information. Background Art
[0002] The XR device relocating technology is the core of achieving accurate alignment between the virtual and real spaces, and its goal is to determine the position and pose of the device in the physical environment in real time through sensors and algorithms.
[0003] Currently, base station positioning and visual positioning methods are commonly used for relocating. Among them, base station positioning installs additional laser sensors or infrared sensors in the environment and installs data receivers on the XR device to achieve millimeter-level positioning accuracy; visual positioning requires prior reconstruction of a 3D point cloud positioning map of the environment, which has the advantages of high positioning accuracy and no need to install additional hardware on the XR device.
[0004] However, base station positioning requires the deployment of fixed base stations, resulting in high hardware costs, large volumes, and limited application scopes, and is usually used for relocating industrial-grade XR devices; in visual positioning, due to the influence of ambient light and environmental textures, not only the reconstruction quality of the 3D point cloud map cannot be guaranteed, but also there are defects such as high acquisition costs and large storage space occupation. Summary of the Invention
[0005] In view of the above analysis, embodiments of the present invention aim to provide a method for relocating an XR device without prior map information to solve the problems of high costs and large storage space occupation caused by the existing relocating method relying on a pre-constructed point cloud map.
[0006] Embodiments of the present invention provide a method for relocating an XR device without prior map information, including the following steps: Construct an offline query library according to the environmental data collected by the XR device; Construct a positioning sequence of the current image frame according to the online data collected by the XR device, and reconstruct a 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame; Obtain multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtain the matching point pairs between each offline image and the 3D point cloud map, and then calculate the pose of each offline image in the XR device coordinate system at the current moment, and select the optimal offline image from them; Relocate the pose of the XR device at the current moment according to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment.
[0007] Based on a further improvement of the above method, the environmental data includes a first image stream and the pose corresponding to each frame of the image stream; the online data includes a second image stream and the pose data corresponding to each frame of the image stream; Perform the following preprocessing on environmental data and online data: Eliminate the images in the first image stream and the second image stream that do not meet the brightness, blurriness, and texture conditions, and select key frames by extracting the feature points of the images; Construct their respective image dictionaries, encode to obtain key values according to the poses of the selected key frames, and if the key values are not in the corresponding image dictionaries, put the key values into the corresponding image dictionaries and mark the key frames.
[0008] Based on a further improvement of the above method, constructing an offline query library according to the environmental data collected by the XR device is to, after preprocessing the environmental data, save the timestamps, corresponding poses, image data, image feature point data, and image vector data of the marked key frames to the storage of the XR device to form an offline query library.
[0009] Based on a further improvement of the above method, constructing a positioning sequence of the current image frame according to the online data collected by the XR device includes: after preprocessing the online data, select multiple key frames with the smallest time or distance difference from the marked key frames as multiple historical key frames; form a positioning sequence of the current image frame by combining the multiple historical key frames and the current image frame.
[0010] Based on a further improvement of the above method, reconstructing a 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame includes: Match the feature points of the historical key frames and the current image frame in the positioning sequence to obtain 2D-2D matching point pairs; according to the 2D-2D matching point pairs, the camera parameters of the XR device, and the poses corresponding to each frame of image in the positioning sequence, use the SFM algorithm to obtain a 3D point cloud map of the XR device coordinate system at the current moment.
[0011] Based on a further improvement of the above method, calculate the pose of each frame of offline image in the XR device coordinate system at the current moment, including: According to the 2D-3D matching point pairs of each frame of offline image and the 3D point cloud map, use the PnP algorithm to perform pose solution for each frame of offline image to obtain the initial pose, inlier rate, and actual number of matching point pairs of each frame of offline image in the XR device coordinate system at the current moment; According to the initial pose of each frame of offline image and the pose in the offline query library, construct the first relative pose constraint and the second relative pose constraint; based on the first relative pose constraint and the second relative pose constraint, use the PGO algorithm to perform positioning optimization on the initial poses of multiple frames of offline images to obtain the optimized poses, total position optimization error, and total attitude optimization error; where the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.
[0012] Based on the further improvement of the above method, the first relative pose constraint and the second relative pose constraint are obtained by calculating the relative pose relationship between the initial poses of every two adjacent offline images and the relative pose relationship between the poses in the offline query library respectively after sorting according to the similarity between each frame of offline image and the current image frame.
[0013] Based on the further improvement of the above method, the optimal offline image is selected as the offline image with the highest inlier rate.
[0014] Based on the further improvement of the above method, according to the inlier rate and the actual number of matching point pairs of each frame of offline image, as well as the total position optimization error and the total attitude optimization error of multiple frames of offline images, the confidence of the pose of the XR device at the current moment is calculated by weighted calculation.
[0015] Based on the further improvement of the above method, according to the poses corresponding to the selected key frames respectively, the key values are obtained by encoding through the following formula: key = str(floor(x / posTh)) + "_" + str(floor(y / posTh)) + "_" + str(floor(z / posTh)) + "_" + str(floor(roll / rotTh)) + "_" + str(floor(pitch / rotTh)) + "_" + str(floor(yaw / rotTh)), where key represents the key value of the key frame, str(·) represents the numerical conversion to character operation, floor(·) represents the floor operation, posTh and rotTh represent the position segmentation threshold and the rotation segmentation threshold respectively; (x, y, z) represent the positions of the x, y, and z axes in the pose respectively, and (roll, pitch, yaw) represent the roll angle, pitch angle, and yaw angle in the pose respectively.
[0016] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects: 1. Utilize the offline query library constructed in the offline processing stage to provide prior data support for the online relocalization stage, with high positioning accuracy and low storage occupancy.
[0017] 2. There is no need to pre - construct a 3D point cloud map, with low deployment and storage costs and strong scalability.
[0018] 3. By constructing an image dictionary to select key frames, redundant calculations are reduced, and the poses of the images are encoded as the key values of the image dictionary, making the key frames evenly distributed in space and improving the accuracy of three - dimensional reconstruction.
[0019] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the following specification, and some advantages can be made obvious from the specification, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained from the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings are only for the purpose of showing specific embodiments, and are not considered as a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components; Figure 1 It is a flowchart of a method for relocating an XR device without map prior in an embodiment of the present invention; Figure 2 It is a schematic diagram of the positioning sequence of the image frame at time t in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0022] The coordinate system of the XR device may shift after restart. A specific embodiment of the present invention discloses a method for relocating an XR device without map prior. By repositioning the position of the XR device in the real world, the coordinate system is aligned to a unified world coordinate system, thereby ensuring the correct placement and interaction of virtual objects. As Figure 1 shown, the method includes the following steps: S1. Construct an offline query library according to the environmental data collected by the XR device; S2. Construct the positioning sequence of the current image frame according to the online data collected by the XR device, and reconstruct the 3D point cloud map of the coordinate system of the XR device at the current moment according to the positioning sequence of the current image frame; S3. Obtain multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtain the matching point pairs between each offline image and the 3D point cloud map, and then calculate the pose of each offline image in the coordinate system of the XR device at the current moment, and select the optimal offline image therefrom; S4. Reposition the pose of the XR device at the current moment according to the pose of the optimal offline image in the offline query library and the pose in the coordinate system of the XR device at the current moment.
[0023] During implementation, the offline query library constructed in the offline processing stage of step S1 provides prior data support for the online relocating stage of steps S2 - S4, with high positioning accuracy and low storage occupancy.
[0024] It should be noted that the XR device in this embodiment is provided with a camera and an IMU sensor, and can be an AR glasses, VR glasses or MR glasses; the XR device uses the camera to collect images and output an image stream, and outputs the six-degree-of-freedom (Six Degrees of Freedom, abbreviated as 6DOF) pose data corresponding to each frame of the image stream through the positioning interface, where the six-degree-of-freedom pose data includes: three rotational degrees of freedom (pitch, yaw, roll) and three translational degrees of freedom (front-back, left-right, up-down movement along the x, y, z axes), realizing dynamic perception of high-precision position and attitude, which can be represented by a 6D vector or a 4×4 transformation matrix; the pose in this embodiment is the 6DOF pose data.
[0025] In step S1, the XR device is used to scan the environment, and the collected environmental data includes the first image stream and the pose corresponding to each frame of the image. Further, in order to obtain high-quality images and improve the processing efficiency, the following preprocessing is performed on the environmental data: ① Images that do not meet the brightness, blur, and texture conditions in the first image stream are removed, and key frames are selected by extracting the feature points of the images.
[0026] It should be noted that the brightness condition is that the occupancy ratio of the overexposed or underexposed area of the image is less than or equal to the occupancy ratio threshold; the blur condition is that the variance of the Laplacian of the image is greater than or equal to the variance threshold; the texture condition is that the proportion of edge pixels in the image is greater than or equal to the edge proportion threshold.
[0027] Specifically, when screening images according to the brightness condition, first perform histogram equalization on the images, then calculate the histogram of the images, and count the number of pixels corresponding to each gray value (0~255); then count the number of pixels in the overexposed and underexposed areas respectively, and divide by the total number of pixels of the image to obtain the occupancy ratio of the overexposed or underexposed area; finally, compare with the occupancy ratio threshold respectively. If any one is greater than the occupancy ratio threshold, the image is removed.
[0028] Exemplarily, the equalizeHist() method in OpenCV is used to perform histogram equalization on the images, and the calcHist() method is used to calculate the histogram of the images; the overexposed area is a gray value greater than 220, the underexposed area is a gray value less than 30; the occupancy ratio threshold is set to 0.25.
[0029] When screening images according to the blur condition, the Laplacian() method of OpenCV is used to calculate the variance of the Laplacian of the image; the variance threshold is set to 80.
[0030] When screening images according to the texture condition, the Sobel() or Canny() method of OpenCV is used to detect the edges of the images; the edge proportion threshold is set to 0.05.
[0031] Further, key frames are selected from the images that meet all the above conditions by extracting the feature points of the images.
[0032] ② Construct an image dictionary, encode to obtain a key value according to the pose of the selected key frame. If the key value is not in the corresponding image dictionary, put the key value into the corresponding image dictionary and mark the key frame.
[0033] It should be noted that in this embodiment, by constructing an image dictionary, the pose of the selected key frame is sequentially encoded to obtain a key value key. Each key value key corresponds to a grid area in the physical space. Query whether the key value key exists in the image dictionary. If it exists, it means that there is already a key frame in this grid area, and the current key frame is not the required key frame, that is, each grid area only corresponds to the key frame that first has this key value; if it does not exist, put the key value key into the image dictionary and mark the key frame as the required key frame. The value value corresponding to each key value key in the image dictionary will not be used and can be set to the same data, such as 1.
[0034] Specifically, according to the pose corresponding to the selected key frame, the key value is encoded through the following formula: key = str(floor(x / posTh)) + "_" + str(floor(y / posTh)) + "_" + str(floor(z / posTh)) + "_" + str(floor(roll / rotTh)) + "_" + str(floor(pitch / rotTh)) + "_" + str(floor(yaw / rotTh)) Among them, key represents the key value of the key frame, str(·) represents the numerical conversion to character operation, floor(·) represents the floor operation, posTh and rotTh respectively represent the position segmentation threshold and the rotation segmentation threshold; (x, y, z) respectively represent the positions of the x, y, and z axes in the pose, and (roll, pitch, yaw) respectively represent the roll angle, pitch angle, and yaw angle in the pose.
[0035] Exemplarily, posTh is set to 0.5m and rotTh is set to 10 degrees.
[0036] It can be seen from the encoding formula of the key value that there will be a situation where multiple key frames calculate the same key value, but only the first key frame corresponding to this key value is required. This method reduces redundant calculations while ensuring sufficient parallax between key frames and making the key frames evenly distributed in space, improving the accuracy of 3D modeling.
[0037] Further, save the timestamps, corresponding poses, image data, image feature point data, and image vector data of the marked key frames to the XR device storage to form an offline query library. Among them, the image feature point data can be obtained by using a feature point extraction model based on deep learning, such as SurperPoint, R2D2, D2Net, SoSNet, Disk, etc.; it can also be obtained by using traditional feature point extraction methods such as ORB and SIFT. The image vector can be obtained by using an image vector extraction model based on deep learning, such as deep learning models like NetVLad and CosPlace; it can also be obtained by using an image vector encoding method based on the bag of words.
[0038] Steps S2 - S4 are the process of the user performing relocalization using the real-time collected online data after turning on the XR device. Among them, step S2 uses the historical data cached online to construct a real-time positioning sequence corresponding to each moment, which is convenient for reconstructing the real-time point cloud map at each moment in step S3, and then in step S4, by calculating the pose of the most similar offline image in the offline query library constructed in step S1 in the XR device coordinate system, the pose of the XR device at each moment is relocalized.
[0039] It should be noted that for the online data collected in real time in step S2, including the second image stream and the pose data corresponding to each frame of the image stream, the same preprocessing method as in step S1 is also used to screen the images in the second image stream and select key frames from them, and mark the required key frames by constructing an image dictionary of the online data.
[0040] When constructing the positioning sequence of the current image frame, select multiple key frames with the smallest time or distance gap from the marked key frames as multiple historical key frames; form the positioning sequence of the current image frame by combining the multiple historical key frames and the current image frame. Exemplarily, the positioning sequence of the image frame at the current t moment is as Figure 2 shown; if 10 historical key frames are selected, then Figure 2 n in is 10, and there are a total of 11 frames of images in the positioning sequence.
[0041] It should be noted that this embodiment does not limit whether the current image frame is a key frame; the smallest time gap from the current image frame means selecting multiple historical key frames with the smallest time gap before the moment where the current image frame is located; the smallest distance gap from the current image frame means that after representing the 6DOF pose with a 6D vector, comparing the distances between the vectors and selecting multiple historical key frames with the smallest gap.
[0042] Further, reconstruct the 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame, including: Match the feature points of the historical key frames and the current image frame in the positioning sequence to obtain 2D-2D matching point pairs; according to the 2D-2D matching point pairs, the camera parameters of the XR device, and the poses corresponding to each frame of the image in the positioning sequence, use the SFM algorithm to obtain the 3D point cloud map in the coordinate system of the XR device at the current moment.
[0043] It should be noted that the camera parameters on the XR device include the camera internal parameters and distortion parameters; if the current image frame is not a key frame, the method for extracting image feature point data in step S1 is used to obtain the feature points of the current image frame; if the current image frame is a key frame, the image feature points have been extracted when selecting the key frame.
[0044] When performing feature point matching, a deep learning feature point matching model can be used, such as LightGlue, SurperGlue and other deep learning-based feature point matching models; or methods such as nearest neighbor matching provided by OpenCV can be used to obtain 2D-2D matching point pairs between image pairs.
[0045] Furthermore, use the SFM (Structure From Motion) algorithm provided by the COLMAP library to realize the reconstruction of the 3D point cloud map of the positioning sequence at the current moment. Among them, when performing reconstruction, the optimization parameters are fixed to the 6DOF pose of the XR device (that is, the poses corresponding to each frame of the image are not modified), and only the positions of the 3D point cloud obtained by triangulation from the 2D-2D matching point pairs are optimized, reducing the number of optimization variables and the computational complexity. Finally, output the 3D point cloud map in the coordinate system of the XR device at the current moment, and the 2D feature points matching the 3D point cloud.
[0046] In step S3, extract the image vector of the current image frame according to the method of extracting the image vector in step S1, and find the top N offline images with the highest similarity to the image vector of the current image frame from the offline query library. Exemplarily, Figure 2 where q1, q2, q3, and q4 are the top 4 offline images with the highest similarity.
[0047] Furthermore, obtain the matching point pairs between each frame of the offline image and the 3D point cloud map. First, match the feature points of each frame of the offline image with the feature points of the current image frame respectively to obtain 2D-2D matching point pairs, and then obtain 2D-3D matching point pairs according to the 3D point cloud matched by the feature points of the current image frame.
[0048] Furthermore, calculating the pose of each frame of the offline image in the coordinate system of the XR device at the current moment includes: pose solution and positioning optimization.
[0049] Among them, the pose solution is to perform pose solution on each frame of offline image according to the 2D-3D matching point pairs between each frame of offline image and the 3D point cloud map, respectively using the PnP algorithm (Perspective-n-Point algorithm) to obtain the initial pose, inlier rate and the number of actual matching point pairs of each frame of offline image in the XR device coordinate system at the current moment.
[0050] It should be noted that the PnP algorithm will combine methods such as RANSAC (Random Sample Consensus) to screen and process the input 2D-3D matching point pairs, and eliminate abnormal matching point pairs. Therefore, the number of actual matching point pairs is the number of matching point pairs actually used in the PnP algorithm; the inlier rate is obtained by dividing the number of actual matching point pairs by the number of input 2D-3D matching point pairs.
[0051] Exemplarily, the solvePnPRansac() method in OpenCV is used for pose solution.
[0052] The positioning optimization is to construct the first relative pose constraint and the second relative pose constraint according to the initial pose of each frame of offline image and the pose in the offline query library; based on the first relative pose constraint and the second relative pose constraint, the PGO algorithm is used to perform positioning optimization on the initial poses of multiple frames of offline images to obtain the optimized pose, the total position optimization error and the total attitude optimization error; among them, the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.
[0053] Exemplarily, it is implemented by using the PGO algorithm provided by the CERES library or the COLMAP library.
[0054] Specifically, the first relative pose constraint and the second relative pose constraint are obtained by calculating the relative pose relationship between the initial poses of every two adjacent frames of offline images and the relative pose relationship between the poses in the offline query library respectively after sorting according to the similarity between each frame of offline image and the current image frame.
[0055] The relative pose relationship is obtained through the following formula: , where represents the pose of the th frame of offline image and the pose of the th frame of offline image The relative pose relationship, that is, the pose of the th frame of offline image relative to the th frame of offline image; "-1" represents the inverse operation. The first relative pose constraint is obtained by substituting the initial pose of each frame of offline image into the above formula, and the second relative pose constraint is obtained by substituting the pose of each frame of offline image in the offline query library into the above formula.
[0056] Finally, select the offline image with the highest inlier rate as the optimal offline image. According to the pose of the optimal offline image in the offline query library and the pose in the coordinate system of the XR device at the current moment, re - locate the pose of the XR device at the current moment through the following formula: , where, denotes the pose of the XR device at time denotes the corresponding pose of the optimal offline image in the offline query library, denotes the pose of the optimal offline image in the coordinate system of the XR device at time
[0057] Furthermore, according to the inlier rate and the actual number of matching point pairs of each frame of offline image, as well as the total error of position optimization and the total error of attitude optimization of multiple frames of offline images, calculate the confidence of the pose of the XR device at the current moment by weighted calculation.
[0058] Specifically, calculate the confidence through the following steps: ① Calculate the average inlier rate and the average number of matching point pairs according to the inlier rate and the actual number of matching point pairs of each frame of offline image, and truncate and normalize the average number of matching point pairs according to the maximum value of the number.
[0059] It should be noted that the value of the average inlier rate meanRatio is in the range of 0 - 1 and does not need to be normalized; the average number of matching point pairs is normalized through the following formula: , where, norm_meanMNum and meanMNum represent the normalized and non - normalized average number of matching point pairs respectively, and MaxNum represents the preset maximum value of the number. Exemplarily, MaxNum is set to 50.
[0060] ② Reverse - truncate and normalize the total error of position optimization and the total error of attitude optimization respectively according to the maximum error of position optimization and the maximum error of attitude optimization.
[0061] It should be noted that both the total error of position optimization and the total error of attitude optimization are reverse indicators. The smaller the value, the more accurate the positioning result. Therefore, a reverse - truncate normalization operation is adopted, and the formula is as follows: , , Among them, norm_errPos and errPos respectively represent the total position optimization errors after and before normalization, norm_errRot and errRot respectively represent the total attitude optimization errors after and before normalization, MaxerrPosNum and MaxerrRotNum respectively represent the preset maximum position optimization error and maximum attitude optimization error, and exemplarily, they are respectively set to 0.002 and 0.0005.
[0062] ③ According to the average inlier rate, the average number of normalized matching point pairs, the total position optimization error, the total attitude optimization error, and their respective weights, the confidence level conf is obtained by weighted summation, and the formula is as follows: , where, , , and respectively represent the weights of the average inlier rate, the average number of normalized matching point pairs, the total position optimization error, and the total attitude optimization error; exemplarily, the weights are sequentially set to: 0.3, 0.2, 0.25, and 0.25.
[0063] The confidence level conf of the pose of the XR device at the current moment obtained by calculation takes values from 0 to 1, and the greater the confidence level, the higher the positioning accuracy.
[0064] Compared with the prior art, a mapless prior XR device relocalization method provided in this embodiment uses the offline query library constructed in the offline processing stage to provide prior data support for the online relocalization stage, has high positioning accuracy, and less storage occupancy; does not require pre-constructing a 3D point cloud map, has low deployment and storage costs, and strong scalability; selects key frames by constructing an image dictionary to reduce redundant calculations, encodes according to the image pose as the key value of the image dictionary, and makes the key frames evenly distributed in space, improving the accuracy of three-dimensional reconstruction.
[0065] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0066] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for relocalizing an XR device without map prior, characterized in that: The following steps are involved: Build an offline query library based on the environmental data collected by XR devices; Build a positioning sequence of the current image frame based on the online data collected by the XR device, and reconstruct the 3D point cloud map of the XR device coordinate system at the current moment based on the positioning sequence of the current image frame; Obtain multiple frames of offline images with the highest similarity to the current image frame from the offline query library, obtain matching point pairs between each frame of the offline image and the 3D point cloud map, and then calculate the pose of each frame of the offline image in the XR device coordinate system at the current moment, and select the best offline image; The pose of the XR device at the current moment is relocated according to the pose of the optimal offline image in the offline query library and the pose of the XR device coordinate system at the current moment.
2. The method for relocalizing an XR device without map prior according to claim 1, characterized in that: The environmental data includes the first image stream and the position and posture data corresponding to each frame image in the image stream; the online data includes the second image stream and the position and posture data corresponding to each frame image in the image stream; The environmental data and the online data are preprocessed as follows: Eliminate images that do not meet brightness, blur and texture conditions in the first image stream and the second image stream, and select key frames by extracting feature points of the images; Construct their own image dictionaries, encode the key values according to the poses of the selected key frames, and if the key values are not in the corresponding image dictionary, put the key values into the corresponding image dictionary and mark the key frames.
3. The method for relocalizing an XR device without map prior according to claim 2, characterized in that: The construction of an offline query library based on the environmental data collected by the XR device is to pre-process the environmental data and save the timestamps of the marked key frames, the corresponding postures, image data, image feature point data and image vector data to the XR device storage to form an offline query library.
4. The method for relocalizing an XR device without map prior according to claim 2, characterized in that: The method of constructing a positioning sequence of the current image frame based on the online data collected by the XR device includes: after preprocessing the online data, selecting multiple key frames with the smallest time or distance difference with the current image frame from the marked key frames as multiple historical key frames; and combining the multiple historical key frames and the current image frame into a positioning sequence of the current image frame.
5. The method for relocalizing an XR device without map prior according to claim 4, characterized in that: The reconstructing the 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame includes: The feature points of the historical key frames and the current image frames in the positioning sequence are matched to obtain 2D-2D matching point pairs. Based on the 2D-2D matching point pairs, the camera parameters of the XR device and the pose corresponding to each frame image in the positioning sequence, the SFM algorithm is used to obtain the 3D point cloud map of the XR device coordinate system at the current moment.
6. The method for relocalizing an XR device without map prior according to claim 1 or 5, characterized in that: The step of calculating the pose of each frame of the offline image in the XR device coordinate system at the current moment includes: According to the 2D-3D matching point pairs between each frame of offline image and the 3D point cloud map, the PnP algorithm is used to solve the pose of each frame of offline image, so as to obtain the initial pose, internal point rate and actual number of matching point pairs of each frame of offline image in the XR device coordinate system at the current moment; According to the initial pose of each frame of offline image and the pose in the offline query library, a first relative pose constraint and a second relative pose constraint are constructed; based on the first relative pose constraint and the second relative pose constraint, the PGO algorithm is used to perform positioning optimization on the initial pose of multiple frames of offline images to obtain the optimized pose, the total position optimization error and the total pose optimization error; the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.
7. The method for relocalizing an XR device without map prior according to claim 6, characterized in that: The first relative posture constraint and the second relative posture constraint are obtained by respectively calculating the relative posture relationship between the initial postures of each two adjacent frames of offline images and the relative posture relationship between the postures in the offline query library after sorting each frame of offline image and the current image frame according to the similarity.
8. The method for relocalizing an XR device without map prior according to claim 6, characterized in that: The selecting of the optimal offline image is selecting the offline image with the highest inlier rate.
9. The method for relocalizing an XR device without map prior according to claim 6, characterized in that: The confidence of the position and pose of the XR device at the current moment is calculated by weighted calculation based on the inlier rate and the number of actual matching point pairs of each frame of offline images, as well as the total error of position optimization and the total error of posture optimization of multiple frames of offline images.
10. The method for relocalizing an XR device without map prior according to claim 2, characterized in that: The key values are obtained by encoding the postures corresponding to the selected key frames using the following formula: key=str(floor(x / posTh))+"_"+str(floor(y / posTh))+"_"+str(floor(z / posTh))+"_"+str(floor(roll / rotTh))+"_"+str(floor(pitch / rotTh))+"_"+str(floor(yaw / rotTh)), Among them, key represents the key value of the keyframe, str(·) represents the operation of converting numeric values to characters, floor(·) represents the operation of rounding down, posTh and rotTh represent the position segmentation threshold and rotation segmentation threshold respectively; (x, y, z) represents the positions of the x, y, and z axes in the pose respectively, and (roll, pitch, yaw) represents the roll angle, pitch angle, and yaw angle in the pose respectively.
Citation Information
Patent Citations
Binocular vision positioning method and binocular vision positioning device for robots, and storage medium
CN107796397A
Image sequence relocation judgment method and device and computer equipment
CN112990003A
Visual positioning method based on point cloud map
CN114723920A
Positioning method and device and electronic equipment
CN116862978A
Unmanned aerial vehicle relative pose positioning method and system
CN118570290A
Cited By
XR equipment image synthesis method and system based on multiple key frames
CN120997055A
A visual positioning reconstruction method for periodic bridge inspection of apparent diseases
CN122617978A
A visual positioning reconstruction method for periodic bridge inspection of apparent diseases
CN122617978B