An XR device relocalization method without map prior

The map-free XR device repositioning method constructs an offline query library from environmental data to optimize pose estimation, addressing high costs and storage issues, achieving precise and cost-effective XR device repositioning.

CN120070818BActive Publication Date: 2025-07-15HANGZHOU HUIJIAN ZHILIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510550.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-15
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing XR equipment relocation technology has problems such as high hardware costs, large storage space occupation and positioning accuracy affected by ambient light and texture. In particular, the base station positioning and visual positioning methods have defects in cost and accuracy.

Method used

Using the XR device relocation method without map priors, the offline query library is constructed, and the offline query library and the positioning sequence of the current image frame is constructed using the environmental data collected by the XR device, the 3D point cloud map is reconstructed, and the optimal offline image is obtained from the offline query library for relocation, reducing redundant calculations, and improving the uniform distribution and positioning accuracy of keyframes.

Benefits of technology

High-precision XR device relocation is achieved, reducing storage and deployment costs, improving the accuracy and scalability of 3D reconstruction, and reducing dependence on ambient light and textures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070818B_ABST
    Figure CN120070818B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for relocating an XR device without map prior, belonging to the technical field of visual positioning, and solves the problems that the existing relocating depends on a pre-constructed point cloud map, resulting in high cost and large storage space occupation. The method includes: constructing an offline query library according to the collected environmental data; constructing a positioning sequence of the current image frame according to the collected online data, and reconstructing a 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence; obtaining multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtaining the matching point pairs between each offline image and the 3D point cloud map, calculating the pose of each offline image in the XR device coordinate system at the current moment, and selecting the optimal offline image; relocating the pose of the XR device at the current moment according to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment. The relocating with low cost and less storage occupation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual positioning, and in particular, to a method for repositioning an XR device without map prior knowledge. Background Art

[0002] The XR device repositioning technology is the core to achieve accurate alignment between the virtual and real spaces, and its goal is to determine the position and orientation of the device in the physical environment in real time through sensors and algorithms.

[0003] Currently, base station positioning and visual positioning methods are usually used for repositioning. Among them, base station positioning installs additional laser sensors or infrared sensors in the environment and installs data receivers on the XR device to achieve millimeter-level positioning accuracy; visual positioning requires pre-reconstructing a 3D point cloud positioning map of the environment, which has the advantages of high positioning accuracy and no need to install additional hardware on the XR device.

[0004] However, base station positioning requires deploying fixed base stations, resulting in high hardware costs, large volumes, and limited application ranges. It is usually used for repositioning industrial-grade XR devices; in visual positioning, due to the influence of ambient light and environmental textures, not only the reconstruction quality of the 3D point cloud map cannot be guaranteed, but also there are defects such as high acquisition costs and large storage space occupancy. Summary of the Invention

[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method for repositioning an XR device without map prior knowledge to solve the problems of high costs and large storage space occupancy caused by the existing repositioning relying on a pre-constructed point cloud map.

[0006] The embodiments of the present invention provide a method for repositioning an XR device without map prior knowledge, including the following steps:

[0007] Construct an offline query library according to the environmental data collected by the XR device;

[0008] Construct a positioning sequence of the current image frame according to the online data collected by the XR device, and reconstruct a 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame;

[0009] Obtain multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtain the matching point pairs between each offline image and the 3D point cloud map, and then calculate the poses of each offline image in the XR device coordinate system at the current moment, and select the optimal offline image from them;

[0010] Reposition the pose of the XR device at the current moment according to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment.

[0011] Based on further improvements to the above method, the environmental data includes the first image stream and the poses corresponding to each frame of the image stream; the online data includes the second image stream and the pose data corresponding to each frame of the image stream.

[0012] Perform the following preprocessing on the environmental data and the online data:

[0013] Eliminate the images in the first image stream and the second image stream that do not meet the brightness, blur, and texture conditions, and select key frames by extracting the feature points of the images.

[0014] Construct their respective image dictionaries, encode the key values according to the poses of the selected key frames, and if the key value is not in the corresponding image dictionary, put the key value into the corresponding image dictionary and mark the key frame.

[0015] Based on further improvements to the above method, constructing an offline query library according to the environmental data collected by the XR device is to, after preprocessing the environmental data, save the timestamps, corresponding poses, image data, image feature point data, and image vector data of the marked key frames to the storage of the XR device to form an offline query library.

[0016] Based on further improvements to the above method, constructing the positioning sequence of the current image frame according to the online data collected by the XR device includes: after preprocessing the online data, select multiple key frames with the smallest time or distance difference from the marked key frames as multiple historical key frames; form the positioning sequence of the current image frame by combining the multiple historical key frames and the current image frame.

[0017] Based on further improvements to the above method, reconstructing the 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame includes:

[0018] Match the feature points of the historical key frames and the current image frame in the positioning sequence to obtain 2D-2D matching point pairs; according to the 2D-2D matching point pairs, the camera parameters of the XR device, and the poses corresponding to each frame of the image in the positioning sequence, use the SFM algorithm to obtain the 3D point cloud map of the XR device coordinate system at the current moment.

[0019] Based on further improvements to the above method, calculating the pose of each frame of offline image in the XR device coordinate system at the current moment includes:

[0020] According to the 2D-3D matching point pairs between each frame of offline image and the 3D point cloud map, use the PnP algorithm to perform pose solution for each frame of offline image respectively, and obtain the initial pose, inlier rate, and actual number of matching point pairs of each frame of offline image in the XR device coordinate system at the current moment.

[0021] Construct the first relative pose constraint and the second relative pose constraint according to the initial pose of each frame of offline image and the pose in the offline query library; based on the first relative pose constraint and the second relative pose constraint, use the PGO algorithm to optimize the positioning of the initial poses of multiple frames of offline images to obtain the optimized poses, the total position optimization error, and the total attitude optimization error; wherein the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.

[0022] Based on a further improvement of the above method, the first relative pose constraint and the second relative pose constraint are obtained by calculating the relative pose relationship between the initial poses of every two adjacent frames of offline images and the relative pose relationship between the poses in the offline query library respectively after sorting according to the similarity between each frame of offline image and the current image frame.

[0023] Based on a further improvement of the above method, the optimal offline image is selected as the offline image with the highest inlier rate.

[0024] Based on a further improvement of the above method, according to the inlier rate and the actual number of matching point pairs of each frame of offline image, as well as the total position optimization error and the total attitude optimization error of multiple frames of offline images, the confidence of the pose of the XR device at the current moment is calculated by weighted calculation.

[0025] Based on a further improvement of the above method, according to the poses corresponding to the selected key frames respectively, the key value is obtained by encoding through the following formula:

[0026] key = str(floor(x / posTh)) + "_" + str(floor(y / posTh)) + "_" + str(floor(z / posTh)) + "_" + str(floor(roll / rotTh)) + "_" + str(floor(pitch / rotTh)) + "_" + str(floor(yaw / rotTh)),

[0027] wherein, key represents the key value of the key frame, str(·) represents the numerical conversion to character operation, floor(·) represents the floor operation, posTh and rotTh represent the position segmentation threshold and the rotation segmentation threshold respectively; (x, y, z) represent the positions of the x, y, and z axes in the pose respectively, and (roll, pitch, yaw) represent the roll angle, pitch angle, and yaw angle in the pose respectively.

[0028] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:

[0029] 1. Utilize the offline query library constructed in the offline processing stage to provide prior data support for the online relocalization stage, with high positioning accuracy and less storage occupancy.

[0030] 2. It is not necessary to pre - construct a 3D point cloud map, with low deployment and storage costs and strong scalability.

[0031] 3. By constructing an image dictionary to select key frames, redundant calculations are reduced. The pose of the image is encoded as the key value of the image dictionary, making the key frames evenly distributed in space and improving the accuracy of 3D reconstruction.

[0032] In the present invention, the above - mentioned technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification. Moreover, some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components;

[0034] Figure 1 It is a flowchart of a method for relocating an XR device without map prior in an embodiment of the present invention;

[0035] Figure 2 It is a schematic diagram of the positioning sequence of the image frame at time t in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0037] The coordinate system of the XR device may shift after restart. A specific embodiment of the present invention discloses a method for relocating an XR device without map prior. By re - locating the position of the XR device in the real world, the coordinate system is aligned to a unified world coordinate system, thereby ensuring the correct placement and interaction of virtual objects. As Figure 1 shown, the method includes the following steps:

[0038] S1. Construct an offline query library according to the environmental data collected by the XR device;

[0039] S2. Construct a positioning sequence of the current image frame according to the online data collected by the XR device, and reconstruct a 3D point cloud map of the coordinate system of the XR device at the current moment according to the positioning sequence of the current image frame;

[0040] S3. Obtain multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtain the matching point pairs between each frame of the offline image and the 3D point cloud map, and then calculate the pose of each frame of the offline image in the XR device coordinate system at the current moment, and select the optimal offline image from them;

[0041] S4. According to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment, re - locate the pose of the XR device at the current moment.

[0042] During implementation, the offline query library constructed in the offline processing stage of step S1 provides prior data support for the online re - location stages of steps S2 - S4, with high positioning accuracy and low storage occupancy.

[0043] It should be noted that the XR device in this embodiment is provided with a camera and an IMU sensor, and can be an AR glasses, VR glasses or MR glasses; the XR device uses the camera to collect images and outputs an image stream, and outputs the six - degree - of - freedom (Six Degrees of Freedom, abbreviated as 6DOF) pose data corresponding to each frame of the image stream through the positioning interface. Among them, the six - degree - of - freedom pose data includes: three rotational degrees of freedom (pitch, yaw, roll) and three translational degrees of freedom (front - back, left - right, up - down movement along the x, y, z axes), realizing dynamic perception of high - precision position and attitude, which can be represented by a 6 - dimensional vector or a 4×4 transformation matrix; the pose in this embodiment is the 6DOF pose data.

[0044] In step S1, the XR device is used to scan the environment, and the collected environmental data includes the first image stream and the pose corresponding to each frame of the image. Further, in order to obtain high - quality images and improve the processing efficiency, the following pre - processing is performed on the environmental data:

[0045] ① Eliminate the images in the first image stream that do not meet the brightness, blurriness and texture conditions, and select key frames by extracting the feature points of the images.

[0046] It should be noted that the brightness condition is that the proportion of over - exposed or under - exposed areas in the image is less than or equal to the proportion threshold; the blurriness condition is that the Laplacian variance of the image is greater than or equal to the variance threshold; the texture condition is that the proportion of edge pixels in the image is greater than or equal to the edge proportion threshold.

[0047] Specifically, when screening images according to the brightness condition, first perform histogram equalization on the image, then calculate the histogram of the image, and count the number of pixels corresponding to each gray value (0 - 255); then respectively count the number of pixels falling in the over - exposed and under - exposed areas, and divide by the total number of pixels of the image to obtain the proportion of the over - exposed or under - exposed area; finally, compare with the proportion threshold respectively. If any one is greater than the proportion threshold, then eliminate the image.

[0048] Exemplarily, the equalizeHist() method in OpenCV is used to perform histogram equalization on the image, and the calcHist() method is used to calculate the histogram of the image; the overexposed region is where the gray value is greater than 220, and the underexposed region is where the gray value is less than 30; the occupancy ratio threshold is set to 0.25.

[0049] When screening images according to the blurriness condition, the Laplacian variance of the image is calculated using the Laplacian() method in OpenCV; the variance threshold is set to 80.

[0050] When screening images according to the texture condition, the Sobel() or Canny() method in OpenCV is used to detect the edges of the image; the edge occupancy ratio threshold is set to 0.05.

[0051] Furthermore, key frames are selected from the images that meet all the above conditions by extracting the feature points of the images.

[0052] ② Construct an image dictionary, encode the key value according to the pose of the selected key frame. If the key value is not in the corresponding image dictionary, then put the key value into the corresponding image dictionary and mark the key frame.

[0053] It should be noted that in this embodiment, by constructing an image dictionary, the key value key is encoded in sequence according to the pose of the selected key frame. Each key value key corresponds to a grid area in the physical space. Query whether the key value key exists in the image dictionary. If it exists, it means that there is already a key frame in this grid area, and the current key frame is not the required key frame, that is, each grid area only corresponds to the key frame that first has this key value; if it does not exist, then put the key value key into the image dictionary and mark the key frame as the required key frame. The value value corresponding to each key value key in the image dictionary will not be used and can be set to the same data, such as 1.

[0054] Specifically, according to the pose corresponding to the selected key frame, the key value is encoded through the following formula:

[0055] key = str(floor(x / posTh)) + "_" + str(floor(y / posTh)) + "_" + str(floor(z / posTh)) + "_" + str(floor(roll / rotTh)) + "_" + str(floor(pitch / rotTh)) + "_" + str(floor(yaw / rotTh))

[0056] Among them, key represents the key value of the key frame, str(·) represents the operation of converting a numerical value to a character, floor(·) represents the floor operation, posTh and rotTh respectively represent the position segmentation threshold and the rotation segmentation threshold; (x, y, z) respectively represent the positions of the x, y, and z axes in the pose, and (roll, pitch, yaw) respectively represent the roll angle, pitch angle, and yaw angle in the pose.

[0057] Exemplarily, posTh is set to 0.5m and rotTh is set to 10 degrees.

[0058] It can be seen from the encoding formula of the key value that there will be cases where multiple key frames calculate the same key value, but only the first key frame corresponding to this key value is required. This method reduces redundant calculations while ensuring sufficient parallax between key frames and making the key frames evenly distributed in space, improving the accuracy of 3D modeling.

[0059] Furthermore, save the timestamps, corresponding poses, image data, image feature point data, and image vector data of the marked key frames to the XR device storage to form an offline query library. Among them, the image feature point data can be obtained by using a feature point extraction model based on deep learning, such as SurperPoint, R2D2, D2Net, SoSNet, Disk, etc.; it can also be obtained by using traditional feature point extraction methods such as ORB and SIFT. The image vector can be obtained by using an image vector extraction model based on deep learning, such as deep learning models like NetVLad and CosPlace; it can also be obtained by using an image vector encoding method based on the bag of words.

[0060] Steps S2 - S4 are the processes for the user to perform relocalization using the online data collected in real time after turning on the XR device. Among them, step S2 uses the historical data cached online to construct a real-time positioning sequence corresponding to each moment, facilitating the reconstruction of the real-time point cloud map at each moment in step S3, and then in step S4, by calculating the pose of the most similar offline image in the offline query library constructed in step S1 in the XR device coordinate system, the pose of the XR device at each moment is relocalized.

[0061] It should be noted that for the online data collected in real time in step S2, including the second image stream and the pose data corresponding to each frame of the image stream, the images in the second image stream are also screened and key frames are selected therefrom according to the preprocessing method in step S1, and the required key frames are marked by constructing an image dictionary of the online data.

[0062] When constructing the positioning sequence of the current image frame, multiple key frames with the smallest time or distance difference from the current image frame are selected from the marked key frames as multiple historical key frames; the multiple historical key frames and the current image frame are combined to form the positioning sequence of the current image frame. Exemplarily, the positioning sequence of the image frame at the current time t is as shown in Figure 2 ; if 10 historical key frames are selected, then Figure 2 n in is 10, and there are 11 frames of images in the positioning sequence.

[0063] It should be noted that this embodiment does not limit whether the current image frame is a key frame; the smallest time difference from the current image frame means selecting multiple historical key frames with the smallest time difference before the time of the current image frame; the smallest distance difference from the current image frame means that after representing the 6DOF pose with a 6D vector, multiple historical key frames with the smallest difference are selected by comparing the distances between the vectors.

[0064] Furthermore, reconstructing the 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame includes:

[0065] Matching the feature points of the historical key frames and the current image frame in the positioning sequence to obtain 2D-2D matching point pairs; according to the 2D-2D matching point pairs, the camera parameters of the XR device, and the poses corresponding to each frame of image in the positioning sequence, the 3D point cloud map of the XR device coordinate system at the current moment is obtained by using the SFM algorithm.

[0066] It should be noted that the camera parameters on the XR device include the camera internal parameters and distortion parameters; if the current image frame is not a key frame, the feature points of the current image frame are obtained by using the extraction method of image feature point data in step S1; if the current image frame is a key frame, the image feature points have been extracted when selecting the key frame.

[0067] When performing feature point matching, a deep learning feature point matching model can be used, such as deep learning-based feature point matching models like LightGlue and SurperGlue; or methods such as nearest neighbor matching provided by OpenCV can be used to obtain 2D-2D matching point pairs between image pairs.

[0068] Furthermore, the SFM (Structure From Motion) algorithm provided by the COLMAP library is used to reconstruct the 3D point cloud map of the current positioning sequence. Among them, when performing reconstruction, the optimization parameters are fixed to the 6DOF pose of the XR device (that is, the pose corresponding to each frame of the image is not modified), and only the positions of the 3D point cloud obtained by triangulation from the 2D-2D matching point pairs are optimized, reducing the number of optimization variables and the computational complexity. Finally, the 3D point cloud map in the coordinate system of the XR device at the current moment, as well as the 2D feature points matching the 3D point cloud, are output.

[0069] In step S3, the image vector of the current image frame is extracted according to the method of extracting the image vector in step S1, and the top N offline images with the highest similarity to the image vector of the current image frame are searched from the offline query library. Exemplarily, Figure 2 where q1, q2, q3, and q4 are the top 4 offline images with the highest similarity.

[0070] Furthermore, the matching point pairs between each frame of the offline image and the 3D point cloud map are obtained respectively. First, the feature points of each frame of the offline image are respectively matched with the feature points of the current image frame to obtain 2D-2D matching point pairs, and then 2D-3D matching point pairs are obtained according to the 3D point cloud matched by the feature points of the current image frame.

[0071] Furthermore, calculating the pose of each frame of the offline image in the coordinate system of the XR device at the current moment includes: pose solution and positioning optimization.

[0072] Among them, the pose solution is to perform pose solution on each frame of the offline image respectively using the PnP algorithm (Perspective-n-Point algorithm) according to the 2D-3D matching point pairs between each frame of the offline image and the 3D point cloud map, and obtain the initial pose, inlier rate, and the number of actual matching point pairs of each frame of the offline image in the coordinate system of the XR device at the current moment.

[0073] It should be noted that the PnP algorithm will combine methods such as RANSAC (Random Sample Consensus) to screen and process the input 2D-3D matching point pairs, and eliminate abnormal matching point pairs. Therefore, the number of actual matching point pairs is the number of matching point pairs actually used in the PnP algorithm; the inlier rate is obtained by dividing the number of actual matching point pairs by the number of input 2D-3D matching point pairs.

[0074] Exemplarily, the solvePnPRansac() method of OpenCV is used for pose solution.

[0075] The positioning optimization constructs the first relative pose constraint and the second relative pose constraint according to the initial pose of each frame of offline image and the pose in the offline query library; based on the first relative pose constraint and the second relative pose constraint, the PGO algorithm is used to perform positioning optimization on the initial poses of multiple frames of offline images to obtain the optimized pose, the total position optimization error, and the total attitude optimization error; where the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.

[0076] Exemplarily, it is implemented by using the PGO algorithm provided by the CERES library or the COLMAP library.

[0077] Specifically, the first relative pose constraint and the second relative pose constraint are obtained by calculating the relative pose relationship between the initial poses of every two adjacent frames of offline images and the relative pose relationship between the poses in the offline query library respectively after sorting according to the similarity between each frame of offline image and the current image frame.

[0078] The relative pose relationship is obtained through the following formula:

[0079] ,

[0080] where, represents the pose of the th frame of offline image and the pose of the th frame of offline image The relative pose relationship, that is, the pose of the th frame of offline image relative to the th frame of offline image; "-1" represents the inverse operation. The first relative pose constraint is obtained by substituting the initial pose of each frame of offline image into the above formula, and the second relative pose constraint is obtained by substituting the pose of each frame of offline image in the offline query library into the above formula.

[0081] Finally, the offline image with the highest inlier rate is selected as the optimal offline image. According to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment, the pose of the XR device at the current moment is repositioned through the following formula:

[0082] ,

[0083] where, represents the pose of the XR device at the moment, represents the corresponding pose of the optimal offline image in the offline query library, represents the pose of the optimal offline image in the XR device coordinate system at the moment.

[0084] Furthermore, according to the inlier rate and the actual number of matching point pairs of each frame of offline image, as well as the total error of position optimization and the total error of attitude optimization of multiple frames of offline images, the confidence of the pose of the XR device at the current moment is calculated by weighted summation.

[0085] Specifically, the confidence is calculated through the following steps:

[0086] ① According to the inlier rate and the actual number of matching point pairs of each frame of offline image, calculate the average inlier rate and the average number of matching point pairs, and truncate and normalize the average number of matching point pairs according to the maximum number.

[0087] It should be noted that the value of the average inlier rate meanRatio is in the range of 0 to 1 and does not need to be normalized; the average number of matching point pairs is normalized through the following formula:

[0088] ,

[0089] where norm_meanMNum and meanMNum represent the average number of matching point pairs after and before normalization respectively, and MaxNum represents the preset maximum number. Exemplarily, MaxNum is set to 50.

[0090] ② According to the maximum error of position optimization and the maximum error of attitude optimization, perform reverse truncation normalization on the total error of position optimization and the total error of attitude optimization respectively.

[0091] It should be noted that both the total error of position optimization and the total error of attitude optimization are reverse indicators. The smaller the value, the more accurate the positioning result. Therefore, a reverse truncation normalization operation is adopted, and the formula is as follows:

[0092] ,

[0093] ,

[0094] where norm_errPos and errPos represent the total error of position optimization after and before normalization respectively, norm_errRot and errRot represent the total error of attitude optimization after and before normalization respectively, MaxerrPosNum and MaxerrRotNum represent the preset maximum error of position optimization and the maximum error of attitude optimization respectively. Exemplarily, they are set to 0.002 and 0.0005 respectively.

[0095] ③ According to the average inlier rate, the normalized average number of matching point pairs, the total error of position optimization, the total error of attitude optimization and their respective weights, perform weighted summation to obtain the confidence conf, and the formula is as follows:

[0096] ,

[0097] Among them, 、 、 and respectively represent the weights of the average inlier rate, the normalized average number of matching point pairs, the total error of position optimization, and the total error of pose optimization; exemplarily, the weights are set in sequence as: 0.3, 0.2, 0.25, and 0.25.

[0098] The confidence level conf of the pose of the XR device at the current moment obtained by calculation takes values from 0 to 1, and the greater the confidence level, the higher the positioning accuracy.

[0099] Compared with the prior art, a mapless prior XR device relocalization method provided by this embodiment uses the offline query library constructed in the offline processing stage to provide prior data support for the online relocalization stage, has high positioning accuracy, and less storage occupancy; it does not require pre-constructing a 3D point cloud map, has low deployment and storage costs, and strong scalability; by constructing an image dictionary to select key frames, redundant calculations are reduced, and the pose of the image is encoded as the key value of the image dictionary, so that the key frames are evenly distributed in space, improving the accuracy of 3D reconstruction.

[0100] Those skilled in the art can understand that all or part of the processes of implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0101] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. An XR device relocalization method without map prior, characterized in that, It includes the following steps: Construct an offline query library based on the environmental data collected by the XR device; Construct a positioning sequence of the current image frame according to the online data collected by the XR device, and reconstruct a 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame; Obtain multiple offline images with the highest similarity to the current image frame from the offline query library, respectively obtain the matching point pairs between each offline image and the 3D point cloud map, and then calculate the poses of each offline image in the XR device coordinate system at the current moment, and select the optimal offline image from them; Relocate the pose of the XR device at the current moment according to the pose of the optimal offline image in the offline query library and the pose in the XR device coordinate system at the current moment; The environmental data includes a first image stream and the poses corresponding to each frame of the image stream; the online data includes a second image stream and the poses corresponding to each frame of the image stream; Perform the following preprocessing on the environmental data and the online data: Eliminate the images in the first image stream and the second image stream that do not meet the brightness, blurriness, and texture conditions, and select key frames by extracting the feature points of the images; Construct their respective image dictionaries, encode the key values according to the poses of the selected key frames, and if the key value is not in the corresponding image dictionary, put the key value into the corresponding image dictionary and mark the key frame; The construction of the offline query library based on the environmental data collected by the XR device is to save the timestamps, corresponding poses of the marked key frames, the images of the marked key frames in the first image stream, the image feature point data, and the image vector data to the XR device storage after preprocessing the environmental data, forming an offline query library.

2. The method for relocating an XR device without map prior according to claim 1, wherein The construction of the positioning sequence of the current image frame according to the online data collected by the XR device includes: after preprocessing the online data, select multiple key frames with the smallest time or distance difference from the marked key frames as multiple historical key frames; form the positioning sequence of the current image frame by combining the multiple historical key frames and the current image frame.

3. The method for relocating an XR device without map prior according to claim 2, wherein The reconstruction of the 3D point cloud map of the XR device coordinate system at the current moment according to the positioning sequence of the current image frame includes: Match the feature points of the historical key frames and the current image frame in the positioning sequence to obtain 2D-2D matching point pairs; according to the 2D-2D matching point pairs, the camera parameters of the XR device, and the poses corresponding to each frame of the image in the positioning sequence, use the SFM algorithm to obtain the 3D point cloud map of the XR device coordinate system at the current moment.

4. The mapless prior-based XR device relocalization method according to claim 1 or 3, characterized in that, The calculation of the pose of each offline image in the XR device coordinate system at the current moment includes: According to the 2D-3D matching point pairs between each offline image and the 3D point cloud map, perform pose solution on each offline image using the PnP algorithm respectively to obtain the initial pose, inlier rate, and actual number of matching point pairs of each offline image in the XR device coordinate system at the current moment; Construct the first relative pose constraint and the second relative pose constraint based on the initial pose of each frame of offline image and the pose in the offline query library; based on the first relative pose constraint and the second relative pose constraint, use the PGO algorithm to optimize the positioning of the initial poses of multiple frames of offline images, and obtain the optimized pose, the total position optimization error, and the total attitude optimization error; wherein the optimized pose is used as the pose of each frame of offline image in the XR device coordinate system at the current moment.

5. The method for relocating an XR device without map prior according to claim 4, wherein The first relative pose constraint and the second relative pose constraint are obtained by calculating the relative pose relationship between the initial poses of every two adjacent frames of offline images and the relative pose relationship between the poses in the offline query library respectively after sorting according to the similarity between each frame of offline image and the current image frame.

6. The method for relocating an XR device without map prior according to claim 4, wherein The selected optimal offline image is the offline image with the highest inlier rate.

7. The method for relocating an XR device without map prior according to claim 4, wherein Calculate the confidence of the pose of the XR device at the current moment by weighted calculation according to the inlier rate and the actual number of matching point pairs of each frame of offline image, as well as the total position optimization error and the total attitude optimization error of multiple frames of offline images.

8. The method for relocating an XR device without map prior according to claim 1, characterized in that, The key value is encoded through the following formula according to the poses corresponding to the respective selected key frames: key = str(floor(x / posTh)) + "_" + str(floor(y / posTh)) + "_" + str(floor(z / posTh)) + "_" + str(floor(roll / rotTh)) + "_" + str(floor(pitch / rotTh)) + "_" + str(floor(yaw / rotTh)), wherein, key represents the key value of the key frame, str(·) represents the numerical conversion to character operation, floor(·) represents the floor operation, posTh and rotTh respectively represent the position segmentation threshold and the rotation segmentation threshold; x, y, and z respectively represent the positions of the x, y, and z axes in the pose, and roll, pitch, and yaw respectively represent the roll angle, pitch angle, and yaw angle in the pose.

Citation Information

Patent Citations

  • Binocular vision positioning method and binocular vision positioning device for robots, and storage medium

    CN107796397A

  • Positioning method and device and electronic equipment

    CN116862978A