Large space positioning method used in mixed reality scene

By acquiring multiple anchor point images from a mixed reality device, constructing an angle-sensitive magnification model, and iteratively optimizing it in conjunction with a global objective function, the efficiency and accuracy issues of establishing a world coordinate system for mixed reality display devices in large-space applications are solved, achieving efficient long-distance positioning and error suppression.

CN121937679APending Publication Date: 2026-04-28BEIJING ZHIHUI HUANYU TECHNOLOGY CULTURE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHIHUI HUANYU TECHNOLOGY CULTURE CO LTD
Filing Date
2026-01-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing mixed reality display devices struggle to establish a world coordinate system efficiently and accurately in large-space applications, resulting in significant errors when displaying distant content, impacting user experience, and making it difficult to correct planar rotation and height deviations.

Method used

By acquiring multiple anchor point images, the initial world origin and initial orientation of the virtual coordinate system are determined, an angle-sensitive magnified model is constructed, and iterative optimization is performed in conjunction with a global objective function. The world origin and orientation are then recursively corrected using the Levenberg-Marquardt algorithm or extended Kalman filtering.

Benefits of technology

It improves positioning accuracy and deployment efficiency in large-scale, long-distance scenarios, shortens the deployment cycle from several days to hours, and effectively suppresses long-distance errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937679A_ABST
    Figure CN121937679A_ABST
Patent Text Reader

Abstract

The invention provides a large space positioning method used in a mixed reality scene. The method comprises the following steps: determining a world initial origin and an initial orientation of a virtual coordinate system based on a plurality of anchor point images; determining a local pose of an enhanced object in the target site based on the world initial origin and the initial orientation, and further determining a predicted world pose of the enhanced object; based on multiple groups of actual measurement parameters, constructing an angle sensitive amplification model; based on an angle sensitive amplification model and a global objective function, performing iterative optimization on respective fine adjustment amounts of the world initial origin and the initial orientation, and determining an optimal fine adjustment amount; and recursively correcting the world origin and orientation based on the respective optimal fine adjustment amounts of the world initial origin and initial orientation, and outputting the final world origin and orientation when the remote comprehensive error meets a preset threshold value. According to the technical scheme, the positioning precision and deployment efficiency of the mixed reality display equipment in a large-space and long-distance scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of mixed reality technology, and in particular to a large-space positioning method for mixed reality scenarios. Background Technology

[0002] In large-space applications, existing mixed reality (MR) display devices often establish a world coordinate system based on single-image recognition. This approach achieves high display accuracy when displaying near-field content in small spaces, but when displaying distant content in larger spaces, such as at a distance of 40m, the constructed world coordinate system often exhibits significant errors, with an error range of approximately 0. . 40m to 1 . This error of 00m is very noticeable when users view content using mixed reality display devices, impacting the user experience. Furthermore, when displaying distant content in a large space, even minute angular errors in pitch or roll can be amplified by the distance into a perceptible lateral offset.

[0003] Related technologies often employ a process of laser ranging combined with manual placement. This involves extensive manual calculations and repeated adjustments to address situations with large errors, resulting in a deployment cycle that can take more than 10 days. Furthermore, maintaining a unified coordinate reference among staff from different shifts is difficult, affecting the accurate construction of the world coordinate system. Additionally, this method of establishing a world coordinate system based on single-image recognition struggles to simultaneously correct for planar rotation and height deviations, leading to an overall tilt or twist of the world coordinate system relative to real space.

[0004] Therefore, how to efficiently and accurately establish the coordinate system of the mixed reality world for mixed reality display devices in large-space applications has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a large-space positioning method for mixed reality scenes, aiming to solve the technical problem in related technologies that it is difficult to efficiently and accurately establish a world coordinate system for mixed reality display devices when facing large spaces.

[0006] In a first aspect, embodiments of this application provide a large-space positioning method for mixed reality scenes, including: Acquire multiple anchor point images for visual identification of the target site; Based on multiple anchor point images, the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located are determined; Based on the world initial origin and the initial orientation, the local pose of the augmented object within the target site is determined, and based on the local pose, the predicted world pose of the augmented object is determined. Multiple sets of actual measurement parameters are obtained, and an angle-sensitive magnification model is constructed based on the multiple sets of actual measurement parameters. Each set of actual measurement parameters includes: the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin. The angle-sensitive magnification model is used to reflect the lateral position error of the augmented object. Based on the angle-sensitive magnification model and the predetermined global objective function, the fine-tuning amounts of the world initial origin and the initial orientation are iteratively optimized until the global objective function is minimized, at which point the optimal fine-tuning amounts of the world initial origin and the initial orientation are output. The Levenburg-Marquardt algorithm or extended Kalman filter is used to recursively correct the world origin and orientation based on the optimal fine-tuning values ​​of the initial world origin and orientation, until the long-distance integrated error meets the preset threshold, and then the final world origin and orientation are output.

[0007] In one embodiment of this application, optionally, determining the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located based on multiple anchor point images includes: Obtain the center coordinates, rotation quaternion, and recognition confidence of each anchor point image, wherein the rotation quaternion is subjected to homogenization processing; Based on the center coordinates and recognition confidence level of each anchor point image, the world initial origin of the virtual coordinate system in which the display interface in the mixed reality device is located is determined, wherein... ; The initial orientation is determined based on the rotation quaternion and recognition confidence of each anchor point image, wherein, , This represents the initial origin of the world, where i is the anchor point image number. Let the center coordinates of the i-th anchor point image be , Let be the recognition confidence score of the i-th anchor point image, and n be the total number of anchor point images. Indicates the initial orientation. The rotation quaternion of the i-th anchor point image q i The resulting rotation matrix It is the proportion of the recognition confidence of the i-th anchor point image in the total confidence of all anchor point images.

[0008] In one embodiment of this application, optionally, determining the local pose of each augmented object within the target site based on the world origin and the initial orientation includes: Based on the world initial origin and the initial orientation, an estimated coordinate system obtained by fusing multiple anchor point images is determined, and the local pose of each augmented object in the target site in the estimated coordinate system is determined, wherein the local pose includes the estimated coordinates, estimated orientation, and augmented object scale of the augmented object in the estimated coordinate system; Determining the predicted world pose of the augmented object based on the local pose includes: Based on the initial world origin and the initial orientation, a homogeneous transformation matrix from the estimated coordinate system to the world coordinate system is determined. Then, based on the homogeneous transformation matrix and the local pose of each augmented object in the estimated coordinate system, the predicted world pose of each augmented object in the world coordinate system is determined. , 'a' is the enhancement object number. The predicted world pose of the a-th augmented object within the target site in the world coordinate system. This represents the homogeneous transformation matrix from the estimated coordinate system to the world coordinate system, determined by the world initial origin and the initial orientation. Let be the local pose of the a-th augmented object in the estimated coordinate system, where , , These represent the estimated coordinates, estimated orientation, and scale of the a-th augmented object within the estimated coordinate system, respectively.

[0009] Optionally, in one embodiment of this application, the angle-sensitive magnification model is: , Where 'a' is the enhancement object number, Let be the lateral position error of the a-th enhanced object. and For model parameters estimated using calibration data, This represents the predicted distance from the a-th augmented object to the world's initial origin. To predict the difference between the orientation and the observed orientation, It is zero-mean noise.

[0010] Optionally, in one embodiment of this application, before acquiring multiple anchor point images for visual recognition of the target site, the method further includes: The global objective function is preset, wherein the global objective function is: G= , G represents the value of the global objective function, reflecting the overall inconsistency between the prediction results of the estimated coordinate system and all observed data; a is the enhancement object number; Σ p For location observation covariance, Let a be the observation position of the a-th augmented object in the world coordinate system. The predicted position of the a-th augmented object in the estimated coordinate system. This represents the initial origin of the world. This is the fine-tuning amount of the world's initial origin in the current iteration. Indicates the initial orientation, ⊕ indicates small spin synthesis. This is the fine-tuning amount of the initial orientation described in the current iteration. Let be the lateral position error of the a-th enhanced object. For model parameters The estimated value in the current iteration, For model parameters The estimated value in the current iteration, This indicates that the a-th augmented object has model parameters of and The predicted distance from the world's initial origin at that time. The variance is obtained from the propagation of uncertainty.

[0011] In one embodiment of this application, optionally, the variance obtained from the uncertainty propagation is calculated as follows: , , The variance obtained from the propagation of the uncertainty is... Let be the difference between the lateral position error of the a-th augmented object and the model residual. It is the Jacobian matrix, representing the effect of the angle-sensitive magnification model on the model parameters. and The partial derivative of the estimated value in the current iteration and The value at that location, yes The transpose of the matrix, Indicates zero-mean noise. This represents the variance of the zero-mean noise estimated based on multiple sets of the actual measurement parameters.

[0012] In one embodiment of this application, optionally, before determining the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located based on multiple anchor point images, the method further includes: For any of the anchor point images, if a predetermined recognition confidence reduction condition is met, the recognition confidence of the anchor point image is reduced, and a robust cost function is introduced as another loss function besides the global objective function. The robust cost function is used to reduce the negative impact of the recognition confidence reduction condition on anchor point image fusion when the recognition confidence reduction condition occurs.

[0013] In one embodiment of this application, optionally, the identification confidence reduction condition includes: the spatial baseline length of the anchor image and any other anchor image is lower than a preset distance threshold, or the spatial position of the anchor image is approximately collinear with at least one other anchor image.

[0014] In a second aspect, embodiments of this application provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in the first aspect above.

[0015] Thirdly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for performing the method described in the first aspect above.

[0016] The above technical solution addresses the challenge of efficiently and accurately establishing a world coordinate system for mixed reality display devices in large spaces. It involves deploying multiple anchor point images at the target site, fusing these images to calculate the initial world origin and orientation of the virtual coordinate system, and directly matching the virtual and real distances. This allows for multiple retests of the augmented object, establishing an angle-sensitive magnification model. Finally, the world coordinate system is iteratively corrected using this model and a global objective function, recursively refining the world origin and orientation based on their optimal fine-tuning values. This improves the positioning accuracy and deployment efficiency of mixed reality display devices in large, long-distance scenarios, effectively suppressing long-distance errors and reducing the traditionally time-consuming manual deployment process to hours. It can be widely applied in fields requiring high-precision immersive experiences, such as museums, cultural tourism displays, and archaeological site restoration. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1A flowchart of a large-space localization method for a mixed reality scene according to an embodiment of this application is shown; Figure 2 A flowchart of a large-space localization method for a mixed reality scene according to another embodiment of this application is shown; Figure 3 A block diagram of a computer device according to one embodiment of this application is shown; Figure 4 A block diagram of a computer device according to another embodiment of this application is shown. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Figure 1 A flowchart of a large-space positioning method for a mixed reality scene according to an embodiment of this application is shown.

[0021] like Figure 1 As shown, the process of a large-space localization method in a mixed reality scene according to an embodiment of this application includes: Step 102: Obtain multiple anchor point images for visual recognition of the target site.

[0022] The target site refers to the actual physical area displayed using mixed reality technology, such as museum exhibition halls, cultural and tourism scenic spots, large-scale historical sites, or indoor and outdoor immersive exhibition spaces. Anchor point images are the basis for mixed reality technology to visually recognize the target site and establish a world coordinate system. In essence, anchor point images are feature images used by mixed reality display devices to visually recognize and establish the correspondence between the virtual coordinate system and the real world. The content of the anchor point images reflects key features such as the spatial location and orientation of the target site. Therefore, setting multiple anchor point images allows mixed reality display devices to match the correspondence between the virtual coordinate system and the real world from multiple perspectives and positions. Compared to the scheme of establishing the real coordinate system of the real world with a single anchor point image, setting multiple anchor images helps mixed reality display devices improve the accuracy of predicting and setting specific features of the world coordinate system by combining multiple perspectives and positions. Especially when facing large-scale scenes with distant augmented objects, setting multiple anchor images can reduce the errors caused by single-image recognition.

[0023] Step 104: Based on multiple anchor point images, determine the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located.

[0024] Since multiple anchor point images reflect the characteristics of the real world from multiple perspectives and positions, the initial origin and initial orientation of the virtual coordinate system can be determined based on these multiple anchor point images. That is, the characteristics of the virtual coordinate system where the display interface is located in the mixed reality device under the combined effect of multiple perspectives and positions of multiple anchor point images.

[0025] Specifically, the center coordinates, rotation quaternions, and recognition confidence scores of each anchor point image can be obtained, wherein the rotation quaternions are subjected to a direction-invariant process; based on the center coordinates and recognition confidence scores of each anchor point image, the world initial origin of the virtual coordinate system in which the display interface in the mixed reality device is located is determined. Specifically, when performing the direction-invariant process on the rotation quaternions, if... Then let And then normalized. ; The initial orientation is determined based on the rotation quaternion and recognition confidence of each anchor point image, wherein, , This represents the initial origin of the world, where i is the anchor point image number. Let the center coordinates of the i-th anchor point image be , Let be the recognition confidence score of the i-th anchor point image, and n be the total number of anchor point images. Indicates the initial orientation. The rotation quaternion of the i-th anchor point image q i The resulting rotation matrix It is the proportion of the recognition confidence of the i-th anchor point image in the total confidence of all anchor point images.

[0026] Calculating the initial world origin and orientation of the virtual coordinate system establishes a unified spatial reference corresponding to the real world for the entire mixed reality scene. This serves as the basis for subsequent spatial relationship matching, enabling the positioning and correction of augmented objects within the target scene in later steps to be performed under correct spatial relationships. In summary, the initial world origin and orientation of the virtual coordinate system are fundamental to achieving virtual-real alignment and high-precision positioning, contributing to improved accuracy in subsequent world coordinate system estimation.

[0027] Step 106: Based on the world initial origin and the initial orientation, determine the local pose of the augmented object in the target site, and based on the local pose, determine the predicted world pose of the augmented object.

[0028] The local pose of an augmented object reflects its position, rotation, and scaling characteristics relative to the initial world coordinate system. Determining the predicted world pose from the local pose of the augmented object involves transforming its pose in the local coordinate system to the global world coordinate system of the mixed reality display device through a fixed spatial transformation relationship. The resulting predicted world pose reflects the expected position and orientation of the augmented object in the global world coordinate system; in other words, it is the theoretical value of the augmented object in the global world coordinate system, which can be used as a benchmark for subsequent comparison with actual observations.

[0029] Specifically, based on the initial world origin and the initial orientation, an estimated coordinate system obtained by fusing multiple anchor point images can be determined, and the local pose of each augmented object within the target site in the estimated coordinate system can be determined. The local pose includes the estimated coordinates, estimated orientation, and scale of the augmented object within the estimated coordinate system. Thus, a more accurate estimated coordinate system is obtained through multi-image fusion, and the local pose of each augmented object in this estimated coordinate system can be precisely defined. This achieves the initial accurate placement of virtual augmented objects in global space, providing a reliable initial spatial reference for subsequent steps.

[0030] Next, when determining the predicted world pose, based on the initial world origin and the initial orientation, a homogeneous transformation matrix from the estimated coordinate system to the world coordinate system is determined. Then, based on the homogeneous transformation matrix and the local pose of each augmented object in the estimated coordinate system, the predicted world pose of each augmented object in the world coordinate system is determined. , 'a' is the enhancement object number. The predicted world pose of the a-th augmented object within the target site in the world coordinate system. This represents the homogeneous transformation matrix from the estimated coordinate system to the world coordinate system, determined by the world initial origin and the initial orientation. Let be the local pose of the a-th augmented object in the estimated coordinate system, where , , These represent the estimated coordinates, estimated orientation, and scale of the a-th augmented object within the estimated coordinate system, respectively.

[0031] The world pose chain is implemented using the right-multiplication convention, that is, the world basis transformation is applied first. Then apply the local pose enhancement to the object. This is to ensure the consistency of coordinate synthesis.

[0032] The homogeneous transformation matrix from the estimated coordinate system to the world coordinate system, i.e., the fixed spatial transformation relationship mentioned above, is used to transform the pose of the augmented object in the local coordinate system to the global world coordinate system of the mixed reality display device. Through the homogeneous transformation matrix obtained by multi-image fusion, a precise mapping from the local reference coordinate system to the global world coordinate system is achieved. This enables the automatic and accurate conversion of the position, orientation, and size of each augmented object into its theoretically predicted pose in the world coordinate system, thereby facilitating automated and high-precision alignment of virtual and real objects in large-scale, long-distance scenes.

[0033] Step 108: Obtain multiple sets of actual measurement parameters, and construct an angle-sensitive magnification model based on the multiple sets of actual measurement parameters. Each set of actual measurement parameters includes: the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin. The angle-sensitive magnification model is used to reflect the lateral position error of the augmented object.

[0034] Multiple sets of actual measurement parameters are obtained, namely, the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin are measured multiple times.

[0035] Optionally, the number of tests may be more than 20.

[0036] Optionally, the evaluation of the difference between the predicted orientation and the observed orientation (i.e., angular error) and the displacement error adopts the quantile control criterion, and the 90% quantile error threshold can be used as the correction and acceptance standard.

[0037] The angle-sensitive magnification model is as follows: , Where 'a' is the enhancement object number, Let be the lateral position error of the a-th enhanced object. and For model parameters estimated using calibration data, This represents the predicted distance from the a-th augmented object to the world's initial origin. To predict the difference between the orientation and the observed orientation, It is zero-mean noise.

[0038] and It can be estimated from multiple sets of actual measured parameters using weighted least squares or robust regression; optionally, it can be estimated according to the identification confidence level. The distance segmentation weights are then weighted. Alternatively, an online adaptive update mechanism can be used, adjusting the values ​​according to the forgetting factor λ∈(0,1). and The estimates are updated iteratively to adapt to factors such as ambient lighting, occlusion, and texture changes.

[0039] By collecting a large amount of retest data and fitting an angle-sensitive amplification model, we can model the situation where a small angle error is amplified into a significant displacement deviation under long-distance conditions. The angle-sensitive amplification model transforms the high-order nonlinear error caused by the coupling of angle and distance, which is difficult to compensate directly, into a calibrable and predictable error parameter. This provides a precise error correction basis for subsequent global optimization algorithms and helps improve the positioning accuracy of mixed reality display devices under long-distance conditions.

[0040] Step 110: Based on the angle-sensitive magnification model and the predetermined global objective function, iteratively optimize the fine-tuning amounts of the world initial origin and the initial orientation until the global objective function is minimized, and then output the optimal fine-tuning amounts of the world initial origin and the initial orientation.

[0041] In a single iteration, the angle-sensitive magnification model is used to transform the angle deviation, which is difficult to optimize directly, into an error parameter with the same dimensions as the position observation data that can be compared. The global objective function is used to identify the difference between the actual observed position and the predicted position of the augmented object. Based on the equivalent displacement error of each augmented object predicted by the angle-sensitive magnification model, a numerical optimization objective reflecting the degree of world coordinate system correction is constructed. When the numerical optimization objective is minimized, the optimal fine-tuning amount of the world initial origin and initial orientation can be obtained.

[0042] Of course, the global objective function needs to be preset before step 102, wherein the global objective function is: G= , G represents the value of the global objective function, reflecting the overall inconsistency between the prediction results of the estimated coordinate system and all observed data; a is the enhancement object number; Σ p For location observation covariance, Let a be the observation position of the a-th augmented object in the world coordinate system. The predicted position of the a-th augmented object in the estimated coordinate system. This represents the initial origin of the world. This is the fine-tuning amount of the world's initial origin in the current iteration. Indicates the initial orientation, ⊕ indicates small spin synthesis. This is the fine-tuning amount of the initial orientation described in the current iteration. Let be the lateral position error of the a-th enhanced object. For model parameters The estimated value in the current iteration, For model parameters The estimated value in the current iteration, This indicates that the a-th augmented object has model parameters of and The predicted distance from the world's initial origin at that time. The variance is obtained from the propagation of uncertainty.

[0043] The variance obtained from the uncertainty propagation is calculated as follows: , , The variance obtained from the propagation of the uncertainty is... Let be the difference between the lateral position error of the a-th augmented object and the model residual. It is the Jacobian matrix, representing the effect of the angle-sensitive magnification model on the model parameters. and The partial derivative of the estimated value in the current iteration and The value at that location, for and covariance, yes The transpose of the matrix, Indicates zero-mean noise. This represents the variance of the zero-mean noise estimated based on multiple sets of the actual measurement parameters.

[0044] Thus, by iterating the global objective function to its minimum and outputting the optimal fine-tuning amount, the world origin and orientation of the initial world coordinate system can be gradually corrected to the optimal value through multiple iterations. Finally, the optimal correction amount compensates for the remaining systematic deviations after the initial alignment by multi-image fusion, especially the angle errors amplified at long distances in large scenes.

[0045] Step 112: Using the Levenberg-Marquardt algorithm or extended Kalman filter, the world origin and orientation are recursively corrected based on the optimal fine-tuning amounts of the initial world origin and the initial orientation until the long-distance integrated error meets the preset threshold, and then the final world origin and orientation are output.

[0046] The Levenberg-Marquardt algorithm is an iterative optimization algorithm for solving nonlinear least squares problems. It can adaptively adjust between gradient descent and the Gauss-Newton method, and features fast convergence and good stability. The Extended Kalman Filter (EKF), on the other hand, is a recursive state estimation method suitable for nonlinear systems. It achieves real-time optimal estimation and correction of the system state through linearization models and covariance updates.

[0047] By recursively correcting the world origin and orientation using the Levenberg-Marquardt algorithm or the extended Kalman filter, high-precision positioning can be achieved in complex, large-scale, long-distance scenarios by iteratively optimizing the solution to continuously approach the theoretical optimal solution. The reason for using recursive correction instead of direct correction is that iterative correction allows for gradual, small-scale adjustments to parameters based on the latest state of each iteration, avoiding parameter oscillations caused by potentially introducing new instabilities with a large-scale, one-time correction, thus ensuring a smooth correction process.

[0048] The above technical solution involves deploying multiple anchor point images at the target site, fusing these images to calculate the initial world origin and orientation of the virtual coordinate system, and using this data for direct matching of virtual and real distances. This allows for multiple retests of the augmented object, establishing an angle-sensitive magnification model. Finally, the world coordinate system is iteratively corrected using this model and a global objective function, until the world origin and orientation are recursively corrected based on their optimal fine-tuning values. This improves the positioning accuracy and deployment efficiency of mixed reality display devices in large-space, long-distance scenarios, effectively suppressing long-distance errors and reducing the traditionally time-consuming manual deployment process to hours. It can be widely applied in fields requiring high-precision immersive experiences, such as museums, cultural tourism displays, and archaeological site restoration.

[0049] Furthermore, prior to step 104, for any of the anchor point images, if a predetermined recognition confidence reduction condition is met, the recognition confidence of the anchor point image is reduced, and a robust cost function is introduced as another loss function besides the global objective function. The robust cost function is used to reduce the negative impact of the recognition confidence reduction condition on anchor point image fusion when the recognition confidence reduction condition occurs. The recognition confidence reduction condition includes: the spatial baseline length of the anchor point image and any other anchor point image is lower than a preset distance threshold, or the spatial position of the anchor point image is approximately collinear with at least one other anchor point image. That is, when the baseline length of two anchor points is lower than the threshold or multiple anchor points are nearly collinear, the recognition confidence of the corresponding anchor point is reduced, and a robust cost function is introduced to suppress the influence of ill-conditioned geometry on the fusion result.

[0050] Meanwhile, when only two anchor point images are identified, the initial world origin can be the midpoint of the center points of the two anchor point images, and the initial orientation can be the weighted average of the rotation quaternions after the two anchor point images are aligned or the spherical linear interpolation.

[0051] In the above technical solution, by using virtual placement within a mixed reality display device to replace laser ranging and manual deployment, the deployment cycle can be shortened from 10 days to 1 hour, greatly improving deployment efficiency. The virtual placement directly matches the real distance with virtual coordinates, which can also reduce manual conversion errors. During acceptance testing, if the display error at a distance of 40m is less than or equal to 0.2m, the final output world origin and orientation are determined to be valid. Otherwise, the step of reducing the recognition confidence of the anchor point image mentioned above is required. After reducing the recognition confidence of the anchor point image and introducing a robust cost function as another loss function besides the global objective function, ill-conditioned points are removed, and the final world origin and orientation are recalculated.

[0052] Figure 2 A flowchart of a large-space positioning method for a mixed reality scene according to another embodiment of this application is shown.

[0053] like Figure 2 As shown, the process of a large-space positioning method in a mixed reality scene according to another embodiment of this application includes: Step 202: Obtain multiple anchor point images for visual recognition of the target site.

[0054] Step 204: Obtain the center coordinates, rotation quaternion, and recognition confidence of each anchor point image, wherein the rotation quaternion is subjected to homogenization processing.

[0055] Step 206: For any of the anchor point images, if the predetermined recognition confidence reduction condition is met, the recognition confidence of the anchor point image is reduced, and a robust cost function is introduced as another loss function in addition to the global objective function.

[0056] Step 208: Based on the center coordinates and recognition confidence of each anchor point image, determine the world initial origin of the virtual coordinate system in which the display interface in the mixed reality device is located.

[0057] Step 210: Determine the initial orientation based on the rotation quaternion and recognition confidence of each anchor point image.

[0058] Step 212: Based on the world initial origin and the initial orientation, determine the estimated coordinate system obtained by fusing multiple anchor point images, and determine the local pose of each augmented object in the target site in the estimated coordinate system, wherein the local pose includes the estimated coordinates, estimated orientation and augmented object scale of the augmented object in the estimated coordinate system.

[0059] Step 214: Based on the world initial origin and the initial orientation, determine the homogeneous transformation matrix from the estimated coordinate system to the world coordinate system, and based on the homogeneous transformation matrix and the local pose of each augmented object in the estimated coordinate system, determine the predicted world pose of each augmented object in the world coordinate system.

[0060] Step 216: Obtain multiple sets of actual measurement parameters, and construct an angle-sensitive magnification model based on the multiple sets of actual measurement parameters. Each set of actual measurement parameters includes: the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin. The angle-sensitive magnification model is used to reflect the lateral position error of the augmented object.

[0061] Step 218: Based on the angle-sensitive magnification model and the predetermined global objective function, iteratively optimize the fine-tuning amounts of the world initial origin and the initial orientation until the global objective function is minimized, and then output the optimal fine-tuning amounts of the world initial origin and the initial orientation.

[0062] Step 220: Using the Levenberg-Marquardt algorithm or extended Kalman filter, the world origin and orientation are recursively corrected based on the optimal fine-tuning amounts of the initial world origin and the initial orientation until the long-distance integrated error meets the preset threshold, and then the final world origin and orientation are output.

[0063] This technical solution has Figure 1 All the technical effects described in the illustrated embodiments will not be repeated here.

[0064] exist Figure 1 and Figure 2 Based on the illustrated embodiment, in a real-world scenario, a three-image fusion + batch placement + offline correction approach was used to deploy three anchor point images A, B, and C in a 2000-square-meter exhibition hall. The baseline was approximately 40 meters, and the included angle was approximately equilateral triangles to avoid collinearity issues. P was identified. A P B P C With q A q B q CBased on this, the initial world origin and initial orientation are determined. At this point, an initial check is performed to see if the display error at a distance of 40m is less than or equal to 0.20m. If the display error at 40m is less than or equal to 0.20m, the main exhibits can be imported, (t, r, s) are set, and the world pose is generated. Next, at least 20 retest samples are collected for each augmented object, and the results are fitted to obtain... and ,as well as and Finally, the variance obtained from the angle-sensitive magnification model and uncertainty propagation is integrated into the global objective function for optimization iteration. A second check is then performed to verify whether the overall error at a distance of 40m is less than or equal to 0.05m. If the overall error at a distance of 40m is less than or equal to 0.05m, the world coordinate system construction is complete, and the final world origin and orientation are output. The total time is approximately 1 hour.

[0065] In another practical scenario, a dual-anchor image approach is used for rapid deployment with robust weights, employing two anchor point images. Take the midpoint of the center points of the two anchor point images. Two anchor point images are used for quaternion weighting to achieve the same orientation. To address the instability of short baselines, a baseline weight ψ(d) = min(1, d / d0) is introduced to correct ωi←ωiψ(d). The remaining steps are the same as in the previous practical scenario. Finally, it is checked whether the overall error at a distance of 40m is less than or equal to 0.05m. If the overall error at a distance of 40m is less than or equal to 0.05m, the world coordinate system construction is completed, and the final world origin and orientation are output.

[0066] In practical scenarios, online extended Kalman filtering recursion can also be performed, setting the state x=[ P0, ] The observations include positional residuals and equivalent displacement terms, where, , , Observation noise covariance The process noise Q is set based on experience with equipment stability and site disturbance. Online recursion can maintain a steady state of long-distance error during user movement.

[0067] Thus, this invention, with calibrable models and covariance weighted optimization as its core, constructs a robust world basis through multi-graph fusion, replaces manual distance measurement with virtual placement, and completes angle-sensitive adaptive correction with the support of multiple retests. Ultimately, it achieves centimeter-level alignment and hour-level deployment in large-space scenarios, demonstrating significant engineering feasibility and general promotion value.

[0068] In another embodiment, this application provides a computer device, which may be a server, and its internal structure diagram may be as follows. Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it can implement the methods described in any of the above embodiments.

[0069] In one embodiment, this application also provides a computer device, which can be a client, and its internal structure diagram can be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it can implement the methods described in any of the above embodiments.

[0070] Any of the computer devices described in the embodiments of this application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0071] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, etc.

[0072] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes: mixed reality display devices, audio and video players, handheld game consoles, e-books, as well as smart toys, wearable devices, and portable in-vehicle navigation devices.

[0073] (4) Server: A device that provides computing services. The components of a server include a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but because they need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0074] (5) Other electronic devices with data interaction functions.

[0075] Additionally, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which are used to perform the following steps: Acquire multiple anchor point images for visual identification of the target site; Based on multiple anchor point images, the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located are determined; Based on the world initial origin and the initial orientation, the local pose of the augmented object within the target site is determined, and based on the local pose, the predicted world pose of the augmented object is determined. Multiple sets of actual measurement parameters are obtained, and an angle-sensitive magnification model is constructed based on the multiple sets of actual measurement parameters. Each set of actual measurement parameters includes: the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin. The angle-sensitive magnification model is used to reflect the lateral position error of the augmented object. Based on the angle-sensitive magnification model and the predetermined global objective function, the fine-tuning amounts of the world initial origin and the initial orientation are iteratively optimized until the global objective function is minimized, at which point the optimal fine-tuning amounts of the world initial origin and the initial orientation are output. The Levenburg-Marquardt algorithm or extended Kalman filter is used to recursively correct the world origin and orientation based on the optimal fine-tuning values ​​of the initial world origin and orientation, until the long-distance integrated error meets the preset threshold, and then the final world origin and orientation are output.

[0076] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0077] The technical solution of this application has been described in detail above with reference to the accompanying drawings. The technical solution of this application improves the positioning accuracy and deployment efficiency of mixed reality display devices in large-space and long-distance scenarios, effectively suppresses long-distance errors, and shortens the traditional manual deployment process that takes several days to hours. It can be widely used in fields that require high-precision immersive experiences, such as museums, cultural tourism displays, and site restoration.

[0078] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0079] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0080] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0081] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0083] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A large-space positioning method for mixed reality scenarios, characterized in that, include: Acquire multiple anchor point images for visual identification of the target site; Based on multiple anchor point images, the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located are determined; Based on the world initial origin and the initial orientation, the local pose of the augmented object within the target site is determined, and based on the local pose, the predicted world pose of the augmented object is determined. Multiple sets of actual measurement parameters are obtained, and an angle-sensitive magnification model is constructed based on the multiple sets of actual measurement parameters. Each set of actual measurement parameters includes: the difference between the predicted orientation and the observed orientation, the observed world pose, and the predicted distance from the augmented object to the world's initial origin. The angle-sensitive magnification model is used to reflect the lateral position error of the augmented object. Based on the angle-sensitive magnification model and the predetermined global objective function, the fine-tuning amounts of the world initial origin and the initial orientation are iteratively optimized until the global objective function is minimized, at which point the optimal fine-tuning amounts of the world initial origin and the initial orientation are output. The Levenburg-Marquardt algorithm or extended Kalman filter is used to recursively correct the world origin and orientation based on the optimal fine-tuning values ​​of the initial world origin and orientation, until the long-distance integrated error meets the preset threshold, and then the final world origin and orientation are output.

2. The method according to claim 1, characterized in that, The step of determining the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located based on multiple anchor point images includes: Obtain the center coordinates, rotation quaternion, and recognition confidence of each anchor point image, wherein the rotation quaternion is subjected to homogenization processing; Based on the center coordinates and recognition confidence level of each anchor point image, the world initial origin of the virtual coordinate system in which the display interface in the mixed reality device is located is determined, wherein... ; The initial orientation is determined based on the rotation quaternion and recognition confidence of each anchor point image, wherein, , This represents the initial origin of the world, where i is the anchor point image number. Let the center coordinates of the i-th anchor point image be , Let be the recognition confidence score of the i-th anchor point image, and n be the total number of anchor point images. Indicates the initial orientation. The rotation quaternion of the i-th anchor point image q i The resulting rotation matrix It is the proportion of the recognition confidence of the i-th anchor point image in the total confidence of all anchor point images.

3. The method according to claim 2, characterized in that, The determination of the local pose of each augmented object within the target site based on the world origin and the initial orientation includes: Based on the world initial origin and the initial orientation, an estimated coordinate system obtained by fusing multiple anchor point images is determined, and the local pose of each augmented object in the target site in the estimated coordinate system is determined, wherein the local pose includes the estimated coordinates, estimated orientation, and augmented object scale of the augmented object in the estimated coordinate system; Determining the predicted world pose of the augmented object based on the local pose includes: Based on the initial world origin and the initial orientation, a homogeneous transformation matrix from the estimated coordinate system to the world coordinate system is determined. Then, based on the homogeneous transformation matrix and the local pose of each augmented object in the estimated coordinate system, the predicted world pose of each augmented object in the world coordinate system is determined. , 'a' is the enhancement object number. The predicted world pose of the a-th augmented object within the target site in the world coordinate system. This represents the homogeneous transformation matrix from the estimated coordinate system to the world coordinate system, determined by the world initial origin and the initial orientation. Let be the local pose of the a-th augmented object in the estimated coordinate system, where , , These represent the estimated coordinates, estimated orientation, and scale of the a-th augmented object within the estimated coordinate system, respectively.

4. The method according to claim 3, characterized in that, The angle-sensitive magnification model is as follows: , Where 'a' is the enhancement object number, Let be the lateral position error of the a-th enhanced object. and For model parameters estimated using calibration data, This represents the predicted distance from the a-th augmented object to the world's initial origin. To predict the difference between the orientation and the observed orientation, It is zero-mean noise.

5. The method according to claim 4, characterized in that, Before acquiring multiple anchor point images for visual recognition of the target site, the method further includes: The global objective function is preset, wherein the global objective function is: G= , G represents the value of the global objective function, reflecting the overall inconsistency between the prediction results of the estimated coordinate system and all observed data; a is the enhancement object number; Σ p For location observation covariance, Let a be the observation position of the a-th augmented object in the world coordinate system. The predicted position of the a-th augmented object in the estimated coordinate system. This represents the initial origin of the world. This is the fine-tuning amount of the world's initial origin in the current iteration. Indicates the initial orientation, ⊕ indicates small spin synthesis. This is the fine-tuning amount of the initial orientation described in the current iteration. Let be the lateral position error of the a-th enhanced object. For model parameters The estimated value in the current iteration, For model parameters The estimated value in the current iteration, This indicates that the a-th augmented object has model parameters of and The predicted distance from the world's initial origin at that time. The variance is obtained from the propagation of uncertainty.

6. The method according to claim 5, characterized in that, The variance obtained from the uncertainty propagation is calculated as follows: , , The variance obtained from the propagation of the uncertainty is... Let be the difference between the lateral position error of the a-th augmented object and the model residual. It is the Jacobian matrix, representing the effect of the angle-sensitive magnification model on the model parameters. and The partial derivative of the estimated value in the current iteration and The value at that location, yes The transpose of the matrix, Indicates zero-mean noise. This represents the variance of the zero-mean noise estimated based on multiple sets of the actual measurement parameters.

7. The method according to any one of claims 2 to 6, characterized in that, Before determining the world initial origin and initial orientation of the virtual coordinate system in which the display interface in the mixed reality device is located based on multiple anchor point images, the method further includes: For any of the anchor point images, if a predetermined recognition confidence reduction condition is met, the recognition confidence of the anchor point image is reduced, and a robust cost function is introduced as another loss function besides the global objective function. The robust cost function is used to reduce the negative impact of the recognition confidence reduction condition on anchor point image fusion when the recognition confidence reduction condition occurs.

8. The method according to claim 7, characterized in that, The confidence reduction conditions include: the spatial baseline length of the anchor image and any other anchor image is lower than a preset distance threshold, or the spatial position of the anchor image is approximately collinear with at least one other anchor image.

9. A computer device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being configured to cause the processor to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions configured to perform the method as described in any one of claims 1 to 8.