A mixed reality large space interaction method, server, medium and product
Patent Information
- Application Number
- CN202610777550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-21
AI Technical Summary
[0005]本申请提供了一种混合现实大空间互动方法、服务器、介质及产品,用于解决MR大空间近场精细交互中因MR设备全局位姿强制校准导致的共享虚拟交互物体位移、画面撕裂与用户视觉眩晕问题
[0025] 1. By adopting the above technical solution, when multiple MR devices are interacting in the near field, the system switches to a relative coordinate system with one of the MR devices as the origin to complete the pose mapping and coordinate transformation of the shared virtual interactive object. The motion increment of the visual inertial odometry is used to achieve smooth iterative deduction of the pose, avoiding the pose jump caused by the scanning of physical anchor points of the MR devices in the global coordinate system. At the same time, by calculating the asynchronous calibration error vector and combining it with the spatial elastic decay function centered on the shared virtual interactive object to generate the environmental deformation compensation field, while maintaining the stability of the interactive pose and virtual object coordinates in the relative coordinate system, differential deformation compensation is only performed on the global static background environment. This fundamentally solves the problems of displacement of shared virtual interactive objects, screen tearing and user visual dizziness caused by the forced global pose calibration of MR devices in large-space near-field fine interaction. While ensuring the spatial alignment accuracy of multiple MR devices, it greatly improves the continuity, immersion and operation stability of close-range collaborative interaction.
Smart Images

Figure CN122614219A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mixed reality technology, and in particular to a mixed reality large-space interactive method, server, medium and product. Background Technology
[0002] Mixed reality (MR) technology is increasingly being used in fields such as multi-user collaboration, large-space immersive gaming, and industrial simulation training. In large-space interactive scenarios, multiple users wearing MR devices move freely and interact within the same physical space. The system needs to sense the spatial position of each user in real time and synchronously render shared virtual content (such as virtual parts or props that are operated together) in their field of vision to ensure a highly consistent and immersive collaborative interactive experience among multiple users.
[0003] To achieve spatial alignment and interaction among multiple users, relevant mixed reality large-space interaction methods typically rely on a unified global coordinate system for positioning and rendering. Specifically, each MR device independently calculates its absolute pose in physical space using its own spatial positioning system (such as SLAM technology). The system collects this pose information and renders shared virtual objects at fixed positions in the global coordinate system. Simultaneously, to eliminate cumulative drift errors caused by long-term device operation, related technologies deploy multiple visual anchor points (such as QR codes or specific images) in physical space. When an MR device scans these physical anchor points during movement, the system directly uses the absolute coordinates provided by the physical anchor points to forcibly reset or hard-correct the current global pose of the MR device, thereby ensuring that all MR devices maintain baseline alignment in the macroscopic global coordinate system.
[0004] However, when two users are close enough for precise near-field interaction (such as two people passing or assembling a virtual object face-to-face), the above methods reveal significant limitations. In precise near-field interaction, users are extremely sensitive to the accuracy and visual coherence of their relative positions and those of shared virtual objects. If one of the MR devices happens to scan a physical anchor point and triggers a forced correction of the global pose, the absolute coordinates of that MR device in the global coordinate system will instantly change. Since the shared virtual object and the virtual avatars of other users are strongly bound to this global coordinate system, this unilateral instantaneous pose correction will directly cause a drastic relative displacement or screen tearing of the shared virtual object and interactive object in the user's field of vision. This visual abrupt change not only instantly disrupts the continuity of close-range collaboration, causing previously precise interactive actions to fail, but also leads to severe visual dizziness and spatial cognitive dissonance in users, greatly reducing the immersion and usability of large-scale mixed reality interaction. Summary of the Invention
[0005] This application provides a mixed reality large-space interaction method, server, medium, and product to solve the problems of displacement of shared virtual interactive objects, screen tearing, and user visual dizziness caused by forced global pose calibration of MR devices in MR large-space near-field fine interaction.
[0006] In a first aspect, this application provides a mixed reality large-space interaction method applied to a server. The method includes: acquiring a first global pose and a second global pose in a global coordinate system based on physical anchor point calibration outputs of a first MR device and a second MR device, and calculating the spatial relative distance between the first MR device and the second MR device; when the spatial relative distance is less than a preset near-field interaction threshold, constructing a relative coordinate system with the first global pose as the origin, and calculating a relative spatial transformation matrix based on the first global pose and the second global pose; mapping the second global pose to the relative coordinate system through the relative spatial transformation matrix to obtain a second relative pose; mapping the target global coordinates of a shared virtual interactive object located between the first MR device and the second MR device to the relative coordinate system to obtain the target relative coordinates; acquiring a first motion increment and a second motion increment output by the first MR device and the second MR device based on their own visual inertial odometry. Motion increment; in the relative coordinate system, the second relative pose is iteratively updated continuously with a preset step size using the second motion increment to obtain a smooth inferred pose; based on the first motion increment and the target relative coordinates, a shared virtual interactive object is synchronously rendered for the first MR device and the second MR device; when the second MR device scans the physical anchor point and generates the latest global calibration pose, the smooth inferred pose is inversely mapped to the global coordinate system to obtain the expected global pose, and the difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector; a spatial elastic decay function is constructed with the shared virtual interactive object as the center and the spatial relative distance as the variable, and the asynchronous calibration error vector is input into the spatial elastic decay function to generate an environmental deformation compensation field; keeping the smooth inferred pose and the target relative coordinates unchanged in the relative coordinate system, the environmental deformation compensation field is applied to the static background environment rendering in the global coordinate system.
[0007] By adopting the above technical solution, when multiple MR devices are interacting in the near field, the system switches to a relative coordinate system with one of the MR devices as the origin to complete the pose mapping and coordinate transformation of the shared virtual interactive object. The motion increment of the visual inertial odometry is used to achieve smooth iterative deduction of the pose, avoiding pose jumps caused by the scanning of physical anchor points of MR devices in the global coordinate system. At the same time, by calculating the asynchronous calibration error vector and combining it with the spatial elastic decay function centered on the shared virtual interactive object to generate an environmental deformation compensation field, while maintaining the stability of the interactive pose and virtual object coordinates in the relative coordinate system, differential deformation compensation is only performed on the global static background environment. This fundamentally solves the problems of displacement of shared virtual interactive objects, screen tearing, and user visual dizziness caused by the forced global pose calibration of MR devices in large-space near-field fine interaction. While ensuring the spatial alignment accuracy of multiple MR devices, it greatly improves the continuity, immersion, and operational stability of close-range collaborative interaction.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, when the second MR device scans the physical anchor point and generates the latest global calibration pose, the smoothed projection pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector. Specifically, this includes: recording the target timestamp when the second MR device scans the physical anchor point; extracting the historical smoothed projection pose corresponding to the target timestamp in the relative coordinate system; inversely mapping the historical smoothed projection pose to the global coordinate system to obtain the expected global pose; calculating the pose transformation matrix between the latest global calibration pose and the expected global pose; decomposing the pose transformation matrix into a translation error vector and a rotation error vector; and combining them to obtain the asynchronous calibration error vector.
[0009] By adopting the above technical solution, the pose deviation of MR equipment caused by anchor point calibration can be accurately quantified, providing accurate and actual interaction timing error data for subsequent environmental deformation compensation, avoiding timing misalignment and numerical deviation in error calculation, ensuring that the generation basis of the environmental deformation compensation field is real and reliable, improving the accuracy and rationality of the compensation effect, and further ensuring the stability and continuity of the virtual screen during near-field interaction.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, after the steps of mapping the smoothed derivation pose inversely to the global coordinate system to obtain the expected global pose when the second MR device scans the physical anchor point and generates the latest global calibration pose, calculating the difference between the latest global calibration pose and the expected global pose to obtain the asynchronous calibration error vector, the method further includes: extracting the magnitude of the translation error vector in the asynchronous calibration error vector; determining whether the magnitude is greater than a preset error tolerance threshold; if the magnitude is greater than the preset error tolerance threshold, increasing the preset step size when iteratively updating the second relative pose according to the ratio of the magnitude to the preset error tolerance threshold.
[0011] By adopting the above technical solution, a delicate and smooth pose update effect is maintained when the error is small, and the pose convergence speed is accelerated when the error is large to quickly reduce the deviation between the inferred pose and the actual pose. This not only ensures the visual smoothness of conventional near-field interaction, but also copes with the pose offset risk caused by large calibration errors, achieving a balance between smooth pose inference and rapid error correction, and continuously ensuring the rendering stability of shared virtual interactive objects and the collaborative accuracy of multi-device interaction.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, a spatial elastic decay function is constructed with the shared virtual interactive object as the center and the spatial relative distance as the variable. Specifically, this includes: obtaining the spatial bounding box of the shared virtual interactive object; dividing the spatial bounding box outward from its geometric center as the origin into a rigid preservation region, an elastic deformation region, and a background fixed region, wherein the boundary of the rigid preservation region is larger than the boundary of the spatial bounding box; within the elastic deformation region, constructing a decay weight function that monotonically decreases from the boundary of the rigid preservation region to the boundary of the background fixed region, with the spatial relative distance as the independent variable, to obtain the spatial elastic decay function.
[0013] By adopting the above technical solution, we can not only provide an absolutely stable rigidity guarantee for the core area of near-field fine interaction, avoiding deformation of shared virtual interactive objects and interactive space, but also reserve reasonable deformation adjustment space for the surrounding background environment. This accurately matches the actual needs of zero disturbance in the core interactive area and moderate compensation of the surrounding environment in MR near-field interaction, making the distribution and intensity of environmental deformation compensation more in line with the user's visual perception habits, and ensuring the compatibility of core interactive stability and global calibration correction needs from the perspective of spatial structure.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the asynchronous calibration error vector is input into a spatial elastic decay function to generate an environmental deformation compensation field. Specifically, this includes: obtaining the global coordinates of each rendered vertex in the static background environment under the global coordinate system; calculating the vertex distance from each rendered vertex to the geometric center of the spatial bounding box; inputting the vertex distance into a decay weight function to obtain the deformation weight corresponding to each rendered vertex; multiplying the asynchronous calibration error vector with the deformation weight corresponding to each rendered vertex to obtain the local compensation vector of each rendered vertex, and the environmental deformation compensation field is composed of the local compensation vectors of each rendered vertex.
[0015] By adopting the above technical solution, the overall calibration error can be finely and gradually distributed according to spatial distance, so that the environmental deformation compensation is highly adapted to the distance from each vertex to the interaction core, avoiding the abruptness of the screen caused by global uniform correction, making the compensation processing of calibration error smoother and more natural. While realizing global pose calibration correction, it minimizes the interference of background environmental deformation on the user's near-field interaction experience, further improving the continuity and immersion of the virtual screen.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the environmental deformation compensation field is applied to the static background environment rendering in the global coordinate system, specifically including: for rendering vertices located within the rigid preservation region, keeping the global coordinates of the rendering vertices unchanged; for rendering vertices located within the elastic deformation region, vector superimposing the global coordinates of the rendering vertices with the corresponding local compensation vectors to obtain updated vertex coordinates; for rendering vertices located within the fixed background region, performing an overall translation of the global coordinates of the rendering vertices according to the asynchronous calibration error vector to obtain updated vertex coordinates; and performing mesh reconstruction and rendering of the static background environment based on the updated vertex coordinates.
[0017] By adopting the above technical solution, the problems of shared object displacement and screen jump caused by the scanning of physical anchor points of MR devices are completely solved. At the same time, the calibration and correction of the global coordinate system can be completed implicitly, achieving a balance between near-field interaction stability and global spatial positioning accuracy, and improving the overall experience and reliability of MR large-space multi-person interaction.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: real-time monitoring and updating the spatial relative distance between the first MR device and the second MR device; when the spatial relative distance is greater than a preset interaction cancellation threshold, stopping the iterative update in the relative coordinate system, and obtaining the latest global calibration pose of the first MR device and the second MR device, wherein the preset interaction cancellation threshold is greater than a preset near-field interaction threshold; constructing an inverse transition matrix from the relative coordinate system to the global coordinate system based on the latest global calibration pose; using the inverse transition matrix, smoothly mapping the shared virtual interactive object back to the global coordinate system, and canceling the environmental deformation compensation field to restore the normal rendering of the static background environment in the global coordinate system.
[0019] By adopting the above technical solution, seamless switching between near-field interaction mode and global normal mode can be achieved. When the user ends close-range fine interaction and resumes free movement in a large space, the normal rendering state of the global coordinate system is automatically restored. This avoids the accumulation of spatial positioning deviation caused by the long-term effect of the relative coordinate system and the environmental deformation compensation field. It ensures that the MR device always maintains accurate positioning, stable image, and smooth interaction when freely switching between near-field interaction and large-space roaming, and fully adapts to the usage needs of the entire scene of large-space interaction in mixed reality.
[0020] In a second aspect, embodiments of this application provide a server comprising: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the server to perform the method described in the first aspect and any possible implementation thereof.
[0021] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a server, cause the server to perform the method described in the first aspect and any possible implementation thereof.
[0022] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a server, cause the server to perform the method described in the first aspect and any possible implementation thereof.
[0023] Understandably, the server provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0025] 1. By adopting the above technical solution, when multiple MR devices are interacting in the near field, the system switches to a relative coordinate system with one of the MR devices as the origin to complete the pose mapping and coordinate transformation of the shared virtual interactive object. The motion increment of the visual inertial odometry is used to achieve smooth iterative deduction of the pose, avoiding the pose jump caused by the scanning of physical anchor points of the MR devices in the global coordinate system. At the same time, by calculating the asynchronous calibration error vector and combining it with the spatial elastic decay function centered on the shared virtual interactive object to generate the environmental deformation compensation field, while maintaining the stability of the interactive pose and virtual object coordinates in the relative coordinate system, differential deformation compensation is only performed on the global static background environment. This fundamentally solves the problems of displacement of shared virtual interactive objects, screen tearing and user visual dizziness caused by the forced global pose calibration of MR devices in large-space near-field fine interaction. While ensuring the spatial alignment accuracy of multiple MR devices, it greatly improves the continuity, immersion and operation stability of close-range collaborative interaction.
[0026] 2. By adopting the above technical solution, the pose deviation of the MR device caused by anchor point calibration can be accurately quantified, providing accurate and actual interaction timing error data for subsequent environmental deformation compensation, avoiding timing misalignment and numerical deviation in error calculation, ensuring that the generation basis of the environmental deformation compensation field is real and reliable, improving the accuracy and rationality of the compensation effect, and further ensuring the stability and continuity of the virtual screen during near-field interaction.
[0027] 3. By adopting the above technical solution, a delicate and smooth pose update effect is maintained when the error is small, and the pose convergence speed is accelerated when the error is large to quickly reduce the deviation between the inferred pose and the real pose. This ensures the visual smoothness of conventional near-field interaction and can cope with the pose offset risk caused by large calibration errors. It achieves a balance between smooth pose inference and rapid error correction, and continuously ensures the rendering stability of shared virtual interactive objects and the collaborative accuracy of multi-device interaction. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating a mixed reality large-space interactive method in an embodiment of this application;
[0029] Figure 2 This is another flowchart illustrating the mixed reality large-space interaction method in the embodiments of this application;
[0030] Figure 3 This is a schematic diagram of the physical device structure of a server in an embodiment of this application. Detailed Implementation
[0031] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0032] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0033] The following describes the process of the method provided in this implementation. Please refer to [link / reference]. Figure 1 This is a flowchart illustrating a mixed reality large-space interactive method in an embodiment of this application.
[0034] S101. Obtain the first global pose and the second global pose of the first MR device and the second MR device in the global coordinate system based on the physical anchor point calibration output, and calculate the spatial relative distance between the first MR device and the second MR device.
[0035] In this context, the first MR device and the second MR device refer to two independent mixed reality headsets or terminal devices participating in large-space mixed reality interaction, such as the MR glasses worn by user A and user B; physical anchor points refer to reference markers deployed in the real physical space for spatial positioning and error elimination, such as QR codes or specific visual images posted on walls or the ground; the global coordinate system refers to a unified absolute spatial reference system in the entire large-space interactive scene, used to uniformly manage the positions of all devices and virtual objects; the first global pose and the second global pose represent the absolute position and orientation (including three-dimensional coordinates and rotation angles) of the first MR device and the second MR device in the aforementioned global coordinate system; and the spatial relative distance refers to the straight-line distance or relative position difference between the first MR device and the second MR device in the real physical space, such as the physical Euclidean distance between the center points of the two MR devices.
[0036] In large-scale, multi-user mixed reality (MR) interactive scenarios, when multiple users wearing MR devices move freely within the same physical space and engage in initial collaboration or roaming, the server needs to monitor the real-time location status of each user for scene synchronization. Specifically, the server receives location data uploaded by the first and second MR devices in real-time via a wireless network. This location data is output by the two MR devices after scanning physical anchor points in the physical space and calibrating it using their own SLAM (Simultaneous Localization and Mapping) systems combined with the absolute coordinates of the anchor points. After obtaining the first and second global poses of the two devices in a unified global coordinate system, the server extracts their 3D translation coordinates and uses geometric calculation methods such as 3D spatial distance formulas to calculate the current relative spatial distance between the first and second MR devices in real time. This provides a precise data foundation for subsequently determining whether the two users have entered a close-range, fine-grained interactive state.
[0037] S102. When the spatial relative distance is less than the preset near-field interaction threshold, construct a relative coordinate system with the first global pose as the origin, and calculate the relative spatial transformation matrix based on the first global pose and the second global pose.
[0038] Among them, the preset near-field interaction threshold refers to the distance threshold set in advance to determine whether two MR devices have entered a close-range fine interaction state, such as setting it to an interaction radius of 1.5 meters or 2 meters; the relative coordinate system refers to a local three-dimensional space reference system that is re-established with the current position of a specific device as the reference center, detached from the global absolute coordinate system, such as a local coordinate system established with the viewpoint center of the first MR device (user A) as the origin; the relative space transformation matrix refers to a mathematical matrix used to describe the translation and rotation relationship between two different coordinate systems (or two poses), such as a 4x4 homogeneous transformation matrix containing rotation matrix and translation vector, used to achieve accurate transformation of three-dimensional space coordinates.
[0039] This step is triggered when the server detects that two users are getting closer and preparing for refined near-field interactions, such as jointly assembling virtual parts or passing virtual props face-to-face. Specifically, the server compares the real-time calculated relative spatial distance with a preset near-field interaction threshold. If the relative spatial distance is found to be less than the preset threshold, it means that the two users have entered a near-field interaction range that requires extremely high spatial accuracy. To avoid pose jumps and screen tearing caused by a physical anchor point being scanned and forcibly calibrated by an MR device in the global coordinate system, the server immediately implements a coordinate system switching strategy, using the first MR device's current first global pose as the origin (0, 0, 0) and reference orientation to construct a completely new relative coordinate system. Subsequently, the server uses matrix inversion and multiplication to calculate the relative spatial transformation matrix from the first global pose to the second global pose. The relative spatial transformation matrix accurately describes the relative position and posture of the second MR device relative to the first MR device, thereby transforming the spatial relationship between the two from an absolute global binding to a relative local binding.
[0040] S103. By using the relative space transformation matrix, the second global pose is mapped to the relative coordinate system to obtain the second relative pose. The target global coordinates of the shared virtual interactive object located between the first MR device and the second MR device are mapped to the relative coordinate system to obtain the target relative coordinates.
[0041] Here, the second relative pose refers to the position and orientation of the second MR device in a relative coordinate system constructed with the first MR device as the origin, such as the specific orientation, angle, and distance of user B relative to user A; the shared virtual interactive object represents a digital 3D model in a mixed reality scene that is observed, operated, or interacted with by both the first and second MR devices, such as a virtual box that two people are about to lift together; the target global coordinates refer to the absolute position information of the shared virtual interactive object in the original global coordinate system; and the target relative coordinates refer to the local position information of the shared virtual interactive object after it has been transformed to the aforementioned relative coordinate system, such as the specific coordinates of the virtual box relative to user A.
[0042] After successfully constructing the relative coordinate system and calculating the relative spatial transformation matrix, to ensure the absolute stability of the relative position between the virtual content and the user during near-field interaction, the server needs to smoothly migrate all relevant interactive elements to this local coordinate system. Specifically, the server first uses the relative spatial transformation matrix calculated in the previous step to mathematically map the second global pose (such as matrix multiplication), converting it into the second relative pose in the relative coordinate system, thereby establishing the relative positions of the two devices in the local space. Next, the server extracts the target global coordinates of the shared virtual interactive object located between the two devices and which the user is about to interact with. Using the same coordinate system transformation logic, it multiplies these coordinates by the corresponding inverse transformation matrix and maps them to the relative coordinate system to generate the target relative coordinates. Through this series of mapping operations, the server successfully "packages" the two devices and the shared virtual interactive object into a local rigid space unaffected by global anchor point calibration, ensuring that the relative positions of the virtual objects seen by the two users during near-field interaction remain absolutely static and continuous, laying a solid foundation for subsequent smooth and stable close-range fine-grained interaction.
[0043] S104. Obtain the first motion increment and the second motion increment output by the first MR device and the second MR device based on their own visual inertial odometry.
[0044] Visual inertial odometry refers to a low-level tracking system that combines visual sensor data (such as image features captured by a camera) and inertial measurement unit data to estimate the motion trajectory and attitude changes of the MR device itself in a short period of time with high frequency and low latency. For example, the core algorithm module inside the MR headset is used to track the slight head sway. The first motion increment and the second motion increment represent the relative displacement and rotation changes of the first MR device and the second MR device in a very short time interval (such as between two adjacent frames). For example, the head of user A and user B moves forward a few millimeters in a few milliseconds with a slight deflection angle.
[0045] After successfully establishing a relative coordinate system and mapping the MR devices and shared virtual interactive objects to this local space, the server performs this step to maintain a high frame rate of relative position updates during near-field interaction and to completely shield against interference from global anchor point calibration. Specifically, the server temporarily ignores potentially abrupt global absolute positioning data and instead receives, via high-frequency network reception, the first and second motion increments calculated and uploaded in real time by the underlying visual inertial odometry of the first and second MR devices. These motion increment data only reflect the continuous physical motion trends of the devices themselves and do not contain any absolute coordinate corrections from the global reference system. The server collects this pure relative motion data as the most fundamental and coherent motion input source, thereby providing reliable data support for subsequent smooth and seamless pose extrapolation within the relative coordinate system.
[0046] S105. In the relative coordinate system, the second relative pose is iteratively updated with a preset step size using the second motion increment to obtain a smooth deduced pose.
[0047] Among them, the preset step size refers to the maximum allowable change range of a single translation or rotation when updating pose data or performing interpolation calculations. For example, the maximum movement distance limit set when updating each frame is used to control the smoothness of the motion trajectory. Iterative update refers to the algorithm process of continuously accumulating and refreshing the current state based on the historical state of the previous moment and the latest incremental data through iterative calculation. For example, the latest coordinates are calculated by adding the coordinates of the previous frame to the motion increment of the current frame. Smoothly extrapolated pose refers to the smooth and non-jumping current pose of the device in a relative coordinate system without relying on global absolute calibration, purely relying on continuous motion increments. For example, the extremely coherent motion trajectory points extrapolated when user B moves continuously in a relative coordinate system.
[0048] After acquiring high-frequency, continuous motion increment data from both MR devices, the server triggers a pose estimation mechanism in this step to reflect the dynamic changes of the second MR device in a real-time and coherent manner within a relative coordinate system, avoiding visual abrupt changes caused by directly using absolute coordinates. Specifically, the server extracts the second relative pose from the previous moment as a reference state in a relative coordinate system with the first MR device as the reference. Then, using the latest acquired second motion increment, it performs continuous mathematical accumulation and interpolation calculations according to a preset step size. If a single motion increment is too large, the preset step size mechanism will forcibly break it down into multiple small steps for smooth transition, thereby completing a smooth iterative update. Through this motion increment-based trajectory estimation method, the server can calculate the smooth estimated pose of the second MR device, ensuring that during near-field fine interaction, the movement trajectory of the second MR device (user B) in the field of vision of the first MR device (user A) is extremely delicate and smooth, completely eliminating any instantaneous stuttering or displacement that may be caused by global calibration.
[0049] S106. Based on the first motion increment and the target relative coordinates, synchronously render a shared virtual interactive object for the first MR device and the second MR device.
[0050] Synchronous rendering refers to the process of simultaneously generating and displaying a virtual 3D image with highly consistent spatial relationships on the display screens of multiple terminal devices, based on a unified spatial coordinate reference and timestamp. For example, it allows user A and user B to see a virtual object between them at the same time, presenting a perfectly matched perspective angle and lighting state. Shared virtual interactive objects are used to represent digital assets that two people jointly pay attention to, operate, or interact with in this near-field interaction scenario, such as a virtual mechanical part that is being assembled by the two people together.
[0051] After completing the pose smoothing inference in the relative coordinate system, in order to provide both parties involved in near-field interaction with a visually consistent and stable immersive collaborative experience, the server needs to transform the calculated relative spatial relationships into the final user visual image. Specifically, the server strictly uses the relative coordinate system as the rendering benchmark, combined with the first motion increment of the first MR device (used to update the small viewpoint changes of the first MR device in the relative coordinate system) and the previously fixed target relative coordinates of the shared virtual interactive object, to construct a complete 3D rendered scene map for the current frame. Subsequently, the server synchronously sends these precise relative spatial data, inferred poses, and rendering commands to the first and second MR devices (or sends the video stream after rendering in the cloud). Because the entire rendering process relies entirely on the relative coordinate system and continuous motion increments, the relative position of the shared virtual interactive object in the field of vision of both people is firmly locked, thereby ensuring the absolute continuity and stability of the image during close-range fine-tuning operations, providing users with a high-quality mixed reality experience without any tearing.
[0052] S107. When the second MR device scans the physical anchor point and generates the latest global calibration pose, the smoothed derivation pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector.
[0053] Among them, the latest global calibration pose represents the current absolute position and attitude output by the second MR device after scanning the physical anchor point and forcibly resetting or correcting it by combining the absolute coordinates of the anchor point with the SLAM system; the smooth extrapolation pose refers to the non-jumping device pose calculated purely by continuous motion increment iteration in the relative coordinate system; the expected global pose is used to represent the theoretical absolute position that should be in the global coordinate system according to the smooth extrapolation trajectory if no forced calibration occurs; the asynchronous calibration error vector refers to the difference between the latest global calibration pose and the expected global pose in three-dimensional space, such as a mathematical vector containing translational and rotational deviations, used to quantify the degree of spatial jump caused by this forced calibration.
[0054] During near-field interaction, when the second MR device accidentally scans a physical anchor point while moving, triggering the absolute coordinate correction mechanism, the server performs this step to quantify the error. Specifically, the server first captures the latest global calibration pose generated by the second MR device due to scanning the physical anchor point. To assess the degree of abrupt change caused by this forced calibration, the server reverse-maps the smoothed projection pose, which maintains coherent interaction in the relative coordinate system, back to the global coordinate system using the inverse of the previous relative space transformation matrix, thereby calculating the expected global pose. Next, the server performs matrix subtraction or vector subtraction between the latest global calibration pose and the expected global pose. By comparing these two poses, the server can accurately extract the translational and rotational deviations caused by asynchronous calibration, thus obtaining the asynchronous calibration error vector, providing accurate data for subsequent smoothing compensation.
[0055] Optionally, under normal circumstances, when the second MR device scans the physical anchor point and generates the latest global calibration pose, the smoothed extrapolation pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector. This can be achieved in the following ways, without limitation: record the target timestamp when the second MR device scans the physical anchor point; extract the historical smoothed extrapolation pose corresponding to the target timestamp in the relative coordinate system; inversely map the historical smoothed extrapolation pose to the global coordinate system to obtain the expected global pose; calculate the pose transformation matrix between the latest global calibration pose and the expected global pose, decompose the pose transformation matrix into translation error vector and rotation error vector, and combine them to obtain the asynchronous calibration error vector.
[0056] Among them, the target timestamp represents the precise moment mark at which the visual sensor of the second MR device successfully identified and resolved the physical anchor point; the historical smoothed pose refers to the specific pose data in the relative coordinate system that precisely matches the target timestamp in the time dimension among the continuous pose records continuously cached by the server in chronological order and calculated frame by frame through the motion increment of visual inertial odometry; the inverse mapping refers to the mathematical operation process of using the inverse matrix of the previously constructed relative space transformation matrix to reverse the pose data in the relative coordinate system back to the global coordinate system; the expected global pose is used to represent the theoretical absolute position that should be in the global coordinate system according to the smoothed trajectory. Attitude; The pose transformation matrix is a 4x4 homogeneous transformation matrix that describes the spatial mapping relationship between the latest global calibration pose and the expected global pose, containing complete rotation and translation transformation information between the two; The translation error vector represents the component extracted from the pose transformation matrix that only involves the three-dimensional spatial position offset, such as the displacement deviation values in the X, Y, and Z axes; The rotation error vector represents the component extracted from the pose transformation matrix that only involves the spatial orientation deflection, such as the angular deviation values around each axis; The asynchronous calibration error vector is a complete mathematical vector formed by combining the translation error vector and the rotation error vector, used to fully quantify the degree of spatial jump caused by this forced calibration.
[0057] During near-field interaction, when the second MR device accidentally scans a physical anchor point due to the user's natural movement, triggering a forced correction of the absolute coordinates of the SLAM system, the server needs to precisely quantify the spatial abrupt change caused by this calibration. Specifically, the server first captures the precise moment when the second MR device scans the physical anchor point and records it as the target timestamp to ensure that subsequent error calculations are strictly aligned with the calibration event in the time dimension. Subsequently, the server performs a precise match search in the continuously maintained smoothed pose cache queue, using the target timestamp as the retrieval key, to extract the historical smoothed pose of the second MR device at that moment, derived purely from motion increments in the relative coordinate system. If the target timestamp falls between two adjacent cache frames, the server reconstructs the corresponding intermediate pose through interpolation. Next, the server uses the inverse matrix of the previously calculated relative spatial transformation matrix to perform a reverse mapping on the historical smoothed pose, transforming it from the relative coordinate system back to the global coordinate system, thereby obtaining the expected global pose, i.e., the theoretical position the device should be in global space if forced calibration does not occur. Finally, the server calculates the pose transformation matrix between the latest global calibration pose and the expected global pose, and decomposes the pose transformation matrix into translation error vector and rotation error vector. The two are combined to form a complete asynchronous calibration error vector, thereby providing accurate and time-matched error data for the subsequent generation of environmental deformation compensation field.
[0058] S108. Construct a spatial elastic decay function centered on the shared virtual interactive object and with spatial relative distance as the variable. Input the asynchronous calibration error vector into the spatial elastic decay function to generate an environmental deformation compensation field.
[0059] Among them, the spatial elastic decay function is used to represent a mathematical model in which the weight value of a specific spatial location gradually decreases with distance from the periphery. For example, the weight is 0 in the core interaction area (no deformation) and gradually increases to 1 (full compensation) in the background area. The environmental deformation compensation field refers to the set of local compensation vectors applied to each vertex in the entire three-dimensional space, calculated based on the asynchronous calibration error vector and the spatial elastic decay function. For example, it is used to guide the background environment mesh to undergo flexible distortion to absorb calibration errors.
[0060] After calculating the asynchronous calibration error vector generated by the physical anchor points scanned by the MR devices, the server triggers this step to absorb this error in the global environment without affecting the core near-field interaction area. Specifically, the server first constructs a spatial elastic decay function with the geometric center of the bounding box of the shared virtual interactive object as the origin and the spatial relative distance between the first and second MR devices as the independent variable. The spatial elastic decay function divides the space into a rigid maintenance region (core interaction area), an elastic deformation region (transition area), and a background fixed region (outer area). Subsequently, the server inputs the asynchronous calibration error vector calculated in the previous step into this spatial elastic decay function. The server traverses each rendering vertex in the static background environment under the global coordinate system, assigns different deformation weights to each rendering vertex based on its distance from the shared virtual interactive object using the spatial elastic decay function, and multiplies the asynchronous calibration error vector by these deformation weights to ultimately generate a continuous and smooth environmental deformation compensation field.
[0061] Optionally, in general, the spatial elastic decay function centered on the shared virtual interactive object and with spatial relative distance as the variable can be constructed in the following way, without limitation: obtain the spatial bounding box of the shared virtual interactive object; with the geometric center of the spatial bounding box as the origin, divide it outwards into a rigid preservation region, an elastic deformation region, and a background fixed region, with the boundary of the rigid preservation region being larger than the boundary of the spatial bounding box; within the elastic deformation region, construct a decay weight function with spatial relative distance as the independent variable, which monotonically decreases from the boundary of the rigid preservation region to the boundary of the background fixed region, to obtain the spatial elastic decay function.
[0062] The bounding box represents the smallest axis-aligned cuboid that can completely enclose the entire geometry of the shared virtual interactive object, used to quickly define the area occupied by the shared virtual interactive object in three-dimensional space, such as a rectangular cube that perfectly encloses the virtual box; the geometric center refers to the three-dimensional space center point determined by the intersection of the diagonals of the bounding box; the rigid preservation region refers to the core spatial region extending outward from the geometric center to a certain boundary range. All rendering elements within this region maintain absolutely unchanged coordinates during the compensation process, and its boundary is set to be larger than the boundary of the bounding box to ensure that the shared virtual interactive object and its adjacent interactive operation space are not affected. Any deformation interference; the elastic deformation region refers to the annular transition space zone located outside the rigid preservation region and inside the background fixed region. Rendered elements within this region will receive gradual deformation compensation from zero to full based on their distance; the background fixed region refers to the outermost spatial range outside the elastic deformation region. Rendered elements within this region will receive full compensation correction according to the asynchronous calibration error vector; the attenuation weight function is a continuous mathematical function within the elastic deformation region, with the relative spatial distance as the independent variable, whose output value monotonically increases from zero at the boundary of the rigid preservation region to one at the boundary of the background fixed region. It is used to control the gradual transition law of deformation compensation intensity as a function of distance.
[0063] After calculating the asynchronous calibration error vector, the server needs to construct a mathematical model that can distribute the error in a gradient according to spatial distance to achieve a differentiated correction strategy of zero disturbance in the core interaction area and full compensation in the peripheral background area. Specifically, the server first reads the 3D model data of the shared virtual interactive object, calculates the minimum axis-aligned bounding box that can completely enclose all its vertices, and takes the geometric center of the bounding box as the reference origin for the division of the entire spatial region. Subsequently, the server divides three concentric spatial regions outward from this geometric center: the innermost layer is a rigid preservation region, the boundary of which is set to be larger than the boundary of the spatial bounding box by a certain margin to ensure that the shared virtual interactive object body and the surrounding space that the user's hand may touch during close operation are completely covered within the rigid protection range; the middle layer is an elastic deformation region, serving as a smooth transition buffer between rigid preservation and full compensation; and the outermost layer is a background fixed region, covering the peripheral environment far from the core of the interaction. Finally, within the elastic deformation region, the server constructs a decay weight function that monotonically decreases from the boundary of the rigid maintenance region to the boundary of the background fixed region, with the spatial relative distance as the independent variable. This makes the compensation weight approach zero when it is close to the interaction core and gradually increase to one when it is far away from the interaction core. This yields a complete spatial elastic decay function, providing a spatial distribution rule for subsequently transforming the asynchronous calibration error vector into a continuous and smooth environmental deformation compensation field.
[0064] Optionally, in general, the asynchronous calibration error vector can be input into the spatial elastic decay function to generate the environmental deformation compensation field in the following ways, which are not limited here: obtain the global coordinates of each rendered vertex in the static background environment in the global coordinate system; calculate the vertex distance from each rendered vertex to the geometric center of the spatial bounding box; input the vertex distance into the decay weight function to obtain the deformation weight corresponding to each rendered vertex; multiply the asynchronous calibration error vector by the deformation weight corresponding to each rendered vertex to obtain the local compensation vector of each rendered vertex, and the environmental deformation compensation field is composed of the local compensation vectors of each rendered vertex.
[0065] In this context, rendering vertices represent the basic geometric points that constitute the 3D mesh model of the static background environment. Each rendering vertex carries 3D spatial coordinate information and is the smallest spatial positioning unit when the graphics rendering engine draws the image, such as a corner point of a triangular facet on a virtual wall mesh. Global coordinates refer to the absolute 3D position values of each rendering vertex in a unified global coordinate system. Vertex distance refers to the Euclidean straight-line distance between each rendering vertex and the geometric center of the shared virtual interactive object's bounding box, used to determine the vertex's position in the spatial elastic decay function and the corresponding compensation intensity. Deformation weight refers to the weight of the area after inputting the vertex distance into the decay weight function. The output scalar value between zero and one is used to control the deformation compensation ratio that the rendered vertex should bear. The closer the value is to zero, the weaker the compensation; the closer it is to one, the stronger the compensation. The local compensation vector is a three-dimensional vector obtained by performing a scalar multiplication operation between the asynchronous calibration error vector and the corresponding deformation weight. It is used to indicate the specific displacement direction and displacement amount that the rendered vertex needs to undergo in three-dimensional space. The environmental deformation compensation field is a continuous spatial vector field composed of the local compensation vectors of all rendered vertices in the static background environment under the global coordinate system. It is used to describe how the background environment should undergo differentiated flexible deformation to absorb asynchronous calibration errors.
[0066] After constructing the spatial elastic decay function, the server needs to combine this function with the asynchronous calibration error vector to calculate the specific compensation amount for each location in the background environment point by point, thereby generating complete deformation data that can be directly applied to the rendering pipeline. Specifically, the server first traverses the rendering engine and obtains the global coordinates of all rendered vertices in the static background environment 3D mesh in the global coordinate system. Then, the server calculates the Euclidean distance from each rendered vertex to the geometric center of the shared virtual interactive object's bounding box, obtaining the corresponding vertex distance. Next, the server substitutes each vertex distance into the previously constructed decay weight function for evaluation, obtaining the deformation weight corresponding to each rendered vertex. Vertices within the rigid preservation region receive weights close to zero, vertices within the elastic deformation region receive intermediate weights that increase with distance, and vertices within the fixed background region receive weights close to one. Finally, the server performs scalar multiplication operations on the asynchronous calibration error vector with the deformation weight of each rendered vertex to generate independent local compensation vectors for each rendered vertex. The local compensation vectors of all rendered vertices are brought together to form a complete environmental deformation compensation field, providing vertex-level accurate compensation data for subsequent differential deformation rendering of static background environments.
[0067] S109. Keep the smooth inference pose and target relative coordinates unchanged in the relative coordinate system, and apply the environmental deformation compensation field to the static background environment rendering in the global coordinate system.
[0068] Among them, static background environment rendering is used to represent the graphic drawing process of non-core elements such as distant views, virtual room walls, and decorations in mixed reality scenes that do not participate in direct interaction, such as rendering virtual mountains in the distance or exhibition hall backgrounds.
[0069] After successfully generating the environmental deformation compensation field, the server needs to apply the error correction to the user's visual experience while ensuring that the current fine-grained near-field interaction is not interrupted by sudden changes. Specifically, when issuing rendering commands, the server strictly maintains the smooth inference pose within the relative coordinate system and the target relative coordinates of shared virtual interactive objects. This ensures that when users observe each other and manipulate virtual objects at close range, all elements within the core field of view remain absolutely rigid, stable, and visually coherent, without any sense of tearing. At the same time, the server applies the generated environmental deformation compensation field to the static background environment rendering in the global coordinate system. For background vertices outside the core area, the server fine-tunes their coordinates or performs a global translation based on the local compensation vectors in the environmental deformation compensation field. Through this differentiated rendering strategy of rigidly locking the core area and flexibly deforming the outer background, the server implicitly completes the global pose calibration and correction almost imperceptibly to the user, completely solving the problems of screen jumps and visual dizziness in large-space near-field interaction in MR.
[0070] Optionally, in general, applying the environmental deformation compensation field to the static background environment rendering in the global coordinate system can be achieved in the following ways, without limitation: For rendering vertices located within the rigid preservation region, keep the global coordinates of the rendering vertices unchanged; for rendering vertices located within the elastic deformation region, superimpose the global coordinates of the rendering vertices with the corresponding local compensation vectors to obtain the updated vertex coordinates; for rendering vertices located within the fixed background region, translate the global coordinates of the rendering vertices according to the asynchronous calibration error vector to obtain the updated vertex coordinates; based on the updated vertex coordinates, perform mesh reconstruction and rendering of the static background environment.
[0071] By adopting the above technical solution, when multiple MR devices are interacting in the near field, the system switches to a relative coordinate system with one of the MR devices as the origin to complete the pose mapping and coordinate transformation of the shared virtual interactive object. The motion increment of the visual inertial odometry is used to achieve smooth iterative deduction of the pose, avoiding pose jumps caused by the scanning of physical anchor points of MR devices in the global coordinate system. At the same time, by calculating the asynchronous calibration error vector and combining it with the spatial elastic decay function centered on the shared virtual interactive object to generate an environmental deformation compensation field, while maintaining the stability of the interactive pose and virtual object coordinates in the relative coordinate system, differential deformation compensation is only performed on the global static background environment. This fundamentally solves the problems of displacement of shared virtual interactive objects, screen tearing, and user visual dizziness caused by the forced global pose calibration of MR devices in large-space near-field fine interaction. While ensuring the spatial alignment accuracy of multiple MR devices, it greatly improves the continuity, immersion, and operational stability of close-range collaborative interaction.
[0072] The following provides a more detailed description of the process of the method provided in this implementation. Please refer to [link / reference]. Figure 2 This is another flowchart illustrating the mixed reality large-space interaction method in the embodiments of this application.
[0073] S201. Obtain the first global pose and the second global pose of the first MR device and the second MR device in the global coordinate system based on the physical anchor point calibration output, and calculate the spatial relative distance between the first MR device and the second MR device.
[0074] For details, please refer to step S101, which will not be repeated here.
[0075] S202. When the spatial relative distance is less than the preset near-field interaction threshold, construct a relative coordinate system with the first global pose as the origin, and calculate the relative spatial transformation matrix based on the first global pose and the second global pose.
[0076] For details, please refer to step S102, which will not be repeated here.
[0077] S203. By using the relative space transformation matrix, the second global pose is mapped to the relative coordinate system to obtain the second relative pose. The target global coordinates of the shared virtual interactive object located between the first MR device and the second MR device are mapped to the relative coordinate system to obtain the target relative coordinates.
[0078] For details, please refer to step S103, which will not be repeated here.
[0079] S204. Obtain the first motion increment and the second motion increment output by the first MR device and the second MR device based on their own visual inertial odometry.
[0080] For details, please refer to step S104, which will not be repeated here.
[0081] S205. In the relative coordinate system, the second relative pose is iteratively updated with a preset step size using the second motion increment to obtain a smooth deduced pose.
[0082] For details, please refer to step S105, which will not be repeated here.
[0083] S206. Based on the first motion increment and the target relative coordinates, synchronously render a shared virtual interactive object for the first MR device and the second MR device.
[0084] For details, please refer to step S106, which will not be repeated here.
[0085] S207. When the second MR device scans the physical anchor point and generates the latest global calibration pose, the smoothed derivation pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector.
[0086] For details, please refer to step S107, which will not be repeated here.
[0087] S208. Extract the magnitude of the translation error vector in the asynchronous calibration error vector; determine whether the magnitude is greater than the preset error tolerance threshold; if the magnitude is greater than the preset error tolerance threshold, increase the preset step size when iterating the second relative pose according to the ratio of the magnitude to the preset error tolerance threshold.
[0088] Among them, the translation error vector represents the part of the asynchronous calibration error vector that only involves the movement deviation of the three-dimensional spatial position (X, Y, Z axes), and is used to isolate the interference of rotation error; the modulus is used to represent the absolute straight-line distance of the translation error vector in three-dimensional space, for example, the calculated real physical offset is 0.5 meters; the preset error tolerance threshold refers to the maximum translation deviation distance that can be tolerated without causing visual dizziness or interaction interruption in the user, for example, a safety buffer limit of 0.1 meters.
[0089] Specifically, the server first decouples the asynchronous calibration error vector containing multi-dimensional information, extracting the rotation component and the translation error vector separately. It then calculates the magnitude of this translation error vector using the Euclidean distance formula, obtaining a scalar form of the absolute deviation distance. Subsequently, the server rigorously compares the calculated magnitude with a preset error tolerance threshold. If the magnitude is within the preset tolerance threshold, the current deviation is small, and the original smooth derivation rhythm can be maintained. However, if the magnitude exceeds the preset tolerance threshold, it indicates a significant deviation between the actual physical position and the deduced position. Continuing to use the conventional step size for slow compensation will lead to prolonged virtual-real misalignment. Therefore, the server calculates the specific ratio of the magnitude to the preset error tolerance threshold (i.e., the over-limit factor) and uses it as a dynamic gain coefficient. Based on this, the server proportionally amplifies the preset step size of the second relative pose in subsequent iterative updates. Through this dynamic acceleration mechanism, the server can accelerate the smooth tracking speed of pose within the limit range that the user cannot perceive obvious lag, so that the virtual performance of the second MR device in the relative coordinate system can converge more quickly and fit its real physical space state.
[0090] S209. Construct a spatial elastic decay function centered on the shared virtual interactive object and with spatial relative distance as the variable. Input the asynchronous calibration error vector into the spatial elastic decay function to generate an environmental deformation compensation field.
[0091] For details, please refer to step S108, which will not be repeated here.
[0092] S210. Keep the smooth inference pose and target relative coordinates in the relative coordinate system unchanged, and apply the environmental deformation compensation field to the static background environment rendering in the global coordinate system.
[0093] For details, please refer to step S109, which will not be repeated here.
[0094] S211. Monitor and update the spatial relative distance between the first MR device and the second MR device in real time.
[0095] Real-time monitoring refers to the dynamic process of continuously reading and calculating data status at an extremely high time frequency (such as 60 times per second or higher).
[0096] Specifically, the server will calculate the difference in three-dimensional coordinates between the first MR device and the second MR device in space in real time based on the pose data received from the two MR devices at high frequency, and use the Euclidean distance formula to calculate the precise straight-line distance between the two devices, which will serve as the core trigger condition for subsequent judgment on whether to exit the near-field interaction mode.
[0097] S212. When the spatial relative distance is greater than the preset interaction cancellation threshold, stop the iterative update in the relative coordinate system and obtain the latest global calibration pose of the first MR device and the second MR device. The preset interaction cancellation threshold is greater than the preset near-field interaction threshold.
[0098] Among them, the preset interaction termination threshold refers to the pre-set spatial distance threshold used to determine that two users have ended close collaboration and moved away from each other, such as a physical distance of 3 meters; iterative update refers to the algorithm process of accumulating and extrapolating poses based on continuous motion increments in a relative coordinate system; the latest global calibration pose is used to represent the true absolute coordinates and pose of the MR device in a unified absolute physical space after correction by the SLAM system and physical anchor points.
[0099] Specifically, the server compares the currently calculated relative spatial distance with a preset interaction cancellation threshold. A hysteresis interval design similar to a Schmitt trigger is used here (i.e., the preset interaction cancellation threshold is greater than a preset near-field interaction threshold, such as 3 meters being greater than 1.5 meters). This is to prevent screen flickering caused by the server frequently switching between the relative and global coordinate systems when the user hovers at the edge of the critical distance. Once it is confirmed that the relative spatial distance is greater than the preset interaction cancellation threshold, the server immediately issues an instruction to stop the pose iteration update based on motion increments in the local relative coordinate system, and simultaneously requests the latest absolute calibration pose of the first and second MR devices in the global large space, preparing data for the subsequent coordinate system switch.
[0100] S213. Based on the latest global calibration pose, construct an inverse transition matrix from the relative coordinate system to the global coordinate system.
[0101] The reverse transition matrix is a mathematical transformation matrix that includes translation, rotation, and scaling parameters. It is used to reverse the mapping of three-dimensional coordinate points in the local relative coordinate system back to the global absolute coordinate system, such as a 4x4 homogeneous transformation matrix.
[0102] Specifically, the server uses the latest global calibration pose of the first MR device as the reference origin, and combines it with its historical state in the relative coordinate system. Through matrix inversion and coordinate transformation algorithms, it calculates an accurate inverse transition matrix. This inverse transition matrix acts as a mathematical bridge between local and global spaces, precisely describing how to seamlessly and accurately restore any 3D point in the relative coordinate system to its corresponding absolute position in the global physical space after rotation and translation operations.
[0103] S214. Using the inverse transition matrix, smoothly map the shared virtual interactive objects back to the global coordinate system, and cancel the environment deformation compensation field to restore the normal rendering of the static background environment in the global coordinate system.
[0104] Among them, the environmental deformation compensation field is used to represent the flexible distortion data field applied to the surrounding background during near-field interaction in order to cover up global calibration errors; the static background environment refers to non-core elements such as distant views and walls in the mixed reality scene that do not participate in direct interaction; and the regular rendering represents the standard graphics drawing process that cancels all local distortions and special relative locking and is performed entirely according to the absolute pose of the device in the global coordinate system.
[0105] Specifically, the server first uses an inverse transition matrix to perform matrix multiplication on the current relative coordinates of the shared virtual interactive object, smoothly and accurately mapping and anchoring it to a specific absolute physical position in the global coordinate system. This ensures that the shared virtual interactive object does not instantly teleport due to coordinate system switching. Simultaneously, the server sends instructions to the rendering engine to gradually reduce the weight of the environmental deformation compensation field until it is completely removed, restoring the background mesh, which had undergone slight distortion to absorb errors, to a rigid state. Finally, the server controls the rendering pipelines of all devices to restore the static background environment to normal rendering in the global coordinate system. At this point, the server perfectly releases near-field relative locking, allowing users to continue freely roaming in the large global space, with the entire exit process visually seamless and without any abrupt changes or distortions.
[0106] The server in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of a server in an embodiment of this application.
[0107] It should be noted that, Figure 3 The server structure shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0108] like Figure 3 As shown, the server includes a CPU 301, which can perform various appropriate actions and processes based on a program stored in the read-only memory ROM 302 or a program loaded from the storage section 308 into the random access memory RAM 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0109] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.
[0110] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by CPU 301, it performs the various functions defined in the present invention.
[0111] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0113] Specifically, the server in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the mixed reality large-space interactive method provided in the above embodiment.
[0114] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above embodiments; or it may exist independently and not assembled into the server. The storage medium carries one or more computer programs that, when executed by a processor of the server, cause the server to implement the mixed reality large-space interactive method provided in the above embodiments.
[0115] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0116] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0117] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A mixed reality large-space interactive method, characterized in that, Applied to a server, the method includes: The first global pose and the second global pose of the first MR device and the second MR device in the global coordinate system based on the physical anchor point calibration output are obtained, and the spatial relative distance between the first MR device and the second MR device is calculated. When the spatial relative distance is less than a preset near-field interaction threshold, a relative coordinate system is constructed with the first global pose as the origin, and the relative spatial transformation matrix is calculated based on the first global pose and the second global pose. The second global pose is mapped to the relative coordinate system through the relative space transformation matrix to obtain the second relative pose. The target global coordinates of the shared virtual interactive object located between the first MR device and the second MR device are mapped to the relative coordinate system to obtain the target relative coordinates. Acquire the first motion increment and the second motion increment output by the first MR device and the second MR device based on their own visual inertial odometry. In the relative coordinate system, the second relative pose is iteratively updated continuously with a preset step size using the second motion increment to obtain a smooth inferred pose; Based on the first motion increment and the target relative coordinates, the shared virtual interactive object is rendered synchronously for the first MR device and the second MR device; When the second MR device scans the physical anchor point and generates the latest global calibration pose, the smoothed inference pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector. Construct a spatial elastic decay function centered on the shared virtual interactive object and with the spatial relative distance as the variable. Input the asynchronous calibration error vector into the spatial elastic decay function to generate an environmental deformation compensation field. While keeping the smoothed inference pose and the target relative coordinates unchanged in the relative coordinate system, the environmental deformation compensation field is applied to the static background environment rendering in the global coordinate system.
2. The method according to claim 1, characterized in that, When the second MR device scans a physical anchor point and generates a latest global calibration pose, the smoothed inference pose is inversely mapped to the global coordinate system to obtain the expected global pose. The difference between the latest global calibration pose and the expected global pose is calculated to obtain the asynchronous calibration error vector, specifically including: Record the target timestamp when the second MR device scans the physical anchor point; In the relative coordinate system, extract the historical smoothed projection pose corresponding to the target timestamp; The historical smoothed pose is inversely mapped to the global coordinate system to obtain the expected global pose; Calculate the pose transformation matrix between the latest global calibration pose and the expected global pose, decompose the pose transformation matrix into a translation error vector and a rotation error vector, and combine them to obtain the asynchronous calibration error vector.
3. The method according to claim 2, characterized in that, After the steps of mapping the smoothed inference pose to the global coordinate system inversely to obtain the expected global pose when the second MR device scans the physical anchor point and generates the latest global calibration pose, and calculating the difference between the latest global calibration pose and the expected global pose to obtain the asynchronous calibration error vector, the method further includes: Extract the magnitude of the translation error vector from the asynchronous calibration error vector; Determine whether the modulus length is greater than a preset error tolerance threshold; If the module length is greater than the preset error tolerance threshold, then the preset step size for iterative updating the second relative pose is increased according to the ratio of the module length to the preset error tolerance threshold.
4. The method according to claim 1, characterized in that, The construction of the spatial elastic decay function centered on the shared virtual interactive object and with the spatial relative distance as the variable specifically includes: Obtain the spatial bounding box of the shared virtual interactive object; With the geometric center of the bounding box as the origin, the area is divided outward into a rigid holding region, an elastic deformation region, and a background fixing region, with the boundary of the rigid holding region being larger than the boundary of the bounding box. Within the elastic deformation region, a decay weight function is constructed with the spatial relative distance as the independent variable, which monotonically decreases from the boundary of the rigid maintenance region to the boundary of the background fixed region, thus obtaining the spatial elastic decay function.
5. The method according to claim 4, characterized in that, The step of inputting the asynchronous calibration error vector into the spatial elastic attenuation function to generate the environmental deformation compensation field specifically includes: Obtain the global coordinates of each rendered vertex in the static background environment under the global coordinate system; Calculate the vertex distance from each of the rendered vertices to the geometric center of the bounding box; Input the vertex distance into the decay weight function to obtain the deformation weight corresponding to each of the rendered vertices; The asynchronous calibration error vector is multiplied by the deformation weight corresponding to each of the rendering vertices to obtain the local compensation vector of each of the rendering vertices. The local compensation vectors of each of the rendering vertices constitute the environmental deformation compensation field.
6. The method according to claim 5, characterized in that, The step of applying the environmental deformation compensation field to the static background environment rendering in the global coordinate system specifically includes: For rendering vertices located within the rigid preservation region, the global coordinates of the rendering vertices remain unchanged; For a rendered vertex located within the elastic deformation region, the global coordinates of the rendered vertex are superimposed with the corresponding local compensation vector to obtain the updated vertex coordinates; For the rendered vertices located within the fixed background area, the global coordinates of the rendered vertices are translated according to the asynchronous calibration error vector to obtain the updated vertex coordinates; Based on the updated vertex coordinates, the static background environment is reconstructed and rendered using a mesh.
7. The method according to claim 1, characterized in that, The method further includes: Real-time monitoring and updating of the spatial relative distance between the first MR device and the second MR device; When the spatial relative distance is greater than a preset interaction cancellation threshold, the iterative update in the relative coordinate system is stopped, and the latest global calibration pose of the first MR device and the second MR device is obtained. The preset interaction cancellation threshold is greater than the preset near-field interaction threshold. Based on the latest global calibration pose, an inverse transition matrix from the relative coordinate system to the global coordinate system is constructed; Using the inverse transition matrix, the shared virtual interactive object is smoothly mapped back to the global coordinate system, and the environmental deformation compensation field is canceled to restore the normal rendering of the static background environment in the global coordinate system.
8. A server, characterized in that, The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the server to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the server, the server causes the server to perform the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, When the computer program product is run on the server, the server performs the method as described in any one of claims 1-7.