Out-in body positioning method and system for in-xr games
Patent Information
- Application Number
- CN202610857303.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-12
AI Technical Summary
[0005]因此,如果XR游戏需要手腕/脚腕信息,就必须要求玩家额外穿戴手腕/脚腕传感器,增加了成本和穿戴负担
[0023]采用本发明的技术方案,在面对面的对抗XR游戏中,利用对方玩家的眼镜摄像头充当外向内传感器,配合每副眼镜自身内向外的6DoF定位,通过三角测量计算出各玩家的眼镜中心、左右手腕、左右脚腕等关键点的世界坐标,不需要在手腕、脚腕等位置额外穿戴任何传感器就可以实现肢体定位,从而减轻穿戴负担,提升游戏体验。
Smart Images

Figure CN122377133B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spatial positioning technology in anti-XR games, specifically relating to a method and system for limb positioning from inside to outside and from outside to inside in anti-XR games. Background Technology
[0002] The most basic XR smart glasses use outward-facing sensors mounted on the ceiling of the venue for positioning. The most common sensors are cameras or LiDAR. These sensors capture / acquire signal sources on the player's body and calculate the 3D position and orientation of these signals. The advantage of outward-facing is that the signal source can be the center of the glasses or any part of the body; typically, the signal source is placed on the wrist or ankle, allowing for games like handball or soccer. Outward-facing cameras can also directly identify the player's joints through machine learning, eliminating the need for sensors on every joint and instead using visual calculations to obtain positional information for the head / glasses, wrists, ankles, etc. Patent number ZL202511269525.X, entitled "Target Device Pose Calculation Method and Spatial Positioning System Based on Outward-Facing Technology" (hereinafter referred to as the "Prop Orientation Patent"), shows that cameras can also calculate the pose of the target device by identifying feature points on props (including glasses). However, the disadvantages of outward-facing are that deployment is expensive and non-relocatable, and the venue needs to precisely locate the position of each camera to calculate the prop's phase.
[0003] In the past five years, most large-space / small-space games have gradually adopted inside-out technology instead of outside-in technology. Inside-out technology uses the sensors of XR smart glasses, such as cameras or TOF LiDAR, to capture surrounding signal sources. These signal sources are typically graphics on walls and floors. By pre-recording the glasses' different positions and angles within the graphic environment, the glasses are trained to determine their own position and orientation. Patent number ZL202510180013.X, entitled "Positioning Method, Head-Mounted Display Device, and Storage Medium for Large-Space Smart Glasses" (hereinafter referred to as the "QR code positioning patent"), allows for rapid positioning of the glasses' location in the world by affixing QR code markers to known locations in the environment. The advantages of inside-out technology are its low cost and lack of equipment installation. Environmental mapping allows the glasses' cameras to remember their position, and the graphics themselves, like the QR code positioning patent, can incorporate position information, allowing the sensors to calculate their position more accurately. However, the downside of inside-out cameras is that they can only calculate the position of the center point of the glasses and cannot see the wrists or ankles. Inside-out cameras or radars cannot calculate the position of other body joints of the wearer. Gamers must wear wrist and ankle sensors to obtain information about their wrists and ankles to achieve limb positioning.
[0004] Existing XR positioning technology presents a dilemma: "you can't have your cake and eat it too." Outside-in technology (ceiling-mounted cameras to view players from the outside): can locate all joints (head, wrists, ankles, etc.), but is expensive to deploy and cannot be moved; Inside-out technology (the glasses' own camera sees the environment and locates itself) is cheap, flexible, and portable, but it can only locate the glasses themselves—it cannot see the wearer's own wrists and ankles (they are not in the field of view of the forward-facing camera).
[0005] Therefore, if XR games require wrist / ankle information, players must wear additional wrist / ankle sensors, increasing costs and the burden of wearing them. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for limb positioning in XR games that can achieve limb positioning without the need for additional wrist / ankle sensors.
[0007] This invention discloses a method for limb positioning in anti-XR games, involving outward and inward movement. It establishes a world coordinate system and sets up at least two physical anchor points with known world coordinates in the game area. The anti-XR game includes at least two players wearing XR glasses, and a server that interacts with all XR glasses. The limb positioning method includes the following steps: Step S1. The XR glasses output their pose in the world coordinate system using the 6DoF self-localization function; Step S2. The player looks at the other player through the glasses camera, captures the other player's camera image, performs 2D human key point detection on each frame of the image, identifies and outputs the pixel coordinates of each key point of the other player in the image; Step S3. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player to the server in real time. The server uses the observations of at least two glasses with known poses and the pixel coordinates of each key point of the player in the image to obtain the world position of each key point through cross-device collaborative triangulation. Step S4. The server distributes world location data of relevant key points to all players' glasses to achieve limb positioning.
[0008] If the game requires information about the direction of items, the game identifies the position of the opposing player holding the item based on the camera footage captured in step S2: (1) At least three positioning points ABC, which are not on the same straight line but on the same plane, are set on the handheld prop. The positioning points refer to signal sources in the outward-to-inward technology. The signal sources include electronic signal transmitters that can be photographed or sensed with built-in ID information, graphics or other identifiable signals. The signal sources are set on the handheld prop in a non-electrical connection manner or implemented in the circuitry of the handheld prop: Establish an XYZ coordinate system for the three positioning points ABC set on the handheld prop. Set the line connecting positioning points A and B as the X-axis, the perpendicular line passing through positioning point C and the X-axis as the Z-axis, set the intersection of the X and Z axes as the origin O, set the orthogonal direction above and below the origin O as the Y-axis, and set the direction of the directed line segment as the forward direction of the handheld prop. (2) Based on the camera footage of the opposite player collected in step S2, obtain the spatial position (X,Y,Z) of each positioning point. (3) The pose data of the handheld prop is obtained by spatial position triangulation of the three positioning points in step (2). The pose data includes yaw angle, pitch angle and roll angle.
[0009] The physical anchor points are placed in asymmetrical positions so that the relative angles of the anchor points are asymmetrical when viewed from any position, thus uniquely determining the orientation of the world coordinate system; the physical anchor points are markers in the corners of the arena, specific patterns on the ground, or other easily visually detectable feature points.
[0010] Step S2 identifies and outputs the pixel positions of each key point (POI) of the opposing player in the image, including pixel coordinates (u,v) and detection uncertainty. Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model: The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. , which represents the uncertainty output in this step; The first branch: when the detector outputs a heatmap At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (6); (7); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (8); in, For the preset reference pixel deviation, Here, is the preset lower bound for the minimum deviation, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (8). The two are equal, and the corresponding pixel covariance matrix is... ; Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K as specified by the camera's factory calibration: (9); in, , Focal length in pixels, , ) are the coordinates of the main point; First, convert the pixel coordinates of the target POI to normalized planar coordinates to eliminate the influence of camera intrinsic parameters: (10); When the Z component is normalized to 1, x_n represents the direction of the pixel of the target POI in the camera coordinate system. Its physical meaning is a vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v). Then, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (11); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (12); Where C_w is the world coordinate of the camera's optical center, and λ is the distance parameter traveled along the ray.
[0011] In step S3, cross-device collaborative triangulation refers to the process where two cameras in different poses observe the same keypoint POI Q, generating two rays. The closest pair of points P1 on these two rays is then determined. * and P2 * The midpoint is taken as the optimal estimate of POI Q, thus obtaining the world position of each keypoint POI. The specific steps are as follows: Based on the ray equation formula (12), the two rays are expressed as: Ray 1: (13); Ray 2: (14); Where C1 and C2 are the world coordinates of the optical centers of the two cameras, λ1 and λ2 are the distance parameters of the two rays along the rays, and r1 and r2 are the direction vectors of the two rays in the world coordinate system. The condition for a nearest point pair (P1*, P2*) is that the line connecting P1* and P2* is perpendicular to the directions of both rays, i.e.: ,and (15); Substituting formulas (13) and (14) into formula (15), the baseline vector is defined. Expanding, we get: (16); make When r1 and r2 are unit vectors, a = c = 1. This system of equations can be written in matrix form. Its coefficient determinant is Solving using Kramer's rule yields: (17); In face-to-face XR game scenarios, when two cameras observe the same point from different positions, the ray directions cannot be parallel, i.e., a·c - b²≠0. Substitute λ1* and λ2* back into formulas (13) and (14) respectively, and take the midpoint. As an estimate of the world coordinates of the key point POI Q: (18).
[0012] When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POI Q. The goal is to obtain the optimal estimate of keypoint POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
[0013] The weighted maximum likelihood estimation for any N rays is specifically as follows: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POIQ. For the i-th ray, with optical center C_i and direction r_i, the squared perpendicular distance from the intersection point Q of multiple rays in three-dimensional space to that ray is: (19); Where I is a 3×3 identity matrix. It is the projection matrix in the direction perpendicular to the ray. The projection matrix is defined. Then formula (19) is ; The goal of multi-ray triangulation is to find a three-dimensional point Q that minimizes the sum of the weighted distances of all rays. (20); Where w_i is the weight; Taking the derivative with respect to intersection point Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. The closed-form solution yields the core formula for cross-device collaborative triangulation: (twenty one); The three-dimensional covariance estimate of the intersection point Q is: (twenty two).
[0014] The theoretically optimal weights w_i are derived using the Fisher information matrix, enabling the least squares estimate WLS to reach the Cramér-Rao lower bound under Gaussian noise, which is the theoretical accuracy limit of the unbiased estimate. If the pixel observation model of the i-th ray is Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ; in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (twenty three); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; The optimal weights that enable least squares estimation (WLS) to reach the Cramér-Rao lower bound (CRLB) are obtained. for: (twenty four); Total Fisher Information Matrix The sum of the information content of each ray: (29); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao.
[0015] In actual deployment, formula (24) is decomposed into the multiplication of three independently estimable terms: (25); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. : (26); Where ε0 is a preset constant to prevent division by zero; The geometric mass weight w_geom is determined by the intersection angle. Control, reflecting the geometric contribution of the current ray pair to triangulation, intersection angle The closer to 90°, the greater the information content of the projection matrix A_i in the depth direction. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two-ray projection matrices A_1 + A_2 is proportional to sin²(α / 2). Therefore: (27); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (28); in, For reprojection residuals, The bandwidth parameter used to control the decay rate is equivalent to the Welsch function kernel in the M-estimator.
[0016] The cross-device collaborative triangulation uses the XR glasses of multiple game players as a distributed multi-camera array, extends the equivalent baseline to the distance between players, and allows multiple game players to exchange ray information to form the total Fisher information matrix. To extend this, a total Fisher information matrix is constructed through cross-device triangulation fusion. Including the player's own Fisher information matrix Fisher information matrix with the opposing player : (32); The covariance of cross-device triangulation is: (33).
[0017] The phrase "after obtaining cross-device triangulation results, continue to apply constraint fusion and temporal filtering to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy" refers to: At each time k, the system first extracts the two-dimensional coordinates of key human body points from the multi-view images. In conjunction with camera geometry, multi-ray triangulation is performed to obtain the initial triangulation results for each key point point (POI). and its covariance ,Should This involves taking several 3D rays from multiple camera viewpoints, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and then using weighted least squares to obtain the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame, along with the corresponding covariance. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the multi-source prior constraint optimization and fusion stage: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeleton connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on... are applied to the elbow, knee, shoulder, and hip joints. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the aforementioned multi-source prior constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, which is then used to obtain the optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable three-dimensional coordinates. and corresponding velocity estimation .
[0018] The aforementioned outlier removal using robust gating and RANSAC refers to introducing hard gating conditions to remove obviously unreliable rays, and using RANSAC purification when the number of rays is sufficient. The robust gating condition is set as follows: (a) The intersection angle α < the preset α_min; (b) The nearest point distance ε > the preset ε_max; (c) The reprojection residual r_proj > the preset τ_proj; Any ray that satisfies any of the above robust gating conditions is marked as an outlier and removed. When the number of available rays is greater than the preset value, first use RANSAC to select the set of interior points, and then perform least squares estimation WLS refinement using only the set of interior points.
[0019] The robust gating threshold is linked with the Fisher information matrix to form an adaptive gating: when the minimum eigenvalue of the total Fisher information matrix is less than the threshold, the robust gating threshold is relaxed to avoid over-removal; when Fisher information is sufficient, the robust gating threshold is tightened to improve accuracy.
[0020] To ensure the system can still provide reasonable output under degraded conditions, and automatically revert to high-precision mode after conditions recover, the system defines normal mode, degraded mode, and recovery mode, as well as the switching method between the three modes: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running a complete cross-device collaborative triangulation, constraint fusion, and Kalman filtering pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only a single camera is visible, intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model. When the second view becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
[0021] Furthermore, this invention discloses an outward-inward limb positioning system for combat XR games, which sets up at least two physical anchor points with known coordinates in the field based on an established world coordinate system; the combat XR game includes at least two players wearing XR glasses; and a server that interacts with all XR glasses. The XR glasses output their own pose through 6DoF self-localization and capture images of the opposing player. They perform 2D human keypoint detection on each frame of the image and output the pixel coordinates of each keypoint in the image. They upload their own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server in real time, and at the same time receive the world position data of the relevant keypoints sent by the server. The server includes a memory and a processor. The memory is connected to the processor and is used to store computer programs and the real-time uploads of each XR glasses of its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player. The processor is used to run computer programs, perform cross-device collaborative triangulation, obtain the world position of each key point, and distribute the relevant key point world position data to all players' XR glasses to achieve limb positioning.
[0022] Furthermore, the present invention discloses a storage medium storing a computer program, which, when executed by a processor, implements the aforementioned method for limb positioning from inside to outside and from outside to inside in an adversarial XR game.
[0023] By employing the technical solution of this invention, in face-to-face XR games, the camera of the opponent's glasses is used as an outward-to-inward sensor, combined with the 6DoF positioning of each pair of glasses from the inside out. The world coordinates of key points such as the center of each player's glasses, left and right wrists, and left and right ankles are calculated through triangulation. This eliminates the need to wear any additional sensors on the wrists, ankles, or other locations to achieve limb positioning, thereby reducing the burden of wearing them and improving the gaming experience. Attached Figure Description
[0024] Figure 1This is a top view of the game field in this invention; Figure 2 This is a flowchart of the limb positioning method in this invention; Figure 3 This is an analysis curve showing the impact of key parameters on positioning accuracy in this invention; Figure 4 This is a schematic diagram of the cross-device fusion collaborative triangulation principle in this invention; Figure 5 This is a flowchart of the constraint fusion and temporal filtering process in this invention; Figure 6 This is a schematic diagram of the three-state mode switching mechanism in this invention; Figure 7 This is a schematic diagram of data broadcasting between player P1 and player P2 in this invention; Figure 8 This is a system block diagram of the limb positioning system in this invention. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0027] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or solution described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Explanation of the principles of this application: 1. Regarding outward-to-inward positioning: The physical architecture for outside-in positioning involves installing sensor arrays at fixed locations within the game venue (usually above the ceiling or walls). Two common sensor types are: Infrared camera array (such as the OptiTrack system): The venue is equipped with 6 to 12 infrared cameras. Players wear infrared reflective marker balls. The cameras capture the infrared reflection of the marker balls. The three-dimensional world coordinates of the marker balls are calculated through multi-camera triangulation with an accuracy of sub-millimeter level. LiDAR (Light Detection and Ranging): It acquires three-dimensional point cloud data of objects in a scene by emitting laser pulses and measuring the reflection time, and extracts human body contours and joint positions from them.
[0028] Full-body joint coverage for location tracking: The sensors are external, with a field of view covering the entire playing area. Therefore, to locate a specific joint, simply place a signal source near that joint; the head (center of the glasses), wrists, ankles, and even fingers can be located. This makes games like handball (requiring collision detection at the wrist) and soccer (requiring kick detection at the ankle) possible.
[0029] Prop Phase Calculation: The prop orientation patent further expands the outward-inward capability. By arranging at least three non-collinear feature points (such as reflective marker balls or visually recognizable patterns) on props (such as rackets or glasses), the outward-inward camera can calculate the prop's complete six-degree-of-freedom posture, which is crucial for games such as fencing (sword tip direction) and table tennis (paddle tilt angle affects the direction of the shot).
[0030] Relying solely on outward-to-inward positioning presents the following technical problems: Deployment costs are extremely high: taking a 6-camera OptiTrack system as an example, the equipment purchase cost ranges from tens of thousands to hundreds of thousands of yuan. During installation, each camera must be fixed at a specific height and angle, and its pose in the world coordinate system must be accurately determined through a professional calibration process (usually using a calibration rod to wave around the site while all cameras are shooting simultaneously). The calibration accuracy requirement is sub-millimeter level. The entire installation and calibration process requires several hours of time from professional technicians.
[0031] Non-relocatable: Once calibration is complete, the camera positions cannot be changed. If a camera is knocked askew or a new game area is desired, all cameras must be recalibrated. This makes outside-in positioning suitable only for permanent, fixed venues, and not for temporary venues or game scenarios that require frequent relocation.
[0032] Therefore, an outward-to-inward positioning is suitable for large, fixed game venues, but not for lightweight, mobile, small competitive spaces.
[0033] 2. Regarding inward-to-outward positioning: Inside-out positioning reverses the sensor's position, placing the sensor on the XR glasses to look outwards at the environment. The most common implementation methods include: Visual Inertial Odometry (VIO) / SLAM: Grayscale or color cameras on glasses capture feature points of the surrounding environment (wall textures, floor patterns, corners of fixed objects, etc.), and combine them with acceleration and angular velocity data from the inertial measurement unit (IMU) to simultaneously build an environmental map and calculate the user's position and orientation on the map using the SLAM algorithm; QR code-assisted positioning: A QR code is pasted at a known location in the environment. The QR code encodes the user's world coordinate information. After the glasses camera in the QR code positioning patent recognizes the QR code, it not only obtains a reference coordinate point, but also directly calculates its own pose through the perspective deformation of the QR code, overcoming the cumulative drift problem of pure SLAM solutions. TOF / Structured Light Depth Sensor: The glasses can also be equipped with a time-of-flight (TOF) or structured light depth sensor, which can directly obtain three-dimensional depth information of the environment to assist SLAM positioning. Since it cannot detect hands outside the field of view, it is only suitable for table tennis games (such as the table tennis game on Quest3 and Apple Glasses), and not suitable for tennis or basketball games.
[0034] Low cost: No external equipment is required. Environmental preparation only requires pasting a few QR code papers in key locations, or even just relying on wall texture (pure SLAM solution). Equipment cost is almost zero.
[0035] Flexible and relocatable: There are no fixed sensors. Just tear off the QR code from site A and stick it to site B. The new site can be deployed in a few minutes, and the game site can be changed at any time.
[0036] No external sensor calibration required: All calculations are performed locally on the glasses, eliminating the need for external camera calibration.
[0037] Relying solely on inside-out positioning presents the following technical problems: The glasses' camera cannot see the wearer's own wrists and ankles: The camera on the glasses faces forward (in the same direction as the eyes' line of sight), and the field of view (FOV) is usually between 60° and 120°. When the wearer is standing naturally, their wrists are located to the side of the body, about 60cm below and to the side of the glasses. This position is completely outside the FOV of the forward-facing camera. Even if the arm is extended forward, the wrist may only enter the edge of the FOV for a limited time, and the angle is extreme, resulting in extremely poor detection accuracy. Furthermore, the wearer's ankles are located about 1.6m directly below the glasses, making it even less likely that they will be within the field of view of the forward-facing camera.
[0038] If the game requires wrist and ankle position information (handball requires wrist collision detection, and soccer requires ankle kick judgment), then players must wear separate sensor devices on their wrists and ankles (such as IMU wristbands / anklebands, or controllers / trackers with independent positioning capabilities such as the HTC VIVE Tracker), which increases the burden of wearing them (players need to wear one device on each wrist and ankle, for a total of 4 additional devices, and slime wearers can even have up to 12). Therefore, inside-out positioning is not suitable for games that require positional information such as wrists and ankles.
[0039] 3. Regarding the outward and inward limb positioning in XR games: To facilitate understanding of the basic principles of this invention, a 1vs1 XR game is used as an example. This explanation does not limit the application of this invention to multi-player scenarios. In a XR game, players face each other, and both players' smart glasses face each other. Player P1's glasses camera is directly facing the entire body of player P2, and player P2's glasses camera is directly facing player P1. This means that player P1 can use player P2's smart glasses camera and radar as an outward-facing sensor to observe key points of interest (POIs) such as their wrists and ankles, thereby obtaining the pixel positions of these POIs in the image. Since this outward-facing sensor is not fixed to the ceiling but moves with the player's head, it does not have an absolute world position. By utilizing the 6DoF self-localization technology of the smart glasses, the real-time pose of this outward-facing sensor can be determined through site anchor points and environmental features. This is equivalent to an outward-facing sensor with a known position, and the world positions of the key points of interest (POIs) can then be calculated using triangulation.
[0040] Therefore, the present invention provides anti-XR game in-game outward and outward-inward limb positioning, such as... Figure 2 As shown, it includes the following steps:
[0041] The outward-inward and inward-outward limb positioning method of this invention for adversarial XR games is applicable to adversarial XR games that require positional information such as wrists and ankles, and defines five key points of interest (POIs):
[0042] The player's POI1 (center point of contact) can be directly obtained from the player's 6DoF autolocalization output. The opposing player's POIs 2 through POIs 5 need to be calculated using the method of this invention, specifically including the following steps: Step S1. The player uses the 6DoF auto-localization function of the glasses to output the pose and orientation of the glasses in the world coordinate system, including the center point of the player's glasses (POI1): First, establish a world coordinate system W, such as... Figure 1 As shown, the origin O of the world coordinate system is located at the center of the game field, with the X-axis pointing horizontally to the right, the Y-axis pointing horizontally forward, and the Z-axis pointing vertically upward.
[0043] At least two physical anchor points O1 and O2 with known world coordinates are placed in the game area. These can be markers in the corners of the area, specific patterns on the ground, or other easily visually detectable feature points. Assuming anchor point O1 is located at (-1.8, 1.8, 0) and anchor point O2 is located at (1.8, -1.8, 0), respectively, at opposite corners of the area, with a Z-coordinate of 0 (ground), placing the anchor points in the corners of the area avoids occlusion and maximizes the geometric distribution of the anchor points to each player's smart glasses.
[0044] Player P1 is located near (-1, 0, 0), facing the +X direction; Player P2 is located near (1, 0, 0), facing the -X direction. The two players are approximately 2 meters apart and are in a face-to-face confrontation.
[0045] Define the global world coordinate system W using the set of anchor points: take the ground as the Z=0 reference plane, with the Z-axis pointing vertically upwards (normal vector n=[0,0,1)). The horizontal direction is determined by the anchor points. Specifically, the ground plane equation and coordinate axes are defined as follows: ,in, , d = 0 (1); , (2); Where z_w = n = [0,0,1] The direction is vertically upward. If the anchor point or the glasses IMU provides the direction of gravity, then z_w is replaced with its normalized opposite direction of gravity. Subscripts are uniformly represented by _* in the text.
[0046] The key design feature of this invention lies in the asymmetrical placement of the anchor points. In face-to-face competitive scenarios, the two players typically stand symmetrically on opposite sides of the arena. If the anchor points are also placed symmetrically (e.g., two anchor points are exactly on the central axis at both ends of the arena), the system may be unable to distinguish when the two players switch positions—resulting in mirror ambiguity. This invention places the anchor points in asymmetrical positions (e.g., O1 at one corner, O2 offset from the opposite corner), ensuring that the relative angles of the anchor points are asymmetrical from any standing position, thus uniquely determining the orientation of the world coordinate system. This design also enhances the observability of the glasses' yaw angle; the asymmetrical difference in the relative orientation of the anchor points provides additional directional constraints. The observability constraints are: (3); Where t_i is the world coordinate position of glasses i. This is the minimum observable angle threshold.
[0047] After establishing a world coordinate system, each pair of glasses uses its built-in 6DoF self-localization function to determine its 6DoF pose and orientation in the coordinate system in real time. For example, a fusion framework of "anchor point PnP absolute correction + IMU / VIO prediction" is adopted, which combines the accuracy of absolute positioning with the continuity of inertial navigation.
[0048] Anchor Point PnP Section: Establishing the 3D world coordinates of the anchor point The camera extrinsic parameters are solved by minimizing the robust reprojection error to establish the correspondence between the 2D pixel coordinates u_k detected in the camera image and the extrinsic parameters u_k. (4); Where K is the camera intrinsic parameter matrix, π(·) is the perspective projection function, ρ(·) is the robust loss function (Huber function) to suppress the influence of false detection, and R_cw is the rotation matrix from the world frame to the camera frame, which is directly obtained by solving PnP of formula (4); Solving for the transformation from the camera frame to the world frame yields the solution. : , (5); Where R_wc is the rotation matrix from the camera frame to the world frame, obtained by transposing R_wc, and C_w is the world coordinate of the camera's optical center. The translation vector to be solved for PnP; The IMU / VIO fusion component provides inertial predictions when anchor points are not visible, preventing pose loss. Extended Kalman filtering (EKF) or sliding window factor maps are used to fuse IMU pre-integration with PnP observations: when anchor points are visible, PnP provides low-frequency absolute correction; when anchor points are temporarily occluded, IMU / VIO maintains short-term pose continuity. In engineering implementation, prediction weights are increased when anchor points are invisible, and relinearization is performed when anchor points are restored to suppress cumulative drift.
[0049] In addition, an online self-calibration mechanism needs to be added to jointly estimate parameters such as binocular extrinsic drift T_LR, time offset δt, and IMU bias (b_a, b_ω) in the low-frequency sliding window, and use anchor point reprojection consistency and bone length stability as constraints to achieve long-term stability and cross-device consistency.
[0050] Step S2. The player looks at the opposing player through the glasses' camera, captures the opposing player's camera view, and performs 2D human keypoint detection on each frame of the image, identifying and outputting the pixel coordinates of each keypoint of the opposing player in the image: Step S1 outputs the pose and orientation of the glasses in the world coordinate system, determining "where the player is looking from". Next, it is necessary to determine "what the player is seeing", that is, the player looks at the opposite player through the glasses' camera, captures the opposite player's camera image, and performs 2D human key point detection on each frame of the image, outputting the pixel coordinates (u,v) of each POI of the opposite player in the image and the detection uncertainty. The human key points include the left and right wrist joints and the left and right ankle joints.
[0051] Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model (used to provide source data for subsequent weight construction and end-to-end uncertainty propagation): The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. This is the uncertainty output in this step, where the first branch: when the detector outputs a heatmap At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (6); (7); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (8); in, The baseline pixel deviation is approximately 3.0 pixels. The minimum deviation lower bound is 0.5 pixels, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation (bias) in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (8). The two are equal, and the corresponding pixel covariance matrix is... In this way, high-confidence detection yields small covariance (high weight), while low-confidence detection yields large covariance (low weight), enabling uncertainty-driven adaptive weight adjustment in subsequent triangulation measurements.
[0052] Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K (given by the camera's factory calibration): (9); in, , The preset focal length in pixels, , The coordinates of the principal point are ideally located at the center of the image. The camera intrinsic parameter matrix K describes the mapping relationship between the three-dimensional spatial orientation and the image pixel coordinates, serving as a bridge connecting the two worlds. For lenses with distortion, distortion correction must be performed before backprojection.
[0053] First, the pixel coordinates of the target POI are converted to normalized planar coordinates to eliminate the influence of intrinsic parameters: (10); Where x_n represents the direction of the pixel of the target POI in the camera coordinate system (Z component normalized to 1), and its physical meaning is the vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v). Next, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (11); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (12); Where C_w is the world coordinate of the camera's optical center, λ is the distance parameter along the ray, and the target POI is located at a certain λ value on this ray. The specific λ value cannot be determined by a single ray alone; multiple rays (at least two) are needed to determine it.
[0054] If the game requires information about the orientation of props, for example, when a player is holding a prop (racket, fencing blade, etc.) during a competition, in addition to knowing the position of the wrist, it is also necessary to know the rotational posture of the prop, such as whether the racket is forehand or backhand, which will affect the direction and effect of the shot. Therefore, based on the camera footage of the opposing player captured in step S2, the posture of the opposing player holding the prop can be identified (this can be achieved using existing technology or by utilizing the technical solution of a prop orientation patent): (1) At least three positioning points ABC, which are not on the same straight line but on the same plane, are set on the handheld prop. The positioning points refer to signal sources in the outward-to-inward technology. The signal sources include electronic signal transmitters that can be photographed or sensed with built-in ID information, graphics or other identifiable signals. The signal sources are set on the handheld prop in a non-electrical connection manner or implemented in the circuitry of the handheld prop: Establish an XYZ coordinate system for the three positioning points ABC set on the handheld prop. Set the line connecting positioning points A and B as the X-axis, the perpendicular line passing through positioning point C and the X-axis as the Z-axis, set the intersection of the X and Z axes as the origin O, set the orthogonal direction above and below the origin O as the Y-axis, and set the direction of the directed line segment as the forward direction of the handheld prop. (2) Based on the camera footage of the opposite player collected in step S2, obtain the spatial position (X,Y,Z) of each positioning point. (3) The pose data of the handheld prop is obtained by spatial position triangulation of the three positioning points in step (2). The pose data includes yaw angle, pitch angle and roll angle.
[0055] Step S3. Each pair of glasses uploads its pose in the world coordinate system and the pixel coordinates of each key point to the server in real time. The server uses the observations of at least two glasses with known poses, as well as the pixel coordinates of each player's key point in the image, to perform cross-device collaborative triangulation to obtain the world position of each key point: The cross-device collaborative triangulation in the server can obtain the world position of key point POI by finding the intersection of two rays, or it can obtain the world position of each key point POI by a multi-ray robust localization method based on Fisher information fusion given in this invention.
[0056] When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POI Q. The goal is to obtain the optimal estimate of keypoint POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
[0057] 1. For ease of intuitive understanding, the derivation is based on the simplest two-ray case (special case N=2) (hereinafter referred to as the "two-ray method"). Assume two cameras in different poses observe the same keypoint POI Q, generating two rays. Ideally, these two rays intersect at point Q. However, due to observation noise (pixel detection error, pose estimation error, etc.), these two rays will be two skew lines and typically will not intersect precisely in space. Therefore, this invention finds the closest pair of points P1 on the two rays. * and P2 * The midpoint is taken as the optimal estimate of POI Q. The specific steps are as follows: Based on the ray equation (Formula (12)), the two rays can be expressed as: Ray 1: (13); Ray 2: (14); Where C1 and C2 are the world coordinates of the optical centers of the two cameras, λ1 and λ2 are the distance parameters along the rays, and r1 and r2 are the direction vectors of the two rays in the world coordinate system. The condition for a nearest point pair (P1*, P2*) is that the line connecting P1* and P2* is perpendicular to the directions of both rays, i.e.: ,and (15); Substituting formulas (13) and (14) into formula (15), and defining the baseline vector Δ = C2 - C1, we can expand to obtain: (16); make When r1 and r2 are unit vectors, a=c=1. The two perpendicular conditions are simplified to a system of two linear equations, which can then be written in matrix form. Its coefficient determinant is Solving using Kramer's rule yields: (17); The solution condition is a·c - b²≠0. By the Cauchy-Schwarz inequality, a·c - b²≥0, and the equality holds if and only if r1∥r2 (the two rays are parallel). In a face-to-face XR game scenario, two cameras observe the same point from different positions, and the ray directions cannot be parallel, so this condition is naturally satisfied.
[0058]
Derivation and verification of formula (17)
[0059] Substitute λ1* and λ2* back into formulas (13) and (14) respectively, and take the midpoint. As an estimate of the world coordinates of the key point POI Q: (18); The distance between the two closest points As a consistency indicator of the positioning quality of this frame, The smaller the value, the closer the two rays are to intersecting, and the more reliable the positioning.
[0060] 2. While the two-ray method is intuitive, it cannot utilize information from more than two rays, nor can it introduce uncertainty weights. This invention extends cross-device collaborative triangulation to a weighted least squares (WLS) estimation framework for N rays, equivalent to maximum likelihood estimation (MLE) under the Gaussian noise assumption, thus enabling multi-ray triangulation: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POIQ. For the i-th ray (optical center C_i, direction r_i), the squared perpendicular distance from the intersection point Q of multiple rays in three-dimensional space to that ray is: (19); Where I is a 3×3 identity matrix. It is the projection matrix (symmetric and idempotent) in the direction perpendicular to the ray. Define the projection matrix. Then formula (19) is .
[0061] The goal of multi-ray triangulation is to find a three-dimensional point Q that minimizes the sum of the weighted distances of all rays. (20); Where w_i is the weight; This is a quadratic optimization problem about point Q. Taking the derivative with respect to point Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. Closed-form solution: (twenty one); The three-dimensional covariance estimate of the intersection point Q is: (twenty two); Formula (21) is the core formula for cross-device collaborative triangulation. It only requires calculating a 3×3 matrix A_i for each ray and accumulating the results, and finally performing a 3×3 matrix inversion to obtain the result. The computational complexity is O(N), which fully meets the requirements of real-time operation on the end side. When N=2 and the weights are equal, formula (21) degenerates into a result equivalent to the midpoint method of the two rays.
[0062] In classical methods, the design of weights w_i relies on engineering experience. This invention, starting from statistical estimation theory, uses the Fisher Information Matrix (FIM) to derive the theoretically optimal weights w_i, enabling the least squares estimation (WLS) to reach the Cramér-Rao lower bound (CRLB) under Gaussian noise, i.e., the theoretical accuracy limit of the unbiased estimate. If the pixel observation model of the i-th ray is Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ,in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (twenty three); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; As can be seen from formula (23), the information content of each ray is determined by three factors: focal length f (the longer the focal length, the more information), distance d_i (the closer the distance, the more information), and detection accuracy. (The more precise the information, the more information there is), and the spatial orientation structure of the information is entirely determined by the projection matrix A_i (which only provides information in the direction perpendicular to the ray).
[0063] Therefore, the optimal weights that allow the least squares estimate (WLS) to reach the Cramér-Rao lower bound (CRLB) are... for: (twenty four); This optimal weighting naturally unifies distance attenuation and detection uncertainty into an information-theoretic framework, eliminating the need for manual selection of attenuation functions or adjustment of hyperparameters. In practical engineering implementations, d_i can be approximated by the estimation result of the previous frame, forming an iterative optimization.
[0064] In actual deployment, formula (24) is further decomposed into three independently estimable terms multiplied together to adapt to the characteristics of each stage in engineering practice: (25); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. (26); Wherein, ε0 is a preset constant to prevent division by zero (generally set to 0.01 pixels²). The geometric mass weight w_geom is determined by the intersection angle. Control, reflecting the geometric contribution of the current ray pair to triangulation, intersection angle The closer to 90°, the greater the information content of the projection matrix A_i in the depth direction. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two-ray projection matrices A_1 + A_2 is proportional to sin²(α / 2), therefore: (27); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair.
[0065] The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (28); in, For reprojection residuals, To control the bandwidth parameter of decay rate, this reprojection consistency weight is equivalent to the Welsch function kernel in the M-estimator, giving the WLS framework robust statistical significance.
[0066] Total Fisher Information Matrix The sum of the information content of each ray: (29); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao; From this, the following accuracy rules can be derived: A. Ray direction error With pixel detection error The relationship is: (30); B. Three-dimensional positioning error The upper limit of the engineering quantity level is: (31); Where B is the equivalent baseline length, D is the target distance, and α is the equivalent intersection angle. For example... Figure 3 As shown, the formula clearly reveals the influence of each parameter on positioning accuracy: the larger the baseline B, the better; the closer the distance D, the better; the closer the intersection angle α is to 90°, the better; and the more accurate the pixel detection, the better.
[0067] Accuracy Analysis and Typical Parameter Estimation: For a single-device binocular baseline B=0.064m and target distance D=1.83m, the actual intersection angle α is only about 2°, and the upper bound of the error given by formula (31) is relatively large. However, this is the worst-case estimate, and the actual error is significantly smaller than the upper bound due to the randomness of the noise direction. More importantly, formula (31) reveals the inherent accuracy bottleneck of single-device narrow baseline triangulation - the error along the depth direction (line of sight) is much larger than the lateral error, and this anisotropy is manifested by the eigenvalue structure of CRLB.
[0068] Typical lateral accuracy estimation: With D=1.83m, f=500 pixels, For example, =2 pixels: ; Since the lateral accuracy is on the order of millimeters, it is sufficient for determining whether a punch has hit the opponent's body (lateral judgment) in combat XR games. The accuracy in the depth direction needs to be improved through cross-device collaborative triangulation fusion and constraint fusion proposed in this invention.
[0069] The baseline of a single-device binocular glasses is only 6.4cm. When the target distance exceeds 1m, the triangulation accuracy in the depth direction is limited by a very small intersection angle (α<3°). The cross-device collaborative triangulation fusion mechanism proposed in this invention uses the XR glasses of two (or even multiple) players as a distributed multi-camera array, extending the equivalent baseline from 6.4cm to the player spacing (1.5m~3m), thereby improving depth accuracy by one to two orders of magnitude.
[0070] like Figure 4 As shown, taking a face-to-face battle between players P1 and P2 as an example, player P1's glasses observe key points (POIs) on player P2, generating a ray set {r_i^(P1)}. Simultaneously, player P2's glasses can also observe symmetry information on player P1 (or on nearby reference points). The two players exchange ray information, which affects the overall Fisher information matrix. To expand this, a cross-device integrated Fisher information matrix is constructed, including the player's own Fisher information matrix. Fisher information matrix with the opposing player : (32) Because the angle difference is close to 180° in face-to-face combat, the two sets of rays from both players come from almost orthogonal directions, and the player's own Fisher information matrix... The lack of information in the depth direction is precisely due to the Fisher information matrix set up by the opposing player. Complete it, and vice versa.
[0071] The covariance of cross-device fusion is: (33); Based on typical parameters: player spacing B_cross = 3m, D = 1.5m, f = 500 pixels, Taking 2 pixels as an example, the equivalent intersection angle is close to 90°: ; Therefore, cross-device collaborative triangulation fusion improves positioning accuracy from the centimeter level to the millimeter level, fully meeting the requirements of fine collision detection.
[0072] The robust gating and RANSAC achieve outlier ray removal: a hard gating condition is introduced to remove obviously unreliable rays, and RANSAC is used for further purification when the number of rays is sufficient; The gating condition is set such that any ray is marked as an outlier and discarded when it meets any of the following conditions: (a) The intersection angle α < the preset α_min (it is recommended that α_min = 5°~10°); (b) The nearest point distance ε > the preset ε_max (it is recommended that ε_max = 0.1m); (c) The reprojection residual r_proj > the preset τ_proj (τ_proj = 5~10 pixels is recommended).
[0073] The RANSAC purification mechanism works as follows: when the number of available rays is ≥3, firstly, an interior point set is selected using RANSAC (two rays are randomly selected for solving, and the remaining rays are voted on), and then least squares estimation (WLS) refinement is performed using only the interior point set. This mechanism is particularly effective in cases of extreme detection errors or abrupt pose changes.
[0074] This invention links the gating threshold with the Fisher information matrix to form adaptive gating: when the total Fisher information matrix... When the minimum eigenvalue is small (poor geometric configuration), the gating threshold should be appropriately relaxed to avoid over-rejection; when Fisher information is sufficient, the gating threshold should be tightened to improve accuracy.
[0075] Triangulation calculates the 3D coordinates of each Point of Interest (POI) independently frame by frame, without considering the temporal continuity between frames and the physical constraints between joints. This invention compensates for this deficiency by fusing multi-source prior constraints and temporal filtering, such as... Figure 5 As shown, a joint optimization along the spatial and temporal dimensions is achieved, specifically as follows: At each time k, the system first extracts the two-dimensional coordinates of human key points from the multi-view images. (and (It means the same thing), and combined with the camera geometry, multi-ray triangulation is performed to obtain the triangulation results of the initial estimate of each key point (POI). and its covariance ,Should It takes several 3D rays from multiple camera views, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and obtains the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame using weighted least squares (WLS). The corresponding covariance is then calculated. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the constraint optimization fusion phase: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeleton connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on the following are applied to joints such as elbows, knees, shoulders, and hips. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the above constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, and solves it to obtain the constrained optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint (POI) to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable 3D coordinates. and corresponding velocity estimation ; Based on this, the system can further adjust according to uncertainty. The collision detection threshold is adaptively adjusted, and the position of the opponent's root node and the body orientation are inferred from points of interest (POIs) such as the left and right hip joints and left and right shoulder joints.
[0076] The ankle ground constraint refers to the fact that when the player is detected to be standing with both feet on the ground, the ankle joint POI is pressed against the ground plane at a height of approximately Z = h0 (h0 ≈ 0.08m, i.e., the anatomical height of the ankle joint POI from the ground). The horizontal coordinate estimate of the ankle joint POI can be obtained by finding the intersection of a single ray with the horizontal plane Z = h0. Depend on Solving for (34); (35); The energy term constrained by the ground is: (36); Where w_j represents the weight of the ground constraint, which is the inverse of the observation covariance, i.e.: , The uncertainty representing the ankle height observation is usually set to a fixed prior value. A smaller value indicates a more reliable ground constraint, corresponding to a higher weight. The larger the size, the higher its proportion in the integration.
[0077] When the triangulation results When present, the variance of triangulation is weighted and fused inversely proportional to the variance of ground constraints: (37); in, The variance of the triangulation (from the inverse of the Fisher information matrix) reflects the uncertainty of the triangulation. The ground constraint variance (fixed prior value, such as σ²_ground≈(0.02m)²) reflects the measurement accuracy of ankle height h0; Formula (37) is weighted and fused by the inverse variance of the two. When the depth variance of the triangulation is large (due to anisotropy caused by narrow baseline), the ground constraint has a higher weight in the height direction; and vice versa. It should be noted that this constraint is only effective when the player is standing with both feet on the ground. When jumping, the ankles leave the ground, and the system automatically detects this through inertial sensors or posture classification and reverts to pure triangulation.
[0078] Triangulation calculations are performed independently for each joint, without considering the physical connections between joints, which may lead to inconsistent bone lengths. This invention introduces a constant bone length constraint, meaning the distance between adjacent joints should be approximately constant (human bone length does not change between frames): (38); Where B is the skeleton connection set, For the calibrated or statistically obtained bone lengths (e.g., shoulder → wrist chain length ≈ 0.55m, hip → ankle chain length ≈ 0.82m), Q_k is the constrained optimized 3D keypoint result, and Σ_{Q_k} is the optimized covariance of the corresponding 3D keypoint Q_k; Note: The covariance of all 3D keypoints after the above constraint optimization is written as Σ_{Q_*}; then it is written as Σ_{Q_k} at a single keypoint, and is abbreviated as Σ_Q when the keypoint subscript is not emphasized.
[0079] If the triangulation results If the constant bone length constraint is violated, a proportional scaling correction is applied along the joint chain direction to obtain the correction value: (39); The relationship between the correction value and the triangulation result is as follows: taking the midpoint of the two joints, center=(Q_j+Q_k) / 2, as the reference, the unit vector d along the bone direction... =(Q_k-Q_j) / ||Q_k-Q_j||, which scales the bone length proportionally to the standard bone length. ,get , This means normalizing the bone length to a standard value while keeping the midpoint of the joint chain unchanged.
[0080] Traditional methods impose joint angle constraints in Euclidean space, but joint rotation is naturally defined in rotation groups. On the manifold, this invention uses geodesic distance on the Riemannian manifold instead of Euclidean distance as the constraint cost function to avoid gimbal lock and singularity problems.
[0081] The basis Joint angle constraints for geodesic distance of rotating manifolds: (40); in, The joint rotation matrix is calculated from the triangulation results. Let the physiological limit boundary of joint j be defined as a submanifold on the rotational manifold SO(3). For the geodesic distance on the rotating manifold SO(3), when the joint rotation is within the physiologically accessible range Inside, = 0, no penalty; when it exceeds the physiologically achievable range, project back to the nearest legal rotation along the geodesic.
[0082] For the elbow joint (flexion-extension range 0°–145°) and the knee joint (flexion-extension range 0°–135°), it can be degenerated into a one-dimensional case: In this case, it is consistent with classical Euclidean constraints; however, for multi-degree-of-freedom ball-and-socket joints such as the shoulder and hip joints, the Riemannian formulation avoids Euler angle singularities and provides geometrically more accurate constraint propagation.
[0083] The above constraint optimization is fused into: (41).
[0084] The aforementioned constraint optimization fusion optimizes the results for each frame in the spatial dimension, but jitter may still exist between frames. This invention maintains a Kalman filter for each keypoint (POI), smooths the output in the temporal dimension, and estimates joint velocity. The state vector output by the Kalman filter contains six components: position and velocity. .
[0085] Observation equations (assuming uniform motion): (42); In the observation equation, the observations are the coordinates of the three-dimensional key points after constraint optimization, and the observation noise covariance is its corresponding three-dimensional covariance Σ_{Q,k} after constraint optimization (which can be simplified to Σ_Q). (43); in, To constrain the optimized 3D covariance, rather than the original triangulation covariance Σ_tri, as the observation noise for the Kalman filter, this means that when the triangulation accuracy is poor in a certain frame ( (Large and insufficient information) The filter automatically relies more on prediction and reduces the update amplitude to achieve measurement noise adaptation. The Kalman filter output velocity estimate v can be used for collision force calculation (fist speed affects damage value) and animation transition without additional calculation.
[0086] According to uncertainty Adaptive adjustment of the collision determination threshold refers to mapping the three-dimensional uncertainty (quantized by CRLB) to an adaptive threshold for collision determination: (44); in, The preset basic collision distance threshold (e.g., 0.05m). This is the scaling factor.
[0087] This means that the more uncertain the positioning ( The larger the threshold, the more conservative the collision detection (the larger the threshold), which can reduce false alarms; the more accurate the positioning, the more sensitive the detection.
[0088] This mechanism can also be used in reverse for hit confidence scoring, that is, upgrading collision determination from a hard threshold to a probability assessment: (45); in, For the standard normal CDF, This represents the POI spacing.
[0089] The inference of hand root node position and body orientation from points of interest (POIs) such as left and right hip joints and left and right shoulder joints refers to: The overall position and orientation of the opponent can be inferred from the set of three-dimensional coordinates of multiple POIs, which is very important for game rendering and tactical judgment.
[0090] Take the midpoint between the left and right hip joints (or left and right shoulder joints) as the root node position (overall position) of the opponent: (46); The yaw angle of body orientation is determined by the intersection of the line connecting the two shoulders with the world horizontal plane. In practical applications, multiple sets of orientation evidence (shoulder, hip, head and chest orientation) can be integrated and weighted by Fisher information weights of each POI to improve the stability of orientation estimation.
[0091] In real-world competitive scenarios, various degradation conditions are unavoidable. This invention proposes... Figure 6 The three-state mode switching mechanism shown here, which switches between normal mode, degraded mode and recovery mode, ensures that the system can still provide reasonable output under degradation conditions, and automatically returns to high-precision mode after the conditions recover. When the intersection angle And the distance between the two closest points Furthermore, when the reprojection residual is qualified, the system operates in normal mode, running the complete triangulation + constraint fusion + Kalman filter pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only one camera is visible (the other party is turned to the side, so only one type of eye can see the joint), intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model (such as SMPL). When the second viewpoint becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
[0092] Furthermore, this invention provides a mechanism to ensure cross-device consistency verification: when communication is available, players P1 and P2 exchange their estimated low-frequency states (position, orientation, confidence) and construct consistency constraints ‖t1-t1'‖ and ‖φ1-φ1'‖. When there is inconsistency, an alarm is triggered: the collision determination weight is reduced, the gating threshold is increased, or a realignment is requested, thereby achieving collaborative verification and anomaly detection.
[0093] Step S4. The server distributes the world location data of relevant keypoints (POIs) to all player glasses. The world location data of the relevant keypoints (POIs) can be used for game rendering or collision detection for player glasses. Each pair of glasses can observe and calculate the key points of interest (POIs) of the opposing player, but game rendering or collision detection requires knowing the POIs of all players (including your own and the opposing player's), such as... Figure 7 The diagram shows the information gathering and distribution for players P1 and P2.
[0094] When this invention is extended from 1v1 to multiplayer (such as 2v2, 3v3), each pair of glasses independently performs joint recognition and triangulation for each opposing player within its FOV; the same joint may be observed by multiple pairs of glasses at the same time, and the more observations, the more overdetermined equations are formed by triangulation, and the higher the accuracy; the server distributes the POI data of each player according to their player ID.
[0095] Example 1 Embodiment 1 of the present invention discloses a method for outward and inward limb positioning in anti-XR games, such as... Figure 1 and 2 As shown, a world coordinate system is established, and at least two physical anchor points with known world coordinates are set up in the game field; the XR game includes at least two players wearing XR glasses; and a server that can interact with all XR glasses. The physical anchor point can be a marker in the corner of the arena, a specific pattern on the ground, or other easily visually detectable feature points; the anchor point is placed in an asymmetrical position so that the relative angle of the anchor point is asymmetrical from any standing position, thereby uniquely determining the orientation of the world coordinate system; the limb positioning method includes the following steps: Step S1. The XR glasses output their pose in the world coordinate system using the 6DoF self-localization function: Step S2. The player looks at the opposing player through the glasses' camera, captures the opposing player's camera view, and performs 2D human keypoint detection on each frame of the image, identifying and outputting the pixel positions of each keypoint (POI) of the opposing player in the image, including pixel coordinates (u,v) and detection uncertainty: Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model: The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. , which represents the uncertainty output in this step; The first branch: when the detector outputs a heatmap At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (6); (7); in, Iterate through all pixel locations in the heatmap. The corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (8); in, The baseline pixel deviation is approximately 3.0 pixels. The minimum deviation lower bound is 0.5 pixels, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (8). The two are equal, and the corresponding pixel covariance matrix is... ; Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K as specified by the camera's factory calibration: (9); in, , Focal length in pixels, , The main point coordinates are ideally located at the center of the image. The camera intrinsic parameter matrix K describes the mapping relationship between the three-dimensional spatial direction and the image pixel coordinates. For lenses with distortion, distortion correction is required before backprojection. First, the pixel coordinates of the target POI are converted to normalized planar coordinates to eliminate the influence of camera intrinsic parameters: (10); When the Z component is normalized to 1, x_n represents the direction of the pixel of the target POI in the camera coordinate system. Its physical meaning is a vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v). Next, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (11); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (12); Where C_w is the world coordinate of the camera's optical center, and λ is the distance parameter traveled along the ray; If the game requires information about the orientation of the props, based on the camera footage of the opposing player captured in step S2, the position of the opposing player holding the props can be identified (this can be achieved using existing technology or by utilizing the technical solution of the prop orientation patent): (1) At least three positioning points ABC, which are not on the same straight line but on the same plane, are set on the handheld prop. The positioning points refer to signal sources in the outward-to-inward technology. The signal sources include electronic signal transmitters that can be photographed or sensed with built-in ID information, graphics or other identifiable signals. The signal sources are set on the handheld prop in a non-electrical connection manner or implemented in the circuitry of the handheld prop: Establish an XYZ coordinate system for the three positioning points ABC set on the handheld prop. Set the line connecting positioning points A and B as the X-axis, the perpendicular line passing through positioning point C and the X-axis as the Z-axis, set the intersection of the X and Z axes as the origin O, set the orthogonal direction above and below the origin O as the Y-axis, and set the direction of the directed line segment as the forward direction of the handheld prop. (2) Based on the camera footage of the opposite player collected in step S2, obtain the spatial position (X,Y,Z) of each positioning point. (3) The pose data of the handheld prop is obtained by spatial position triangulation of the three positioning points in step (2). The pose data includes yaw angle, pitch angle and roll angle.
[0096] Step S3. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player to the server in real time. The server uses the observations of at least two glasses with known poses and the pixel coordinates of each key point of the player in the image to obtain the world position of each key point through cross-device collaborative triangulation. The cross-device collaborative triangulation refers to the process where two cameras in different poses observe the same keypoint POI Q, generating two rays. The world position of the keypoint POI is obtained by finding the intersection of these two rays. This invention provides a two-ray midpoint method, which finds the closest pair of points P1 on the two rays. * and P2 * The midpoint is taken as the optimal estimate of POI Q. The specific steps are as follows: Based on the ray equation formula (12), the two rays are expressed as: Ray 1: (13); Ray 2: (14); Where C1 and C2 are the world coordinates of the optical centers of the two cameras, λ1 and λ2 are the distance parameters of the two rays along the rays, and r1 and r2 are the direction vectors of the two rays in the world coordinate system. The condition for a nearest point pair (P1*, P2*) is that the line connecting P1* and P2* is perpendicular to the directions of both rays, i.e.: ,and (15); Substituting formulas (13) and (14) into formula (15), and defining the baseline vector Δ = C2 - C1, we can expand to obtain: (16); make When r1 and r2 are unit vectors, a = c = 1. This system of equations can be written in matrix form. Its coefficient determinant is Solving using Kramer's rule yields: (17); In face-to-face XR game scenarios, when two cameras observe the same point from different positions, the ray directions cannot be parallel, i.e., a·c - b²≠0. Substitute λ1* and λ2* back into formulas (13) and (14) respectively, and take the midpoint. As an estimate of the world coordinates of the key point POI Q: (18); The distance between the two closest points As a consistency indicator of the positioning quality of this frame, The smaller the value, the closer the two rays are to intersecting, and the more reliable the localization. When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POI Q. The goal is to obtain the optimal estimate of keypoint POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
[0097] The weighted maximum likelihood estimation for any N rays is as follows: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POIQ. For the i-th ray (optical center C_i, direction r_i), the squared perpendicular distance from the intersection point POIQ of multiple rays in 3D space to ray i is: (19); Among them, the definition Let be the projection matrix in the direction perpendicular to the ray. I is a 3×3 identity matrix; The goal of multi-ray triangulation is to find the point of intersection (POI) in three-dimensional space that minimizes the sum of the weighted distances of all rays. (20); Where w_i is the weight; Taking the derivative of the keypoint POI Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. The closed-form solution yields the core formula for cross-device collaborative triangulation: (twenty one); The three-dimensional covariance estimate of the intersection point Q is: (twenty two); When N=2 and the weights are equal, formula (21) degenerates into a result equivalent to the midpoint method of the two rays.
[0098] The theoretically optimal weight w_i is derived using the Fisher Information Matrix (FIM), enabling the least squares estimate WLS to reach the Cramere-Rao lower bound (CRLB) under Gaussian noise, i.e., the theoretical accuracy limit of the unbiased estimate: if the pixel observation model of the i-th ray is... Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ,in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point POI Q is: (twenty three); Where f is the focal length. Let be the distance from the intersection point POI Q to the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; The optimal weights that allow the least squares estimate (WLS) to reach the Cramér-Rao lower bound (CRLB) for: (twenty four); In practical engineering implementation, d_i can be approximated by the estimation result of the previous frame, forming an iterative optimization; In actual deployment, formula (24) is further decomposed into the multiplication of three independently estimable terms: (25); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. : (26); Wherein, ε0 is a preset constant to prevent division by zero (generally set to 0.01 pixels²). The geometric mass weight w_geom is determined by the intersection angle. Control, reflecting the geometric contribution of the current ray pair to triangulation, intersection angle The closer to 90°, the greater the information content of the projection matrix A_i in the depth direction. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two-ray projection matrices A_1 + A_2 is proportional to sin²(α / 2). Therefore: (27); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (28); in, For reprojection residuals, The bandwidth parameter used to control the decay rate is equivalent to the Welsch function kernel in the M-estimator. The sum of the information content of each ray is the total Fisher information matrix. : (29); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao; From this, the following accuracy rules can be derived: A. Ray direction error With pixel detection error The relationship is: (30); B. Three-dimensional positioning error The upper limit of the engineering quantity level is: (31); Where B is the equivalent baseline length, D is the target distance, and α is the equivalent intersection angle; Lateral accuracy estimation: ; By using the XR glasses of both (or even multiple) players as a distributed multi-camera array, the equivalent baseline is extended to the distance between players, improving depth accuracy by one to two orders of magnitude. Figure 4 As shown, taking a face-to-face battle between players P1 and P2 as an example, player P1's glasses observe key points (POIs) on player P2, generating a ray set {r_i^(P1)}. Simultaneously, player P2's glasses can also observe symmetry information on player P1 (or on nearby reference points). The two players exchange ray information, which affects the overall Fisher information matrix. To extend this, a total Fisher information matrix is constructed through cross-device triangulation fusion. Including the player's own Fisher information matrix Fisher information matrix with the opposing player : (32); Because the angle difference is close to 180° in face-to-face combat, the two sets of rays from both players come from almost orthogonal directions, and the player's own Fisher information matrix... The lack of information in the depth direction is precisely due to the Fisher information matrix of the opposing player. Complete it, and vice versa; The covariance of cross-device triangulation is: (33); After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy. Specifically: At each time k, the system first extracts the two-dimensional coordinates of human key points from the multi-view images. (and (It means the same thing), and combined with the camera geometry, multi-ray triangulation is performed to obtain the triangulation results of the initial estimate of each key point (POI). and its covariance ,Should It takes several 3D rays from multiple camera views, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and obtains the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame using weighted least squares (WLS). The corresponding covariance is then calculated. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the multi-source prior constraint optimization and fusion stage: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeleton connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on... are applied to joints such as elbow, knee, shoulder, and hip... Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the aforementioned multi-source prior constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, which is then used to obtain the optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint (POI) to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable 3D coordinates. and corresponding velocity estimation ; Based on this, the system can further adjust according to uncertainty. The collision detection threshold is adaptively adjusted, and the position of the opponent's root node and the body orientation are inferred from points of interest (POIs) such as the left and right hip joints and left and right shoulder joints.
[0099] The ankle ground constraint refers to the fact that when the player is detected to be standing with both feet on the ground, the ankle joint POI is close to the ground plane, with a height of approximately Z = h0 (the anatomical height of the ankle joint POI from the ground). The horizontal coordinate estimate of the ankle joint POI can be obtained by finding the intersection of a single ray with the horizontal plane Z = h0. Depend on Solving for (34); (35); The energy term constrained by the ground is: (36); Where w_j represents the weight of the ground constraint, which is the inverse of the observation covariance, i.e.: , The uncertainty representing the ankle height observation is usually set to a fixed prior value. A smaller value indicates a more reliable ground constraint, corresponding to a higher weight. The larger the size, the higher its proportion in the integration.
[0100] When the triangulation results When present, the variance of triangulation is weighted and fused inversely proportional to the variance of ground constraints: (37); in, The variance of the triangulation is derived from the inverse of the Fisher information matrix, reflecting the uncertainty of the triangulation, with fixed prior values. The ground constraint variance reflects the measurement accuracy of ankle height h0; Formula (37) is weighted and fused according to the inverse variance of the two. When the depth variance of the triangulation is large (anisotropy caused by the narrow baseline), the ground constraint has a higher weight in the height direction, and vice versa. This constraint is only effective when the player is standing and both feet are on the ground. When jumping, the ankles leave the ground and the system automatically detects and reverts to pure triangulation through inertial sensors or posture classification. The constant bone length constraint means that the distance between adjacent joints should be approximately constant. (38); Where B is the skeleton connection set, For the calibrated or statistically obtained bone lengths (e.g., shoulder → wrist chain length ≈ 0.55m, hip → ankle chain length ≈ 0.82m), Q_k is the constrained optimized 3D keypoint result, and Σ_{Q_k} is the constrained optimized covariance of the corresponding 3D keypoint Q_k; If the triangulation results If the constant bone length constraint is violated, a proportional scaling correction is applied along the joint chain direction to obtain the correction value: (39); The relationship between the correction value and the triangulation result is as follows: taking the midpoint of the two joints, center=(Q_j+Q_k) / 2, as the reference, the unit vector d along the bone direction... =(Q_k-Q_j) / ||Q_k-Q_j||, which scales the bone length proportionally to the standard bone length. ,get , This means normalizing the bone length to a standard value while keeping the midpoint of the joint chain unchanged.
[0101] The basis Joint angle constraints for geodesic distance of rotating manifolds: (40); in, The joint rotation matrix is calculated from the triangulation results. Let the physiological limit boundary of joint j be defined as a submanifold on the rotational manifold SO(3). For the geodesic distance on the rotating manifold SO(3), when the joint rotation is within the physiologically accessible range Inside, = 0, no penalty; when it exceeds the physiologically achievable range, project back to the nearest legal rotation along the geodesic; The above constraint optimization is fused into: (41); The Kalman filter is applied to each keypoint (POI) to smooth the output over time and estimate the joint velocity. The state vector output by the Kalman filter contains six components: position and velocity. ; Observation equations (assuming uniform motion): (42); In the observation equation, the observations are the coordinates of the three-dimensional key points after constraint optimization, and the observation noise covariance is its corresponding three-dimensional covariance Σ_{Q,k} after constraint optimization (which can be simplified to Σ_Q). (43); in, To constrain the optimized 3D covariance, rather than the original triangulation covariance Σ_tri, as the observation noise for the Kalman filter, this means that when the triangulation accuracy is poor in a certain frame ( (large, insufficient information) The filter automatically relies more on prediction and reduces the update amplitude to achieve measurement noise adaptation. The Kalman filter output velocity estimate v can be used for collision force calculation and animation transition without additional calculation. According to uncertainty Adaptive adjustment of the collision determination threshold refers to mapping the three-dimensional uncertainty (quantized by CRLB) to an adaptive threshold for collision determination: (44); in, This is a preset basic collision distance threshold. This is the scaling factor; the more uncertain the positioning ( The larger the threshold, the more conservative the collision detection (the larger the threshold, the less likely it is to be a false alarm); the more accurate the positioning, the more sensitive the detection. This mechanism can also be used in reverse for hit confidence scoring, that is, upgrading collision determination from a hard threshold to a probability assessment: (45); in, For the standard normal CDF, This represents the POI spacing.
[0102] The inference of hand root node position and body orientation from points of interest (POIs) such as left and right hip joints and left and right shoulder joints refers to: Take the midpoint between the left and right hip joints (or left and right shoulder joints) as the root node position (overall position) of the opponent: (46); The yaw angle of body orientation is determined by the intersection of the line connecting the two shoulders with the world horizontal plane. In practical applications, multiple sets of orientation evidence (shoulder, hip, head and chest directions) can be integrated and weighted by Fisher information weights of each POI to improve the stability of orientation estimation. The robust gating and RANSAC achieve outlier ray removal: a hard gating condition is introduced to remove obviously unreliable rays, and RANSAC is used for further purification when the number of rays is sufficient; The robust gating condition is set as follows: (a) The intersection angle α < the preset α_min (it is recommended that α_min = 5°~10°); (b) The nearest point distance ε > the preset ε_max (it is recommended that ε_max = 0.1m); (c) The reprojection residual r_proj > the preset τ_proj (τ_proj = 5~10 pixels is recommended); Any ray that satisfies any of the above robust gating conditions is marked as an outlier and removed. The RANSAC purification process involves first selecting an interior point set using RANSAC (randomly selecting 2 rays for solution and voting on the remaining rays), and then performing least squares estimation (WLS) refinement using only the interior point set. This invention links a robust gating threshold with a Fisher information matrix to form an adaptive gating system: when the total Fisher information matrix... When the minimum eigenvalue is small (poor geometric configuration), the robust gating threshold should be appropriately relaxed to avoid over-rejection; when Fisher information is sufficient, the robust gating threshold should be tightened to improve accuracy. To ensure the system can still provide reasonable output under degraded conditions and automatically return to high-precision mode after conditions recover, this invention defines a normal mode, a degraded mode, and a recovery mode, as well as a method for switching between the three modes: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running the complete triangulation + constraint fusion + Kalman filter pipeline, outputting high-precision results, including α_min. The preset robust gating threshold; When only a single camera is visible, intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model (such as SMPL). When the second view becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. When the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode. This invention provides a mechanism to ensure cross-device consistency verification: when communication is available, players P1 and P2 exchange their estimated low-frequency states (position, orientation, confidence) and construct consistency constraints ‖t1-t1'‖ and ‖φ1-φ1'‖. When there is inconsistency, an alarm is triggered, and the collision judgment weight is reduced, the gating threshold is increased, or a realignment is requested, thereby achieving collaborative verification and anomaly detection. Step S4. The server distributes world location data of relevant key point points (POIs) to all player glasses to achieve limb localization. The world location data of relevant key point points (POIs) can also be used for game rendering or collision detection of player glasses.
[0103] Each pair of glasses can observe and calculate the key points of interest (POIs) of the opposing player, but game rendering or collision detection requires knowing the POIs of all players (including both your own and the opponent's), such as... Figure 7 The diagram shows the information gathering and distribution for players P1 and P2.
[0104] When this invention is extended from 1v1 to multiplayer (such as 2v2, 3v3), each pair of glasses independently performs joint recognition and triangulation for each opposing player within its FOV; the same joint may be observed by multiple pairs of glasses at the same time, and the more observations, the more overdetermined equations are formed by triangulation, and the higher the accuracy; the server distributes the POI data of each player according to their player ID.
[0105] Example 2 Embodiment 2 of this application provides a body positioning system 100 for in-game and out-in-game positioning in XR games, such as... Figure 8 As shown, based on the established world coordinate system, at least two physical anchor points 200 with known coordinates are set up in the field; the XR game includes at least two players wearing XR glasses 300; and a server 400 that interacts with all XR glasses. The XR glasses 300 outputs its own pose through 6DoF self-localization and captures images of the opposing player. It performs 2D human keypoint detection on each frame of the image and outputs the pixel coordinates of each keypoint in the image. It uploads its own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server 400 in real time. At the same time, it also receives the world position data of relevant keypoints sent by the server 400. The world position data of relevant keypoints can also be used for game rendering or collision detection. The server 400 includes a memory 401 and a processor 402. The memory 401 is connected to the processor 402 and is used to store the computer program and the real-time upload of its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player. It also executes the cross-device collaborative triangulation method described in Embodiment 1 to obtain the world position of each key point and distributes the relevant key point world position data to the XR glasses 300 of all players to achieve limb positioning.
[0106] If the game requires information about the orientation of items, based on the camera footage of the opposing player captured by the XR glasses 300, the game further identifies the position of the opposing player holding the item and uploads it to the server 400. (1) At least three positioning points ABC, which are not on the same straight line but on the same plane, are set on the handheld prop. The positioning points refer to signal sources in the outward-to-inward technology. The signal sources include electronic signal transmitters that can be photographed or sensed with built-in ID information, graphics or other identifiable signals. The signal sources are set on the handheld prop in a non-electrical connection manner or implemented in the circuitry of the handheld prop: Establish an XYZ coordinate system for the three positioning points ABC set on the handheld prop. Set the line connecting positioning points A and B as the X-axis, the perpendicular line passing through positioning point C and the X-axis as the Z-axis, set the intersection of the X and Z axes as the origin O, set the orthogonal direction above and below the origin O as the Y-axis, and set the direction of the directed line segment as the forward direction of the handheld prop. (2) The spatial position (X,Y,Z) of each positioning point is obtained based on the camera image of the opposite player captured by the XR glasses 300. (3) The pose data of the handheld prop is obtained by spatial position triangulation of the three positioning points in step (2). The pose data includes yaw angle, pitch angle and roll angle.
[0107] The physical anchor points are placed in asymmetrical positions so that the relative angles of the anchor points are asymmetrical when viewed from any position, thus uniquely determining the orientation of the world coordinate system; the physical anchor points are markers in the corners of the arena, specific patterns on the ground, or other easily visually detectable feature points.
[0108] In some embodiments of this application, the processor 402 may include, but is not limited to: General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0109] In some embodiments of this application, the memory 401 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be a read-only memory (Read-Only Memory). Volatile memory can be ROM, Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), or Flash memory. It can also be Random Access Memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0110] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 401 and executed by the processor 402 to complete the cross-device triangulation method of Embodiment 1 provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program on the server 400.
[0111] Example 3 Embodiment 3 of the present invention also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the method in Embodiment 1 above.
[0112] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for in-between and out-in-between limb positioning in anti-XR games, comprising: establishing a world coordinate system; setting up at least two physical anchor points with known world coordinates in the game field; including at least two players wearing XR glasses in the anti-XR game; and a server that interacts with all XR glasses; the XR glasses using 6DoF self-localization function, utilizing the physical anchor points and environmental features of the field, outputting the pose of the glasses in the world coordinate system; characterized in that... It also includes the following steps: In the game, player P1's glasses camera is pointed directly at the entire body of player P2, and player P2's glasses camera is pointed directly at player P1's entire body. Player P1 uses player P2's XR glasses camera and radar as outward-facing sensors to observe key points on their wrists and ankles from the outside, thereby obtaining the pixel positions of each key point in the image. In other words, the player looks at the opponent through the glasses camera, captures the opponent's camera image, performs 2D human key point detection on each frame of the image, identifies and outputs the pixel coordinates of each key point of the opponent in the image. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposing player to the server in real time. The server uses the observations of at least two XR glasses with known poses. At this time, the XR glasses are equivalent to an outward-inward sensor with a known position. Combined with the pixel coordinates of each key point of the opposing player in the image, the server obtains the world position of each key point of the opposing player through triangulation. The server distributes world location data of relevant key points to all players' glasses to achieve limb positioning.
2. The method for limb positioning from inward to outward in anti-XR games according to claim 1, characterized in that, If the game requires information about the direction of items, the game identifies the position of the opposing player holding the item based on the camera footage captured in step S2: (1) At least three positioning points ABC, which are not on the same straight line but on the same plane, are set on the handheld prop. The positioning points refer to signal sources in the outward-to-inward technology. The signal sources include electronic signal transmitters that can be photographed or sensed with built-in ID information, graphics or other identifiable signals. The signal sources are set on the handheld prop in a non-electrical connection manner or implemented in the circuitry of the handheld prop: Establish an XYZ coordinate system for the three positioning points ABC set on the handheld prop. Set the line connecting positioning points A and B as the X-axis, the perpendicular line passing through positioning point C and the X-axis as the Z-axis, set the intersection of the X and Z axes as the origin O, set the orthogonal direction above and below the origin O as the Y-axis, and set the direction of the directed line segment as the forward direction of the handheld prop. (2) Based on the camera footage of the opposite player collected in step S2, obtain the spatial position (X,Y,Z) of each positioning point. (3) The pose data of the handheld prop is obtained by spatial position triangulation of the three positioning points in step (2). The pose data includes yaw angle, pitch angle and roll angle.
3. The method for limb positioning from inward to outward in anti-XR games according to claim 1, characterized in that, The physical anchor points are placed in asymmetrical positions so that the relative angles of the anchor points are asymmetrical when viewed from any position, thus uniquely determining the orientation of the world coordinate system; the physical anchor points are markers in the corners of the arena, specific patterns on the ground, or other easily visually detectable feature points.
4. The method for limb positioning from inward to outward in anti-XR games according to claim 1, characterized in that, Step S2 identifies and outputs the pixel positions of each key point (POI) of the opposing player in the image, including pixel coordinates (u,v) and detection uncertainty. Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model: The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. , which represents the uncertainty output in this step; The first branch: when the detector outputs a heatmap At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (6); (7); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (8); in, For the preset reference pixel deviation, Here, is the preset lower bound for the minimum deviation, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (8). The two are equal, and the corresponding pixel covariance matrix is... ; Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K as specified by the camera's factory calibration: (9); in, , Focal length in pixels, , ) are the coordinates of the main point; First, convert the pixel coordinates of the target POI to normalized planar coordinates to eliminate the influence of camera intrinsic parameters: (10); When the Z component is normalized to 1, x_n represents the direction of the pixel of the target POI in the camera coordinate system. Its physical meaning is a vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v). Then, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (11); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (12); Where C_w is the world coordinate of the camera's optical center, and λ is the distance parameter traveled along the ray.
5. The method for limb positioning from inward to outward in anti-XR games according to claim 4, characterized in that, In step S3, cross-device collaborative triangulation refers to the process where two cameras in different poses observe the same keypoint POI Q, generating two rays. The closest pair of points P1 on these two rays is then determined. * and P2 * The midpoint is taken as the optimal estimate of POI Q, thus obtaining the world position of each keypoint POI. The specific steps are as follows: Based on the ray equation formula (12), the two rays are expressed as: Ray 1: (13); Ray 2: (14); Where C1 and C2 are the world coordinates of the optical centers of the two cameras, λ1 and λ2 are the distance parameters of the two rays along the rays, and r1 and r2 are the direction vectors of the two rays in the world coordinate system. The condition for a nearest point pair (P1*, P2*) is that the line connecting P1* and P2* is perpendicular to the directions of both rays, i.e.: ,and (15); Substituting formulas (13) and (14) into formula (15), the baseline vector is defined. Expanding, we get: (16); make When r1 and r2 are unit vectors, a = c = 1. This system of equations can be written in matrix form. Its coefficient determinant is Solving using Kramer's rule yields: (17); In face-to-face XR game scenarios, when two cameras observe the same point from different positions, the ray directions cannot be parallel, i.e., a·c - b²≠0. Substitute λ1* and λ2* back into formulas (13) and (14) respectively, and take the midpoint. As an estimate of the world coordinates of the key point POI Q: (18)。 6. The method for limb positioning from inward to outward in anti-XR games according to claim 4, characterized in that, When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POI Q. The goal is to obtain the optimal estimate of keypoint POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
7. The method for limb positioning from inward to outward in anti-XR games according to claim 6, characterized in that, The weighted maximum likelihood estimation for any N rays is specifically as follows: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POI Q. For the i-th ray, with optical center C_i and direction r_i, the squared perpendicular distance from the intersection point Q of multiple rays in three-dimensional space to that ray is: (19); Where I is a 3×3 identity matrix. It is the projection matrix in the direction perpendicular to the ray. The projection matrix is defined. Then formula (19) is ; The goal of multi-ray triangulation is to find a three-dimensional point Q that minimizes the sum of the weighted distances of all rays. (20); Where w_i is the weight; Taking the derivative with respect to intersection point Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. The closed-form solution yields the core formula for cross-device collaborative triangulation: (21); The three-dimensional covariance estimate of the intersection point Q is: (22)。 8. The method for limb positioning from inward to outward in anti-XR games according to claim 7, characterized in that, The theoretically optimal weights w_i are derived using the Fisher information matrix, enabling the least squares estimate WLS to reach the Cramér-Rao lower bound under Gaussian noise, which is the theoretical accuracy limit of the unbiased estimate. If the pixel observation model of the i-th ray is Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ; in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (23); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; The optimal weights that enable least squares estimation (WLS) to reach the Cramér-Rao lower bound (CRLB) are obtained. for: (24); Total Fisher Information Matrix The sum of the information content of each ray: (29); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao.
9. The method for limb positioning from inward to outward in anti-XR games according to claim 8, characterized in that, In actual deployment, formula (24) is decomposed into the multiplication of three independently estimable terms: (25); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. : (26); Where ε0 is a preset constant to prevent division by zero; The geometric mass weight w_geom is determined by the intersection angle. Control, reflecting the geometric contribution of the current ray pair to triangulation, intersection angle The closer to 90°, the greater the information content of the projection matrix A_i in the depth direction. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two-ray projection matrices A_1 + A_2 is proportional to sin²(α / 2). Therefore: (27); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (28); in, For reprojection residuals, The bandwidth parameter used to control the decay rate is equivalent to the Welsch function kernel in the M-estimator.
10. The method for limb positioning from inward to outward in anti-XR games according to claim 8, characterized in that, The cross-device collaborative triangulation uses the XR glasses of multiple game players as a distributed multi-camera array, extends the equivalent baseline to the distance between players, and allows multiple game players to exchange ray information to form the total Fisher information matrix. To extend this, a total Fisher information matrix is constructed through cross-device triangulation fusion. Including the player's own Fisher information matrix Fisher information matrix with the opposing player : (32); The covariance of cross-device triangulation is: (33)。 11. The method for limb positioning from inward to outward in anti-XR games according to claim 6, characterized in that, The phrase "after obtaining cross-device triangulation results, continue to apply constraint fusion and temporal filtering to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy" refers to: At each time k, the system first extracts the two-dimensional coordinates of key human body points from the multi-view images. In conjunction with camera geometry, multi-ray triangulation is performed to obtain the initial triangulation results for each key point point (POI). and its covariance ,Should This involves taking several 3D rays from multiple camera viewpoints, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and then using weighted least squares to obtain the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame, along with the corresponding covariance. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the multi-source prior constraint optimization and fusion stage: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeleton connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on... are applied to the elbow, knee, shoulder, and hip joints. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the aforementioned multi-source prior constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, which is then used to obtain the optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable three-dimensional coordinates. and corresponding velocity estimation .
12. The method for limb positioning from inward to outward in anti-XR games according to claim 6, characterized in that, The aforementioned outlier removal using robust gating and RANSAC refers to introducing hard gating conditions to remove obviously unreliable rays, and using RANSAC purification when the number of rays is sufficient. The robust gating condition is set as follows: (a) The intersection angle α < the preset α_min; (b) The nearest point distance ε > the preset ε_max; (c) The reprojection residual r_proj > the preset τ_proj; Any ray that satisfies any of the above robust gating conditions is marked as an outlier and removed. When the number of available rays is greater than the preset value, first use RANSAC to select the set of interior points, and then perform least squares estimation WLS refinement using only the set of interior points.
13. The method for limb positioning from inward to outward in anti-XR games according to claim 12, characterized in that, The robust gating threshold is linked with the Fisher information matrix to form an adaptive gating: when the minimum eigenvalue of the total Fisher information matrix is less than the threshold, the robust gating threshold is relaxed to avoid over-removal; when Fisher information is sufficient, the robust gating threshold is tightened to improve accuracy.
14. The method for limb positioning from inward to outward in anti-XR games according to claim 12, characterized in that, To ensure the system can still provide reasonable output under degraded conditions, and automatically revert to high-precision mode after conditions recover, the system defines normal mode, degraded mode, and recovery mode, as well as the switching method between the three modes: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running a complete cross-device collaborative triangulation, constraint fusion, and Kalman filtering pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only a single camera is visible, intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model. When the second view becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
15. A body positioning system for an adversarial XR game, comprising: an outward-to-inward limb positioning system based on an established world coordinate system with at least two known physical anchor points in the playing field; the adversarial XR game including at least two players wearing XR glasses; and a server that interacts with all XR glasses; characterized in that... The XR glasses, through 6DoF self-localization, utilize the physical anchor points and environmental features of the venue to output the pose of the glasses in the world coordinate system, and capture images of the opposing player. It performs 2D human keypoint detection on each frame of the image and outputs the pixel coordinates of each keypoint in the image. It uploads its own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server in real time, and at the same time receives the world position data of relevant keypoints sent by the server. The server includes a memory and a processor. The memory, connected to the processor, stores the computer program and the real-time uploads of each XR glasses' pose in the world coordinate system and the pixel coordinates of key points of the opposing player. The processor runs the computer program, using observations from at least two glasses with known poses (each XR glasses acting as an outward-facing sensor with a known position), combined with the pixel coordinates of key points of the opposing player in the image, to obtain the world position of each key point of the opposing player through triangulation, and distributes the relevant world position data of key points to all players' XR glasses to achieve limb localization.
16. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the in-out and out-in limb positioning method for adversarial XR games as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Positioning method of large-space intelligent glasses, head-mounted display equipment and storage medium
CN119672268A
Target device pose calculation method based on outward-in technique and spatial positioning system
CN120742230B
Colocation gesture estimation in shared artificial reality environment
CN114730210A
Efficient somatosensory interaction prop and physical space calibration method thereof
CN119633353A
Collaborative augmented reality eyewear with ego motion alignment
US11587255B1