Cross-device cooperative positioning method and system based on fisher information fusion and storage medium
Patent Information
- Application Number
- CN202610857332.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-12
AI Technical Summary
单设备三角测量固有的短基线约束导致了深度方向定位精度的理论瓶颈,无法满足多人XR对战等场景对毫米级、各向同性、高动态定位能力的需求
[0017] This invention proposes a cross-device collaborative triangulation positioning method that uses the XR glasses of multiple game players as a distributed multi-camera array for collaborative observation of dynamic human key points. A theoretical framework based on the Fisher information matrix is constructed, the optimal weights for cross-device triangulation are derived, and a Cramér-Rao lower bound is established as the theoretical accuracy limit. The equivalent baseline is extended to the distance between players, improving positioning accuracy from centimeters to millimeters. This successfully solves the long-standing technical bottleneck in the XR/AR industry of "poor positioning accuracy in line-of-sight (depth) direction under single-device/narrow baseline conditions," making refined multi-player interaction (such as limb collision detection) possible.
Smart Images

Figure CN122415966B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spatial positioning technology in XR games, specifically relating to a cross-device collaborative positioning method, device, and storage medium based on FISHER information fusion. Background Technology
[0002] In the fields of augmented reality (AR), virtual reality (VR), and mixed reality (MR), achieving high-precision spatial interaction between users and virtual content or remote collaborators is a core technological requirement. Self-localization technology, which estimates the user's own pose (position and posture) in the environment, and technology for spatial localization and tracking of key points of handheld devices or others, are the foundation for achieving immersive interaction.
[0003] Existing visual positioning methods based on single-device triangulation estimate the 3D coordinates of a target point using parallax principles through multiple cameras on the same device (such as binocular or multi-view cameras) or image sequences captured by the same camera at consecutive moments (i.e., moving monocular vision). Single-device triangulation has advantages such as high hardware integration, no need for external base stations, and relatively low cost, and is therefore widely used in consumer AR / VR glasses, smartphone AR, and handheld 3D scanners. For example, many glasses achieve 6DoF (six degrees of freedom) self-localization through built-in binocular cameras and can estimate the coarse spatial position of objects on the hand or table. Each user's device is only used for its own localization and environmental perception, and there is no collaborative measurement between devices. Even if there are multiple user devices in the venue, each device is still limited by its own short baseline and cannot utilize the long baseline observation information provided by other devices to improve its own positioning accuracy. This is a technological status quo of information silos. The inherent short baseline constraint of single-device triangulation leads to a theoretical bottleneck in depth direction positioning accuracy, which cannot meet the requirements of multi-player XR battles and other scenarios for millimeter-level, isotropic, and high dynamic positioning capabilities. Summary of the Invention
[0004] The purpose of this invention is to provide a cross-device collaborative positioning method, device, and storage medium based on FISHER information fusion. This method can actively utilize multiple mobile devices to construct an ultra-long virtual baseline and achieve a new positioning framework that realizes theoretically optimal fusion from an information theory perspective. This improves positioning accuracy from the centimeter level to the millimeter level, thereby meeting the needs of game scenarios with positioning accuracy requirements or fine collision detection requirements.
[0005] This invention discloses a cross-device collaborative positioning method based on FISHER information fusion. Based on an established world coordinate system, at least two physical anchor points with known world coordinates are deployed in the field, including at least two players wearing XR glasses, and a server that interacts with all XR glasses. The positioning method includes the following steps: Step S1. Each pair of glasses uses its built-in 6DoF self-localization function to determine its 6DoF pose in the coordinate system in real time; Step S2. Capture the camera image of the opposite player through the glasses camera, and perform 2D human key point detection on each frame of the image, outputting the pixel coordinates (u,v) of each key point POI of the opposite player in the image and the detection uncertainty; Step S3. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player to the server in real time. The server uses the observations of at least two glasses with known poses and the pixel coordinates of each key point of the player in the image to obtain the world position of each key point through cross-device collaborative triangulation. When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POI Q. The goal is to obtain the optimal estimate of keypoint POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
[0006] Step S2 specifically involves: Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model: The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. This represents the uncertainty output in this step, where the first branch: when the detector outputs a heatmap... At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (3); (4); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (5); in, For the preset reference pixel deviation, Here, is the preset lower bound for the minimum deviation, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (5). The two are equal, and the corresponding pixel covariance matrix is... ; Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K as specified by the camera's factory calibration: (6); in, , The preset focal length in pixels, , ) are the coordinates of the main point; The pixel coordinates of the target POI are converted to normalized planar coordinates to eliminate the influence of intrinsic parameters. (7); Where x_n represents the direction of the pixel of the target POI in the camera coordinate system, and its physical meaning is the vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v); Next, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (8); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (9); Where C_w is the world coordinate of the camera's optical center, and λ is the distance parameter traveled along the ray.
[0007] The weighted maximum likelihood estimation for any N rays is specifically as follows: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POIQ. For the i-th ray, with optical center C_i and direction r_i, the squared perpendicular distance from the intersection point Q of multiple rays in three-dimensional space to that ray is: (10); Where I is a 3×3 identity matrix. It is the projection matrix in the direction perpendicular to the ray. The projection matrix is defined. Then formula (10) is ; Find the 3D point Q that minimizes the sum of the weighted distances of all rays: (11); Where w_i is the weight; Taking the derivative with respect to intersection point Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. The closed-form solution yields the core formula for cross-device collaborative triangulation: (12); The three-dimensional covariance estimate of the intersection point Q is: (13).
[0008] The theoretically optimal weight w_i is derived using the Fisher information matrix, enabling the least squares estimation WLS to reach the Cramere-Rao lower bound under Gaussian noise, i.e., the theoretical accuracy limit of the unbiased estimation: if the pixel observation model of the i-th ray is... Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ,in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (14); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; The optimal weights that allow least squares estimation of WLS to reach the Cramér-Rao lower bound for: (15); Total Fisher Information Matrix The sum of the information content of each ray: (20); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao.
[0009] In actual deployment, formula (15) is decomposed into three independently estimable components multiplied together to adapt to the characteristics of each stage in engineering practice: (16); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. (17); Where ε0 is a preset constant to prevent division by zero; The geometric mass weight w_geom is determined by the intersection angle. The control reflects the geometric contribution of the current ray pair to the triangulation. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two ray projection matrices, A_1 + A_2, is proportional to sin²(α / 2). Therefore: (18); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (19); in, For reprojection residuals, The bandwidth parameter used to control the decay rate is equivalent to the Welsch function kernel in the M-estimator.
[0010] During the cross-device collaborative triangulation, the XR glasses of all players are used as a distributed multi-camera array, the equivalent baseline is extended to the distance between players, and the two players exchange ray information to calculate the total Fisher information matrix. To expand this, a cross-device integrated Fisher information matrix is constructed, including the player's own Fisher information matrix. Fisher information matrix with the opposing player : (twenty three); The player's own Fisher information matrix The lack of information in the depth direction is precisely due to the Fisher information matrix of the opposing player. Complete it, and vice versa; The covariance of cross-device fusion is: (twenty four).
[0011] The aforementioned combination of robust gating and RANSAC for outlier ray removal refers to introducing hard gating conditions to remove obviously unreliable rays, and performing RANSAC purification when the number of rays is sufficient. The gating condition is set such that any ray is marked as an outlier and removed when it meets any of the following conditions: (a) The intersection angle α < the preset α_min; (b) The nearest point distance ε > the preset ε_max; (c) The reprojection residual r_proj > the preset τ_proj; The RANSAC purification process involves selecting an interior point set using RANSAC when the number of available rays exceeds a preset value, and then performing least squares estimation (WLS) refinement using only the interior point set.
[0012] The gating threshold is linked with the Fisher information matrix to form an adaptive gating: when the minimum eigenvalue of the total Fisher information matrix is small, the gating threshold is relaxed to avoid over-removal; when Fisher information is sufficient, the gating threshold is tightened to improve accuracy.
[0013] The phrase "after obtaining cross-device triangulation results, continue to apply constraint fusion and temporal filtering to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy" refers to: At each time k, the system first extracts the two-dimensional coordinates of human key points from the multi-view images. Multi-ray triangulation was performed in conjunction with camera geometry to obtain the initial triangulation results for each key point of interest (POI). and its covariance ,Should It takes several 3D rays from multiple camera views, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and obtains the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame using weighted least squares (WLS). The corresponding covariance is then calculated. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the constraint optimization fusion phase: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeletal connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on the elbow, knee, shoulder, and hip joints are applied. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the above constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, and solves it to obtain the constrained optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint (POI) to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable 3D coordinates. and corresponding velocity estimation .
[0014] To ensure the system can still provide reasonable output under degraded conditions, and automatically revert to high-precision mode after conditions recover, three modes are defined: normal mode, degraded mode, and recovery mode, along with their switching methods: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running the complete triangulation, constraint fusion, and Kalman filtering pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only a single camera is visible, intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model. When the second view becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
[0015] Furthermore, this invention discloses a cross-device collaborative positioning system based on FISHER information fusion. Based on an established world coordinate system, at least two physical anchor points with known coordinates are deployed in the field; the system includes at least two players wearing XR glasses; and a server that interacts with all XR glasses. The XR glasses output their own pose through 6DoF self-localization and capture images of the opposing player. They perform 2D human keypoint detection on each frame of the image and output the pixel coordinates of each keypoint in the image. They upload their own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server in real time, and also receive the world position data of relevant keypoints sent by the server. The server includes a memory and a processor. The memory is connected to the processor and is used to store the pose of the computer program and the XR glasses in the world coordinate system in real time, as well as the pixel coordinates of each key point of the opposite player. The server executes the cross-device collaborative positioning method based on FISHER information fusion to obtain the world position of each key point and distributes the relevant key point world position data to the XR glasses of all players.
[0016] Furthermore, the present invention discloses a storage medium storing a computer program, which, when executed by a processor, implements the aforementioned cross-device collaborative positioning method based on FISHER information fusion.
[0017] This invention proposes a cross-device collaborative triangulation positioning method that uses the XR glasses of multiple game players as a distributed multi-camera array for collaborative observation of dynamic human key points. A theoretical framework based on the Fisher information matrix is constructed, the optimal weights for cross-device triangulation are derived, and a Cramér-Rao lower bound is established as the theoretical accuracy limit. The equivalent baseline is extended to the distance between players, improving positioning accuracy from centimeters to millimeters. This successfully solves the long-standing technical bottleneck in the XR / AR industry of "poor positioning accuracy in line-of-sight (depth) direction under single-device / narrow baseline conditions," making refined multi-player interaction (such as limb collision detection) possible.
[0018] Furthermore, this invention transforms the triangulation problem into a weighted least squares problem and provides a closed-form solution. It only requires calculating a 3×3 matrix A_i for each ray and accumulating the results. Finally, the result can be obtained by performing a 3×3 matrix inversion once. The computational complexity is O(N), which fully meets the requirements of real-time operation on the end side.
[0019] Furthermore, in this invention, the game parties exchange ray information to form a cross-device fused total Fisher information matrix. The information missing in the depth direction of a player's own Fisher information matrix is precisely filled by the Fisher information matrix of the opposing player, and vice versa, thereby compensating for the information missing in the depth direction (view direction).
[0020] Furthermore, this invention integrates multiple physical constraints (geodetic distance constraints on the SO(3) manifold, including ground, bone length, and joint angles) and temporal filtering (Kalman filtering) to form a complete space-time joint optimization framework, which is used to improve depth accuracy. This solves the problem that when triangulation is performed, the 3D coordinates of each key point POI are calculated independently frame by frame without considering the temporal continuity between frames and the physical constraint relationship between joints.
[0021] Furthermore, this invention also designs three modes—normal, degraded, and recovery—and their switching mechanisms to ensure the robustness of the system in real-world complex scenarios. Attached Figure Description
[0022] Figure 1 This is a flowchart of the method in this invention; Figure 2 This is an analysis curve showing the impact of key parameters on positioning accuracy in this invention; Figure 3 This is a schematic diagram of the cross-device fusion collaborative triangulation principle in this invention; Figure 4 This is a flowchart of the constraint fusion and temporal filtering process in this invention; Figure 5 This is a schematic diagram of the three-state mode switching mechanism in this invention; Figure 6 This is a schematic diagram of data broadcasting between player P1 and player P2 in this invention; Figure 7 This is a system block diagram of the cross-device collaborative positioning system of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0025] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or solution described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Example 1 In Embodiment 1, this invention discloses a cross-device collaborative positioning method based on FISHER information fusion, such as... Figure 1 As shown, the system covers binocular, multi-view, and multi-frame fusion scenes; a world coordinate system W is established, with its origin O located at the center of the game arena, the X-axis pointing horizontally to the right, the Y-axis pointing horizontally forward, and the Z-axis pointing vertically upward. At least two physical anchor points O1 and O2 with known world coordinates are placed in the arena; these can be markers in the corners of the arena, specific patterns on the ground, or other easily visually detectable feature points; the system includes at least two players wearing XR glasses; and a server that interacts with all XR glasses; the positioning method includes the following steps: Step S1. After establishing the world coordinate system, each pair of glasses uses its built-in 6DoF self-localization function to determine its 6DoF pose and orientation in the coordinate system in real time. Define the global world coordinate system W using the set of anchor points: take the ground as the Z=0 reference plane, with the Z-axis pointing vertically upwards (normal vector n=[0,0,1)). The horizontal direction is determined by the anchor points. Specifically, the ground plane equation and coordinate axes are defined as follows: ,in, , d = 0 (1); , (2); Where z_w = n = [0, 0, 1] The vertically upward direction is indicated by z_w. If the anchor point or the IMU in the glasses provides the direction of gravity, then z_w is replaced with its normalized opposite direction of gravity. Subscripts are uniformly represented by _* in the text. Step S2. Look at the opposite player through the glasses camera, capture the camera image of the opposite player, and perform 2D human key point detection on each frame of the image, outputting the pixel coordinates (u,v) of each POI of the opposite player in the image and the detection uncertainty; Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model (used to provide source data for subsequent weight construction and end-to-end uncertainty propagation): The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. This represents the uncertainty output in this step, where the first branch: when the detector outputs a heatmap... At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (3); (4); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (5); in, The baseline pixel deviation is approximately 3.0 pixels. The minimum deviation lower bound is 0.5 pixels, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation (bias) in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (5). The two are equal, and the corresponding pixel covariance matrix is... In this way, high-confidence detection yields small covariance (high weight), while low-confidence detection yields large covariance (low weight), enabling uncertainty-driven adaptive weight adjustment in subsequent triangulation measurements.
[0026] Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K (given by the camera's factory calibration): (6); in, , The preset focal length in pixels, , The coordinates of the principal point are ideally located at the center of the image. The camera intrinsic matrix K describes the mapping relationship between the three-dimensional spatial orientation and the image pixel coordinates, serving as a bridge connecting the two worlds. For lenses with distortion, distortion correction must be performed before backprojection.
[0027] First, the pixel coordinates of the target POI are converted to normalized planar coordinates to eliminate the influence of intrinsic parameters: (7); Where x_n represents the direction of the pixel of the target POI in the camera coordinate system (Z component normalized to 1), and its physical meaning is the vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v). Next, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (8); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (9); Where C_w is the world coordinate of the camera's optical center, λ is the distance parameter along the ray, and the target POI is located at a certain λ value on this ray. The specific λ value cannot be determined by a single ray alone; multiple rays (at least two) are needed to determine it.
[0028] Step S3. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player to the server in real time. The server uses the observations of at least two glasses with known poses and the pixel coordinates of each key point of the player in the image to obtain the world position of each key point through cross-device collaborative triangulation. When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint (POI) Q. The goal is to obtain the optimal estimate of POI Q. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation (N≥2). Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint (POI) is obtained, specifically: In cross-device collaborative (multi-ray) triangulation, N cameras or cameras in different poses observe the same key point POI Q. For the i-th ray, with optical center C_i and direction r_i, the squared perpendicular distance from the intersection point Q of the multiple rays in three-dimensional space to the ray is: (10); Where I is a 3×3 identity matrix. It is the projection matrix (symmetric and idempotent) in the direction perpendicular to the ray. Define the projection matrix. Then formula (10) is ; Find the 3D point Q that minimizes the sum of the weighted distances of all rays: (11); Where w_i is the weight; This is a quadratic optimization problem about Q. Taking the derivative of Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. Closed-form solution: (12); The three-dimensional covariance estimate of the intersection point Q is: (13); Formula (12) is the core formula for cross-device collaborative triangulation. It only requires calculating a 3×3 matrix A_i for each ray and accumulating the results, and finally performing a 3×3 matrix inversion once to obtain the result. The computational complexity is O(N), which fully meets the requirements of real-time operation on the end side.
[0029] In classical methods, the design of weights w_i relies on engineering experience. This invention, starting from statistical estimation theory, uses the Fisher Information Matrix (FIM) to derive the theoretically optimal weights w_i, enabling the least squares estimation (WLS) to reach the Cramér-Rao lower bound (CRLB) under Gaussian noise, i.e., the theoretical accuracy limit of unbiased estimation. If the pixel observation model of the i-th ray is... Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ,in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (14); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; As can be seen from formula (14), the information content of each ray is determined by three factors: focal length f (the longer the focal length, the more information), distance d_i (the closer the distance, the more information), and detection accuracy. (The more precise the information, the more information there is), and the spatial orientation structure of the information is entirely determined by the projection matrix A_i (which only provides information in the direction perpendicular to the ray).
[0030] Therefore, the optimal weights that allow the least squares estimate (WLS) to reach the Cramér-Rao lower bound (CRLB) are... for: (15); The optimal weight By unifying distance attenuation and detection uncertainty within an information-theoretic framework, there is no need to manually select the attenuation function or adjust hyperparameters. In practical engineering implementations, d_i can be approximated by the estimation result of the previous frame, forming an iterative optimization.
[0031] In actual deployment, formula (15) is further decomposed into three independently estimable components multiplied together to adapt to the characteristics of each stage in engineering practice: (16); The definitions of each item are as follows: The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. (17); Wherein, ε0 is a preset constant to prevent division by zero (generally set to 0.01 pixels²). Geometric mass weight w_geom: determined by the intersection angle Control, reflecting the geometric contribution of the current ray pair to triangulation, intersection angle The closer to 90°, the greater the information content of the projection matrix A_i in the depth direction. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two-ray projection matrices A_1 + A_2 is proportional to sin²(α / 2), therefore: (18); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; Reprojection consistency weight w_reproj: using triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (19); in, For reprojection residuals, To control the bandwidth parameter of decay rate, this reprojection consistency weight is equivalent to the Welsch function kernel in the M-estimator, giving the WLS framework robust statistical significance.
[0032] Total Fisher Information Matrix The sum of the information content of each ray: (20); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao, CRLB; From this, the following accuracy rules can be derived: A. Ray direction error With pixel detection error The relationship is: (twenty one); B. Three-dimensional positioning error The upper limit of the engineering quantity level is: (twenty two); Where B is the equivalent baseline length, D is the target distance, and α is the equivalent intersection angle. For example... Figure 2 As shown, the formula clearly reveals the influence of each parameter on positioning accuracy: the larger the baseline B, the better; the closer the distance D, the better; the closer the intersection angle α is to 90°, the better; and the more accurate the pixel detection, the better.
[0033] Accuracy Analysis and Typical Parameter Estimation: For a single-device binocular baseline B=0.064m and target distance D=1.83m, the actual intersection angle α is only about 2°, and the upper bound of the error given by formula (22) is relatively large. However, this is a worst-case estimate, and the actual error is significantly smaller than the upper bound due to the randomness of the noise direction. More importantly, formula (22) reveals the inherent accuracy bottleneck of single-device narrow baseline triangulation—the error along the depth direction (line of sight) is much larger than the lateral error, and this anisotropy is manifested by the eigenvalue structure of CRLB.
[0034] Typical lateral accuracy estimation: With D=1.83m, f=500 pixels, For example, =2 pixels: ; Since the lateral accuracy is on the order of millimeters, it is sufficient for determining whether a punch has hit the opponent's body (lateral judgment) in combat XR games. The accuracy in the depth direction needs to be improved through cross-device collaborative triangulation fusion and constraint fusion proposed in this invention. The baseline of a single-device binocular glasses is only 6.4cm. When the target distance exceeds 1m, the triangulation accuracy in the depth direction is limited by a very small intersection angle (α<3°). The cross-device collaborative triangulation fusion mechanism proposed in this invention uses the XR glasses of two (or even multiple) players as a distributed multi-camera array, extending the equivalent baseline from 6.4cm to the player spacing (1.5m~3m), thereby improving depth accuracy by one to two orders of magnitude.
[0035] like Figure 3 As shown, taking a face-to-face battle between players P1 and P2 as an example, player P1's glasses observe key points (POIs) on player P2, generating a ray set {r_i^(P1)}. Simultaneously, player P2's glasses can also observe symmetry information on player P1 (or on nearby reference points). The two players exchange ray information, which affects the overall Fisher information matrix. To expand this, a cross-device integrated Fisher information matrix is constructed, including the player's own Fisher information matrix. Fisher information matrix with the opposing player : (twenty three); The player's own Fisher information matrix The lack of information in the depth direction is precisely due to the Fisher information matrix of the opposing player. Complete it, and vice versa; The covariance of cross-device fusion is: (twenty four); Based on typical parameters: player spacing B_cross = 3m, D = 1.5m, f = 500 pixels, Taking 2 pixels as an example, the equivalent intersection angle is close to 90°: ; Cross-device collaborative triangulation improves positioning accuracy from the centimeter level to the millimeter level, fully meeting the needs of fine collision detection.
[0036] The robust gating and RANSAC achieve outlier ray removal: a hard gating condition is introduced to remove obviously unreliable rays, and RANSAC is used for further purification when the number of rays is sufficient; The gating condition is set such that any ray is marked as an outlier and removed when it meets any of the following conditions: (a) The intersection angle α < the preset α_min (it is recommended that α_min = 5°~10°); (b) The nearest point distance ε > the preset ε_max (it is recommended that ε_max = 0.1m); (c) The reprojection residual r_proj > the preset τ_proj (τ_proj = 5~10 pixels is recommended).
[0037] The RANSAC purification mechanism works as follows: when the number of available rays exceeds a preset value, RANSAC is first used to select an interior point set (two rays are randomly selected for solving, and the remaining rays are voted on), and then only the interior point set is used to perform least squares estimation (WLS) refinement. This mechanism is particularly effective in cases of extreme detection errors or abrupt pose changes.
[0038] This invention links the gating threshold with the Fisher information matrix to form adaptive gating: when the total Fisher information matrix... When the minimum eigenvalue is small (poor geometric configuration), the gating threshold should be appropriately relaxed to avoid over-rejection; when Fisher information is sufficient, the gating threshold should be tightened to improve accuracy.
[0039] In traditional triangulation, the 3D coordinates of each Point of Interest (POI) are calculated independently frame by frame, without considering the temporal continuity between frames or the physical constraints between joints. This invention addresses this deficiency by fusing multi-source prior constraints and temporal filtering. The temporal filtering and constraint fusion, as described above... Figure 4 As shown, a joint optimization along the spatial and temporal dimensions is achieved, specifically as follows: At each time k, the system first extracts the two-dimensional coordinates of human key points from the multi-view images. (and (It means the same thing), and combined with the camera geometry, multi-ray triangulation is performed to obtain the triangulation results of the initial estimate of each key point (POI). and its covariance ,Should It takes several 3D rays from multiple camera views, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and obtains the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame using weighted least squares (WLS). The corresponding covariance is then calculated. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the constraint optimization fusion phase: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeleton connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on the following are applied to joints such as elbows, knees, shoulders, and hips. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the above constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, and solves it to obtain the constrained optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint (POI) to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable 3D coordinates. and corresponding velocity estimation ; Based on this, the system can further adjust according to uncertainty. The collision detection threshold is adaptively adjusted, and the position of the opponent's root node and the body orientation are inferred from points of interest (POIs) such as the left and right hip joints and left and right shoulder joints.
[0040] The ankle ground constraint refers to the fact that when the player is detected to be standing with both feet on the ground, the ankle joint POI is pressed against the ground plane at a height of approximately Z = h0 (h0 ≈ 0.08m, i.e., the anatomical height of the ankle joint POI from the ground). The horizontal coordinate estimate of the ankle joint POI can be obtained by finding the intersection of a single ray with the horizontal plane Z = h0. Depend on Solving for (25); (26); The energy term constrained by the ground is: (27); Where w_j represents the weight of the ground constraint, which is the inverse of the observation covariance, i.e.: , The uncertainty representing the ankle height observation is usually set to a fixed prior value. A smaller value indicates a more reliable ground constraint, corresponding to a higher weight. The larger the size, the higher its proportion in the integration.
[0041] When the triangulation results When present, the variance of triangulation is weighted and fused inversely proportional to the variance of ground constraints: (28); in, The variance of the triangulation (from the inverse of the Fisher information matrix) reflects the uncertainty of the triangulation. The ground constraint variance (fixed prior value, such as σ²_ground≈(0.02m)²) reflects the measurement accuracy of ankle height h0; Formula (28) is weighted and fused by the inverse variance of the two. When the depth variance of the triangulation is large (due to anisotropy caused by narrow baseline), the ground constraint has a higher weight in the height direction; and vice versa. It should be noted that this constraint is only effective when the player is standing with both feet on the ground. When jumping, the ankles leave the ground, and the system automatically detects this through inertial sensors or posture classification and reverts to pure triangulation.
[0042] Triangulation calculations are performed independently for each joint, without considering the physical connections between joints, which may lead to inconsistent bone lengths. This invention introduces a constant bone length constraint, meaning the distance between adjacent joints should be approximately constant (human bone length does not change between frames): (29); Where B is the skeleton connection set, For the calibrated or statistically obtained bone lengths (e.g., shoulder → wrist chain length ≈ 0.55m, hip → ankle chain length ≈ 0.82m), Q_k is the constrained optimized 3D keypoint result, and Σ_{Q_k} is the optimized covariance of the corresponding 3D keypoint Q_k; Note: The covariance of all 3D keypoints after the above constraint optimization is written as Σ_{Q_*}; then it is written as Σ_{Q_k} at a single keypoint, and is abbreviated as Σ_Q when the keypoint subscript is not emphasized.
[0043] If the triangulation results If the constant bone length constraint is violated, a proportional scaling correction is applied along the joint chain direction to obtain the correction value: (30); The relationship between the correction value and the triangulation result is as follows: taking the midpoint of the two joints, center=(Q_j+Q_k) / 2, as the reference, the unit vector d along the bone direction... =(Q_k-Q_j) / ||Q_k-Q_j||, which scales the bone length proportionally to the standard bone length. ,get , This means normalizing the bone length to a standard value while keeping the midpoint of the joint chain unchanged.
[0044] Traditional methods impose joint angle constraints in Euclidean space, but joint rotation is naturally defined in rotation groups. On the manifold, this invention uses geodesic distance on the Riemannian manifold instead of Euclidean distance as the constraint cost function to avoid gimbal lock and singularity problems.
[0045] The basis Joint angle constraints for geodesic distance of rotating manifolds: (31); in, The joint rotation matrix is calculated from the triangulation results. Let the physiological limit boundary of joint j be defined as a submanifold on the rotational manifold SO(3). For the geodesic distance on the rotating manifold SO(3), when the joint rotation is within the physiologically accessible range Inside, = 0, no penalty; when it exceeds the physiologically achievable range, project back to the nearest legal rotation along the geodesic.
[0046] For the elbow joint (flexion-extension range 0°–145°) and the knee joint (flexion-extension range 0°–135°), it can be degenerated into a one-dimensional case: In this case, it is consistent with the classical Euclidean constraint; however, for multi-degree-of-freedom ball-and-socket joints such as the shoulder and hip joints, the Riemannian formulation avoids Euler angle singularities and provides geometrically more accurate constraint propagation.
[0047] The above constraint optimization is fused into: (32).
[0048] The aforementioned constraint optimization fusion optimizes the results for each frame in the spatial dimension, but jitter may still exist between frames. This invention maintains a Kalman filter for each keypoint (POI), smooths the output in the temporal dimension, and estimates joint velocity. The state vector output by the Kalman filter contains six components: position and velocity. .
[0049] Observation equations (assuming uniform motion): (33); In the observation equation, the observations are the coordinates of the three-dimensional key points after constraint optimization, and the observation noise covariance is its corresponding three-dimensional covariance Σ_{Q,k} after constraint optimization (which can be simplified to Σ_Q). (34); in, To constrain the optimized 3D covariance, rather than the original triangulation covariance Σ_tri, as the observation noise for the Kalman filter, this means that when the triangulation accuracy is poor in a certain frame ( (Large and insufficient information) The filter automatically relies more on prediction and reduces the update amplitude to achieve measurement noise adaptation. The Kalman filter output velocity estimate v can be used for collision force calculation (fist speed affects damage value) and animation transition without additional calculation.
[0050] According to uncertainty Adaptive adjustment of the collision determination threshold refers to mapping the three-dimensional uncertainty (quantized by CRLB) to an adaptive threshold for collision determination: (35); in, The preset basic collision distance threshold (e.g., 0.05m). This is the scaling factor.
[0051] This means that the more uncertain the positioning ( The larger the threshold, the more conservative the collision detection (the larger the threshold), which can reduce false alarms; the more accurate the positioning, the more sensitive the detection.
[0052] This mechanism can also be used in reverse for hit confidence scoring, that is, upgrading collision determination from a hard threshold to a probability assessment: (36); in, For the standard normal CDF, POI spacing; The inference of hand root node position and body orientation from points of interest (POIs) such as left and right hip joints and left and right shoulder joints refers to: The overall position and orientation of the opponent can be inferred from the set of three-dimensional coordinates of multiple POIs, which is very important for game rendering and tactical judgment.
[0053] Take the midpoint between the left and right hip joints (or left and right shoulder joints) as the root node position (overall position) of the opponent: (37); The yaw angle of body orientation is determined by the intersection of the line connecting the two shoulders with the world horizontal plane. In practical applications, multiple sets of orientation evidence (shoulder, hip, head and chest orientation) can be integrated and weighted by Fisher information weights of each POI to improve the stability of orientation estimation.
[0054] In real-world competitive scenarios, various degradation conditions are unavoidable. To ensure that the system can still provide reasonable output under degradation conditions and automatically revert to high-precision mode after the conditions recover, this invention defines, for example... Figure 5 The three modes and their switching methods are shown below: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running the complete triangulation, constraint fusion, and Kalman filtering pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only one camera is visible (the other party is turned to the side, so only one type of eye can see the joint), intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate information of visible joints to invisible joints; (3) optionally, the 3D pose is directly recovered from monocular 2D detection as a weak observation using a parameterized human model (such as SMPL). When the second viewpoint becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
[0055] Furthermore, this invention provides a mechanism to ensure cross-device consistency verification: when communication is available, players P1 and P2 exchange their estimated low-frequency states (position, orientation, confidence) and construct consistency constraints ‖t1-t1'‖ and ‖φ1-φ1'‖. When there is inconsistency, an alarm is triggered: the collision determination weight is reduced, the gating threshold is increased, or a realignment is requested, thereby achieving collaborative verification and anomaly detection.
[0056] Step S4. The server distributes the world location data of relevant keypoints (POIs) to all player glasses. This world location data of keypoints (POIs) can also be used for game rendering or collision detection for player glasses. Each pair of glasses can observe and calculate the world positions of the opponent's key points of interest (POIs), but game rendering or collision detection requires knowing the world positions of all players' POIs (including their own and the opponent's), such as... Figure 6 The diagram shows the information gathering and distribution for players P1 and P2.
[0057] When this invention is extended from 1v1 to multiplayer (such as 2v2, 3v3), each pair of glasses independently performs key point POI identification and triangulation for each opposing player within its FOV; the same joint may be observed by multiple pairs of glasses at the same time, and the more observations, the more overdetermined equations are formed by triangulation, and the higher the accuracy; the server distributes the world position data of each player's relevant key point POI according to the player ID.
[0058] Example 2 Embodiment 2 of this application provides a cross-device collaborative positioning system 100 based on FISHER information fusion, such as... Figure 7 As shown, based on the established world coordinate system, at least two physical anchor points 200 with known coordinates are set up; at least two players wearing XR glasses 300 are included; and a server 400 that interacts with all XR glasses; The XR glasses 300 outputs its own pose through 6DoF self-localization and captures images of the opposing player. It performs 2D human keypoint detection on each frame of the image and outputs the pixel coordinates of each keypoint in the image. It uploads its own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server 400 in real time. At the same time, it also receives the world position data of relevant keypoints sent by the server 400 for game rendering or collision detection. The server 400 includes a memory 401 and a processor 402. The memory 401 is connected to the processor 402 and is used to store the computer program and the real-time upload of its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player. The server executes the cross-device collaborative positioning method described in Embodiment 1 to obtain the world position of each key point and distributes the relevant key point world position data to the XR glasses 300 of all players.
[0059] The physical anchor points are placed in asymmetrical positions so that the relative angles of the anchor points are asymmetrical when viewed from any position, thus uniquely determining the orientation of the world coordinate system; the physical anchor points are markers in the corners of the arena, specific patterns on the ground, or other easily visually detectable feature points.
[0060] In some embodiments of this application, the processor 402 may include, but is not limited to: General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0061] In some embodiments of this application, the memory 401 includes, but is not limited to, volatile memory and / or non-volatile memory. The non-volatile memory may be a read-only memory (Read-Only Memory). Volatile memory can be ROM, Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), or Flash memory. It can also be Random Access Memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0062] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 401 and executed by the processor 402 to complete the cross-device collaborative positioning method of Embodiment 1 provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program on the server 400.
[0063] Example 3 Embodiment 3 of the present invention also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the method in Embodiment 1 above.
[0064] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A cross-device collaborative positioning method based on FISHER information fusion, which, based on an established world coordinate system, deploys at least two physical anchor points with known world coordinates in the field, including at least two players wearing XR glasses, and a server that interacts with all XR glasses; characterized in that, The positioning method includes the following steps: Step S1. Each pair of glasses uses its built-in 6DoF self-localization function to determine its 6DoF pose in the coordinate system in real time; Step S2. Capture the camera image of the opposite player through the glasses camera, and perform 2D human key point detection on each frame of the image, outputting the pixel coordinates (u,v) of each key point POI of the opposite player in the image and the detection uncertainty; Step S3. Each pair of glasses uploads its own pose in the world coordinate system and the pixel coordinates of each key point of the opposite player to the server in real time. The server uses the observations of at least two glasses with known poses and the pixel coordinates of each key point of the player in the image to obtain the world position of each key point through cross-device collaborative triangulation. When performing cross-device collaborative triangulation, N cameras or cameras in different poses observe the same keypoint POIQ. The goal is to obtain the optimal estimate of the keypoint POIQ. Weighted maximum likelihood estimation is performed on any N rays to achieve multi-ray triangulation, where N≥2. Then, robust gating and RANSAC are used to remove outliers, and the theoretically optimal weights are derived using the Fisher information matrix to establish the Cramér-Rao lower bound quantization theoretical accuracy limit. After obtaining the cross-device triangulation results, constraint fusion and temporal filtering are applied to form a joint optimization of the spatial and temporal dimensions to compensate for depth accuracy. Finally, the world position of each keypoint POI is obtained.
2. The cross-device collaborative positioning method based on FISHER information fusion according to claim 1, characterized in that, Step S2 specifically involves: Step S21. In multi-person scenarios, use the existing pedestrian re-identification model to track pedestrian IDs to ensure cross-frame consistency. For each detected POIm of the target pedestrian, the pixel coordinates are u^(m) = (u,v), and the detection confidence is s, s∈[0,1]. Step S22. Detect uncertainty using the constructed uncertainty estimation model: The uncertainty estimation model consists of two branches, both of which output pixel covariance matrices. This is the uncertainty output in this step, where the first branch: when the detector outputs a heatmap At that time, the pixel covariance matrix is calculated from the second moment of the heatmap. : (3); (4); in, Iterate through all pixel locations in the heatmap. For the corresponding probability value; Second branch: When only the detection confidence s is available, construct an isotropic approximate pixel covariance matrix. : (5); in, For the preset reference pixel deviation, Here, is the preset lower bound for the minimum deviation, and s is the detection confidence level. This represents the standard deviation of pixels in the u-direction at pixel coordinates (u, v). The pixel standard deviation in the v direction of pixel coordinates (u,v) is calculated by the isotropic approximation of formula (5). The two are equal, and the corresponding pixel covariance matrix is... ; Step S23. For a given pixel coordinate (u,v) of a target POI, calculate the equation of the ray equation originating from the camera optical center and passing through the corresponding spatial direction of the target POI pixel: Define the camera intrinsic parameter matrix K as specified by the camera's factory calibration: (6); in, , The preset focal length in pixels, , () are the coordinates of the main point; The pixel coordinates of the target POI are converted to normalized planar coordinates to eliminate the influence of intrinsic parameters. (7); Where x_n represents the direction of the pixel of the target POI in the camera coordinate system, and its physical meaning is the vector pointing from the camera optical center to the spatial direction corresponding to pixel (u,v); Next, the orientation in the camera coordinate system is transformed to the world coordinate system using the rotation matrix R_wc: (8); Where, normalize(·) is the normalization operation, r is the direction vector of the ray in the world coordinate system, and R_wc is the rotation matrix from the camera frame to the world frame; The equation of the ray in the world coordinate system is: (9); Where C_w is the world coordinate of the camera's optical center, and λ is the distance parameter traveled along the ray.
3. The cross-device collaborative positioning method based on FISHER information fusion according to claim 2, characterized in that, The weighted maximum likelihood estimation for any N rays is specifically as follows: In cross-device collaborative triangulation, N cameras or cameras in different poses observe the same key point POI Q. For the i-th ray, with optical center C_i and direction r_i, the squared perpendicular distance from the intersection point Q of multiple rays in three-dimensional space to that ray is: (10); Where I is a 3×3 identity matrix. It is the projection matrix in the direction perpendicular to the ray. The projection matrix is defined. Then formula (10) is ; Find the 3D point Q that minimizes the sum of the weighted distances of all rays: (11); Where w_i is the weight; Taking the derivative with respect to intersection point Q and setting the gradient to zero, we obtain the estimated world coordinates of the keypoint POI Q. The closed-form solution yields the core formula for cross-device collaborative triangulation: (12); The three-dimensional covariance estimate of the intersection point Q is: (13)。 4. The cross-device collaborative positioning method based on FISHER information fusion according to claim 3, characterized in that, The theoretically optimal weight w_i is derived using the Fisher information matrix, enabling the least squares estimation WLS to reach the Cramere-Rao lower bound under Gaussian noise, i.e., the theoretical accuracy limit of the unbiased estimation: if the pixel observation model of the i-th ray is... Where n_i represents the pixel observation noise term, and Then the pixel observation value u_i of the i-th ray follows a formula with the true projection point as the mean and the mean as the mean. The pixel covariance matrix is a two-dimensional Gaussian distribution: ,in The pixel covariance matrix estimated in step S22; Through Jacobi propagation, the Fisher information matrix of the i-th ray to point Q is: (14); Where f is the focal length. Let Q be the distance from the optical center of the i-th camera. Let be the pixel covariance matrix of the i-th detection. Representative matrix The sum of the diagonal elements; The optimal weights that allow least squares estimation of WLS to reach the Cramér-Rao lower bound for: (15); Total Fisher Information Matrix The sum of the information content of each ray: (20); in, It is a two-dimensional pixel covariance matrix Scalared equivalent pixel variance; Establish the theoretical lower bound for multi-ray triangulation. Any unbiased estimator The covariance matrix satisfies the matrix inequality: That is, the lower bound of Cramér-Rao.
5. The cross-device collaborative positioning method based on FISHER information fusion according to claim 4, characterized in that, In actual deployment, formula (15) is decomposed into three independently estimable components multiplied together to adapt to the characteristics of each stage in engineering practice: (16); Where w_conf is the detection confidence weight, w_geom is the geometric quality weight, and w_reproj is the reprojection consistency weight; The detection confidence weight w_conf is driven by the pixel covariance Σ_uv output in step S22 and directly corresponds to the detection accuracy factor in the Fisher information. (17); Where ε0 is a preset constant to prevent division by zero; The geometric mass weight w_geom is determined by the intersection angle. The control reflects the geometric contribution of the current ray pair to the triangulation. From the perspective of the Fisher information matrix, the minimum eigenvalue of the superposition of the two ray projection matrices, A_1 + A_2, is proportional to sin²(α / 2). Therefore: (18); When the intersection angle α is too small, w_geom approaches zero, automatically reducing the contribution of the ray pair; The reprojection consistency weight w_reproj: utilizes triangulation results Backprojecting back to the image plane and checking for consistency with the original detection achieves iterative soft outlier suppression. (19); in, For reprojection residuals, The bandwidth parameter used to control the decay rate is equivalent to the Welsch function kernel in the M-estimator.
6. The cross-device collaborative positioning method based on FISHER information fusion according to claim 4, characterized in that, During the cross-device collaborative triangulation, the XR glasses of all players are used as a distributed multi-camera array, the equivalent baseline is extended to the distance between players, and the two players exchange ray information to calculate the total Fisher information matrix. To expand this, a cross-device integrated Fisher information matrix is constructed, including the player's own Fisher information matrix. Fisher information matrix with the opposing player : (23); The player's own Fisher information matrix The lack of information in the depth direction is precisely due to the Fisher information matrix of the opposing player. Complete it, and vice versa; The covariance of cross-device fusion is: (24)。 7. The cross-device collaborative positioning method based on FISHER information fusion according to claim 1, characterized in that... The aforementioned outlier removal using robust gating and RANSAC refers to introducing hard gating conditions to remove obviously unreliable rays, and performing RANSAC purification when the number of rays is sufficient. The gating condition is set such that any ray is marked as an outlier and discarded when it meets any of the following conditions: (a) The intersection angle α < the preset α_min; (b) The nearest point distance ε > the preset ε_max; (c) The reprojection residual r_proj > the preset τ_proj; The RANSAC purification process involves selecting an interior point set using RANSAC when the number of available rays exceeds a preset value, and then performing least squares estimation (WLS) refinement using only the interior point set.
8. The cross-device collaborative positioning method based on FISHER information fusion according to claim 7, characterized in that, The gating threshold is linked with the Fisher information matrix to form an adaptive gating: when the minimum eigenvalue of the total Fisher information matrix is small, the gating threshold is relaxed to avoid over-removal; when Fisher information is sufficient, the gating threshold is tightened to improve accuracy.
9. The cross-device collaborative positioning method based on FISHER information fusion according to claim 1, characterized in that, The phrase "after obtaining cross-device triangulation results, continue to apply constraint fusion and temporal filtering to form a joint optimization in the spatial and temporal dimensions to compensate for depth accuracy" refers to: At each time k, the system first extracts the two-dimensional coordinates of human key points from the multi-view images. Multi-ray triangulation was performed in conjunction with camera geometry to obtain the initial triangulation results for each key point of interest (POI). and its covariance ,Should It takes several 3D rays from multiple camera views, weighted by their respective detection confidence, intersection geometry quality, and reprojection consistency, and obtains the optimal closed-form estimate of the 3D coordinates of each point of interest (POI) in the current frame using weighted least squares (WLS). The corresponding covariance is then calculated. The corresponding three-dimensional uncertainty; simultaneously, the system reads the state of the previous frame after Kalman filtering. Generate the position prior of the current frame, which is used as the initialization of constraint optimization and the time-stability term; Subsequently, the system enters the constraint optimization fusion phase: when the player is detected to be standing with both feet on the ground, ground constraints are applied to the ankle POI, and weighted fusion is performed according to the inverse ratio of the triangulation variance to the ground constraint variance; constant bone length constraints are applied to all skeletal connection pairs to suppress non-physical jumps in the distance between adjacent joints; and constraints based on the elbow, knee, shoulder, and hip joints are applied. Joint angle constraints on the rotating manifold geodesic distance ensure that the posture falls within the physiologically attainable range of the human body; The system combines the above constraints with triangulation observations and prior terms from the previous frame to form a unified objective function, and solves it to obtain the constrained optimized 3D keypoint results for the current frame. ;Will As the observation input, and As the observation noise covariance, Kalman filtering is performed on each keypoint (POI) to achieve adaptive temporal smoothing based on measurement confidence, ultimately outputting stable 3D coordinates. and corresponding velocity estimation .
10. The cross-device collaborative positioning method based on FISHER information fusion according to claim 7, characterized in that, To ensure the system can still provide reasonable output under degraded conditions, and automatically revert to high-precision mode after conditions recover, three modes are defined: normal mode, degraded mode, and recovery mode, along with their switching methods: When the intersection angle And the distance between the two closest points When the reprojection residual is qualified, the system operates in normal mode, running the complete triangulation, constraint fusion, and Kalman filtering pipeline, outputting high-precision results, including α_min, The preset robust gating threshold; When only a single camera is visible, intersection angle When pose drift is caused by anchor point loss, the system automatically switches to degraded mode: the system performs (1) the prediction step of the Kalman filter to provide short-term position continuation; (2) the skeleton and ground constraints propagate the information of visible joints to invisible joints; and (3) the parameterized human body model is used to directly recover the 3D pose from monocular 2D detection as a weak observation. When the second view becomes visible again or the anchor point is detected again, the system enters recovery mode. The system performs anchor point PnP recalibration and Kalman filter relinearization to eliminate the drift accumulated during the degraded mode. After the reprojection residual converges and the conditions of normal mode are met, the system returns to normal mode.
11. A cross-device collaborative positioning system based on FISHER information fusion, comprising: establishing a world coordinate system; deploying at least two physical anchor points with known coordinates in the field; including at least two players wearing XR glasses; and a server that interacts with all XR glasses; characterized in that: The XR glasses output their own pose through 6DoF self-localization and capture images of the opposing player. They perform 2D human keypoint detection on each frame of the image and output the pixel coordinates of each keypoint in the image. They upload their own pose and the pixel coordinates of each keypoint of the opposing player in the image to the server in real time, and also receive the world position data of relevant keypoints sent by the server. The server includes a memory and a processor. The memory is connected to the processor and is used to store the pose of the computer program and the XR glasses in the world coordinate system and the pixel coordinates of each key point of the opposite player, which are uploaded in real time. The server executes the cross-device collaborative positioning method based on FISHER information fusion as described in any one of claims 1 to 10 to obtain the world position of each key point and distribute the world position data of the relevant key points to the XR glasses of all players.
12. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the cross-device collaborative positioning method based on FISHER information fusion as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Binocular camera calibration method and system, electronic equipment and storage medium
CN113808220A
Gesture track interaction method, intelligent glasses and storage medium
CN116820251A