Low-latency optical motion capture and markerless real-time co-located positioning method and system for multi-person XR large spaces
By combining array-deployed optical sensors and edge computing nodes, and utilizing convolutional neural networks to achieve label-free skeletal keypoint extraction and multi-view triangulation, the latency and accuracy issues of optical motion capture and collaborative localization in large-space XR for multiple users are resolved, improving the accuracy and consistency of interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU OGESTE INFORMATION TECH CO LTD
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-31
AI Technical Summary
In existing large-space extended reality (XR) applications involving multiple users, optical motion capture systems require users to wear markers, which affects user experience and is prone to data loss; markerless motion capture methods are difficult to extract features in multi-user scenarios and have poor real-time performance; collaborative localization methods have high network latency, making it difficult to meet the requirements of low-latency interaction, and multi-sensor synchronization and data fusion are highly complex.
The system employs an array of optical sensors to synchronously capture images, utilizes edge computing nodes for real-time preprocessing, combines convolutional neural networks to extract unlabeled skeletal key points, and constructs a three-dimensional human skeletal model through multi-view triangulation to achieve low-latency multi-person motion capture and real-time collaborative localization.
It improves the accuracy of 3D positioning, eliminates positional conflicts in multi-person interactions, ensures spatial consistency, and enhances the overall accuracy of interaction.
Smart Images

Figure CN122492998A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative positioning technology, and in particular to a low-latency optical motion capture and markerless real-time collaborative positioning method and system for multi-person XR large spaces. Background Technology
[0002] Currently, optical motion capture and real-time collaborative positioning are key technologies in large-space applications of extended reality (XR) for multiple users. Existing optical motion capture systems typically require users to wear markers, which not only affects the user experience but also easily leads to data loss due to occlusion. Markerless motion capture methods use computer vision, but in multi-person scenarios, feature extraction is difficult due to mutual occlusion and complex backgrounds, resulting in poor real-time performance. Moreover, existing collaborative localization methods mostly adopt a centralized processing architecture, requiring all sensor data to be transmitted to a central server, resulting in high network latency and difficulty in meeting the requirements for low-latency interaction. In addition, the high complexity of multi-sensor synchronization and data fusion in large spaces means that existing methods often sacrifice accuracy for speed, thereby greatly reducing the localization effect. Therefore, in order to overcome the above-mentioned defects, the present invention provides a low-latency optical motion capture and markerless real-time collaborative positioning method and system for multi-person XR large spaces. Summary of the Invention
[0003] This invention provides a low-latency optical motion capture and label-free real-time collaborative localization method and system for multi-person XR large spaces. It uses an array of deployed optical sensors to synchronously capture images of multiple people, and utilizes edge computing nodes for real-time preprocessing to reduce data transmission latency and ensure real-time processing. It combines convolutional neural networks to extract label-free skeletal key points, and then constructs a three-dimensional human skeletal model through multi-view triangulation. This achieves low-latency multi-person motion capture and real-time collaborative localization, improves three-dimensional localization accuracy, eliminates positional conflicts in multi-person interaction, ensures spatial consistency, and improves the overall accuracy of interaction.
[0004] This invention provides a low-latency optical motion capture and markerless real-time cooperative localization method for large-scale multi-user XR applications, comprising: Step 1: Based on the optical sensors deployed in an array in a large space, multi-person images are captured synchronously in real time, and the obtained multi-person images are time-stamp aligned and image enhancement processed based on each edge computing node to obtain pre-processed multi-person images; Step 2: Extract features from the preprocessed multi-person images using a convolutional neural network to obtain unlabeled 2D information of skeletal key points; Step 3: Perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and construct the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. Step 4: Determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and perform collaborative localization correction on the three-dimensional position and posture of different human bodies in a large space to obtain target localization data; Step 5: Transmit the target location data to the XR rendering engine in real time for rendering.
[0005] Preferably, a low-latency optical motion capture and label-free real-time collaborative localization method for large-space XR involving multiple users includes, in step 1, synchronously and in real-time capturing images of multiple users based on an array of optical sensors deployed in the large space, and performing timestamp alignment and image enhancement processing on the obtained images based on each edge computing node to obtain preprocessed images of multiple users, including: Establish communication links between optical sensors deployed in arrays in a large space and edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. The optical sensors deployed in the array are synchronously controlled based on periodic trigger signals, and raw multi-person images containing multiple human targets are captured based on the synchronous control results. Based on the trigger signal of each cycle, a cycle number corresponding to the corresponding cycle is generated. At the same time, the timestamp when the optical sensor is started is obtained, and the cycle number and timestamp are added to the original multi-person image. Based on the added results, the original multi-person image is sent to the edge computing node according to the communication link, and the edge computing node performs timestamp alignment and image enhancement processing on the original multi-person image to obtain a preprocessed multi-person image.
[0006] Preferably, a low-latency optical motion capture and label-free real-time collaborative localization method for large-space multi-person XR involves performing timestamp alignment and image enhancement processing on the original multi-person image based on edge computing nodes to obtain a preprocessed multi-person image, including: The edge computing node receives raw multi-person images uploaded by optical sensors and groups the raw multi-person images uploaded by each optical sensor into the same batch based on the period number carried in the raw multi-person images. Extract the timestamps of each original multi-person image in the same batch group. At the same time, access the historical parameters of the global synchronization clock source, extract the time base corresponding to the cycle number, and perform time offset compensation on the timestamps of each original multi-person image based on the time base. Based on the time offset compensation results, the original multi-person image is timestamped and then the human target area and background area in the original multi-person image are separated based on the timestamping alignment results. Based on the separation results, the pixel grayscale of the human target area and the background area are determined respectively. When the pixel grayscale does not meet the preset standard, the original multi-person image is enhanced based on the preset contrast curve. Preprocessed multi-person images are obtained based on timestamp alignment results and image enhancement processing results.
[0007] Preferably, in a low-latency optical motion capture and label-free real-time collaborative localization method for large-space multi-person XR, step 2 involves extracting features from preprocessed multi-person images using a convolutional neural network to obtain label-free skeletal keypoint two-dimensional information, including: The preprocessed multi-person images are obtained, and multi-level downsampling convolution processing is performed on the preprocessed multi-person images based on convolutional neural networks to obtain multi-level image feature maps with different spatial resolutions. Based on multi-level image feature maps, deep semantic features and shallow detail features of preprocessed multi-person images are extracted, and the deep semantic features and shallow detail features are fused to obtain an enhanced multi-scale feature map. Convolutional neural networks are used to perform convolution and deconvolution processing on the enhanced multi-scale feature map to determine the existence probability distribution of each human instance in the enhanced multi-scale feature map for each preset skeletal key point type in the two-dimensional space of the image, and a two-dimensional probability heat map corresponding to each preset skeletal key point type is obtained based on the existence probability distribution. Meanwhile, based on the results of convolution and deconvolution, a two-dimensional vector field corresponding to the preprocessed multi-person image is extracted, and the direction information of any pixel in the image space of the preprocessed multi-person image pointing to the adjacent key points inside the human body instance is determined based on the two-dimensional vector field. Based on the two-dimensional probability heat map, the candidate key point positions corresponding to each preset skeletal key point type are determined, and the candidate key points corresponding to the candidate key point positions in the preprocessed multi-person images are clustered based on the direction information. Based on the clustering results, the candidate key point set corresponding to different independent human instances is obtained. Two-dimensional coordinate data of each candidate keypoint in the candidate keypoint set corresponding to different independent human instances are extracted to obtain unlabeled skeletal keypoint two-dimensional information.
[0008] Preferably, a low-latency optical motion capture and label-free real-time collaborative localization method for large-scale XR involving multiple users includes, in step 3, multi-view triangulation of the 2D skeletal keypoint information corresponding to each edge computing node to obtain the 3D skeletal keypoints corresponding to each human body, and constructing a human skeletal model corresponding to each human body based on the 3D skeletal keypoints, including: Obtain the 2D information of the skeletal key points corresponding to each edge computing node, and extract the timestamp, type code, detection confidence and 2D coordinates corresponding to the 2D information of the skeletal key points; Based on the timestamp and the viewpoint corresponding to the optical sensor, the 2D information of the skeletal key points of each edge computing node is time-aligned and grouped, and multi-view 2D data is obtained based on the grouping results. Based on multi-view two-dimensional data, cross-view human instance identity matching is performed on different human instances, and based on the human instance identity matching results, two-dimensional data belonging to the same independent human instance under different views are associated with the same human identity identifier to obtain a multi-view data group for each independent human instance. Based on the intrinsic parameter matrix of the optical sensor under different viewpoints and the extrinsic parameter matrix relative to the global world coordinate system, the projection matrix that projects the three-dimensional world coordinates to the two-dimensional image coordinates under different viewpoints is determined. Based on the type coding corresponding to the two-dimensional information of the skeletal key points, the two-dimensional coordinates and detection confidence of each skeletal key point under each type coding in the multi-view data group are extracted, and the observation angle between the view and each skeletal key point is determined based on the two-dimensional coordinates. Based on the observation angle quantification index and the detection confidence of each skeletal key point under each viewpoint, the target weight of each skeletal key point under different viewpoints is determined, and the three-dimensional spatial coordinates of each skeletal key point are determined based on the projection matrix under each viewpoint and the target weight of each skeletal key point under different viewpoints. Based on the three-dimensional spatial coordinates of each skeletal key point, the corresponding three-dimensional skeletal key points of each human body are obtained, and a human skeletal model corresponding to each human body is constructed based on the three-dimensional skeletal key points.
[0009] Preferably, a low-latency optical motion capture and markerless real-time collaborative localization method for multi-person XR large spaces obtains the three-dimensional skeletal key points corresponding to each human body based on the three-dimensional spatial coordinates of each skeletal key point, and constructs a human skeletal model corresponding to each human body based on the three-dimensional skeletal key points, including: Extract the three-dimensional spatial coordinates of each three-dimensional bone key point, and determine the instantaneous three-dimensional length of the bone segment formed by adjacent three-dimensional bone key points based on the three-dimensional spatial coordinates. Meanwhile, based on prior knowledge of the inherent length ratio of human skeletons, the range of values for the length of human skeleton segments is determined, and the instantaneous three-dimensional length of the skeleton segment is compared with the range of values. If the instantaneous three-dimensional length of a bone segment is outside the range of values, then it is determined that at least one of the two three-dimensional bone keypoints corresponding to the current bone segment is an abnormal bone keypoint. At the same time, the instantaneous lengths of the associated bone segments formed by two 3D skeletal keypoints and other adjacent 3D skeletal keypoints are determined respectively, and the instantaneous lengths of the associated bone segments are compared with the corresponding standard value ranges. The preprocessed multi-person images in a continuous frame sequence are traversed, and the position trajectory smoothness of two 3D skeletal key points is determined based on the traversal results. The reliability index of each 3D skeletal key point is determined based on the position trajectory smoothness. The key points of the skeleton to be corrected are determined based on the reliability index and the comparison results of the instantaneous length of the associated skeletal segment with the corresponding standard value range. Based on the bone type of the current bone segment, the corresponding historical length average and stable position information of adjacent 3D bone key points are obtained from the database, and the position of the bone key points to be corrected is corrected based on the historical length average and stable position information of adjacent 3D bone key points. The initial human skeleton model is obtained from the model library and then initialized. Based on the position correction results and initialization results, the instantaneous three-dimensional lengths of each human bone segment are mapped to the same position in the initial human bone model to obtain the initial lengths of all bone segments in the initial human bone model. The obtained initial length is proportionally aligned with a predefined standard skeleton model, and the initial scale parameters of the initial human skeleton model are obtained based on the proportional alignment result. The human skeleton model corresponding to each human body is obtained based on the initial scale parameters.
[0010] Preferably, a low-latency optical motion capture and label-free real-time collaborative localization method for multi-person XR large spaces includes, in step 4, determining the 3D position and pose of different human bodies based on a human skeletal model, and performing collaborative localization correction on the 3D position and pose of different human bodies in the large space to obtain target localization data, including: Obtain human skeleton models corresponding to different human bodies, parse the human skeleton models, extract the three-dimensional spatial coordinates of the root node of the human skeleton model, and use the three-dimensional spatial coordinates of the root node as the three-dimensional position of the current human body. At the same time, the rotation parameters of different joints of the current human body are determined based on the human skeleton model, and the rotation parameters of different joints of the current human body are used as the current posture of the human body. The three-dimensional positions and postures of different human bodies at the same time stamp are summarized to obtain a set of original pose states of multiple people, and the three-dimensional positions and postures of different human bodies in the set of original pose states of multiple people are transformed into a predefined unified world coordinate system. Based on a pre-set human kinematics model, the three-dimensional position and posture of different human bodies in a unified world coordinate system are analyzed to predict the target state of each human body at the next moment. The spatial relative relationship between the original pose data of different human bodies at the same moment is used as a constraint to correct the target state. Based on the correction results, the final three-dimensional position and attitude parameter set of each human body is obtained, and the final three-dimensional position and attitude parameter set of each human body is used as the target positioning data of each human body.
[0011] Preferably, a low-latency optical motion capture and markerless real-time collaborative localization method for large-scale multi-user XR environments includes step 5, in which the target localization data is transmitted in real-time to the XR rendering engine for rendering processing, including: Based on the target positioning data, the identification, three-dimensional position and posture parameters of each human body are extracted, and the extracted identification, three-dimensional position and posture parameters are encapsulated into a data packet to be transmitted; The data packets to be transmitted are sent to the data receiving port corresponding to the XR rendering engine based on the low-latency network interface, and the data packets to be transmitted are inserted into the rendering queue after the XR rendering engine receives the data packets to be transmitted. The insertion result-driven XR rendering engine maps the 3D position and pose parameters according to the identity identifier in the data packet to be transmitted and drives the corresponding virtual avatar skeleton model in the scene. Based on the driving results and combined with the virtual scene content, the pre-processed multi-person images are rendered and output using the XR rendering engine.
[0012] This invention provides a low-latency optical motion capture and markerless real-time cooperative localization system for large-scale multi-person XR applications, comprising: The image acquisition and preprocessing module is used to synchronously and in real time capture multi-person images based on optical sensors deployed in an array in a large space, and to perform timestamp alignment and image enhancement processing on the obtained multi-person images based on each edge computing node to obtain preprocessed multi-person images. The feature extraction module is used to extract features from preprocessed multi-person images based on convolutional neural networks to obtain unlabeled two-dimensional information of skeletal key points; The model building module is used to perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and to build the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. The positioning and correction module is used to determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and to perform collaborative positioning and correction of the three-dimensional position and posture of different human bodies in a large space to obtain target positioning data. The rendering module is used to transmit target positioning data to the XR rendering engine in real time for rendering processing.
[0013] Preferably, a low-latency optical motion capture and markerless real-time collaborative localization system for large-space XR applications involving multiple users includes an image acquisition and preprocessing module comprising: Image acquisition unit, used for: Establish communication links between optical sensors deployed in arrays in a large space and edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. The optical sensors deployed in the array are synchronously controlled based on periodic trigger signals, and raw multi-person images containing multiple human targets are captured based on the synchronous control results. Image preprocessing unit, used for: Based on the trigger signal of each cycle, a cycle number corresponding to the corresponding cycle is generated. At the same time, the timestamp when the optical sensor is started is obtained, and the cycle number and timestamp are added to the original multi-person image. Based on the added results, the original multi-person image is sent to the edge computing node according to the communication link, and the edge computing node performs timestamp alignment and image enhancement processing on the original multi-person image to obtain a preprocessed multi-person image.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: Multi-person images are captured synchronously by optical sensors deployed in an array, and edge computing nodes are used for real-time preprocessing to reduce data transmission latency and ensure real-time processing. Combined with convolutional neural networks to extract unlabeled skeletal key points, and then multi-view triangulation to construct a three-dimensional human skeleton model, low-latency multi-person motion capture and real-time collaborative positioning are achieved. This improves the accuracy of three-dimensional positioning, eliminates positional conflicts in multi-person interaction, ensures spatial consistency, and improves the overall accuracy of interaction.
[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.
[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a low-latency optical motion capture and markerless real-time collaborative localization method for multi-person XR large spaces, as described in an embodiment of the present invention. Figure 2 This is a flowchart of step 1 in a low-latency optical motion capture and markerless real-time collaborative localization method for multi-person XR large spaces according to an embodiment of the present invention; Figure 3This is a structural diagram of a low-latency optical motion capture and markerless real-time collaborative positioning system for multi-person XR large spaces, as described in an embodiment of the present invention. Detailed Implementation
[0018] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0019] Example 1: This example provides a low-latency optical motion capture and markerless real-time cooperative localization method for large-scale XR environments with multiple users, such as... Figure 1 As shown, it includes: Step 1: Based on the optical sensors deployed in an array in a large space, multi-person images are captured synchronously in real time, and the obtained multi-person images are time-stamp aligned and image enhancement processed based on each edge computing node to obtain pre-processed multi-person images; Step 2: Extract features from the preprocessed multi-person images using a convolutional neural network to obtain unlabeled 2D information of skeletal key points; Step 3: Perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and construct the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. Step 4: Determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and perform collaborative localization correction on the three-dimensional position and posture of different human bodies in a large space to obtain target localization data; Step 5: Transmit the target location data to the XR rendering engine in real time for rendering.
[0020] In this embodiment, low-latency optical motion capture refers to using optical sensors to capture motion and optimizing the processing flow to minimize time delay.
[0021] In this embodiment, the markerless real-time collaborative positioning method refers to a method that calculates and coordinates the positions and postures of multiple people in three-dimensional space in real time without the need to attach markers to the human body.
[0022] In this embodiment, preprocessed multi-person images refer to multi-person image data that has been time-stamped and enhanced, and is used for subsequent feature extraction.
[0023] In this embodiment, the two-dimensional information of skeletal key points refers to the position information of human joints or key points in a two-dimensional image coordinate system extracted from the preprocessed image through a convolutional neural network.
[0024] In this embodiment, the three-dimensional skeleton key points refer to the coordinate points in three-dimensional space obtained by converting the two-dimensional skeleton key point information through multi-view triangulation.
[0025] In this embodiment, the human skeleton model refers to a three-dimensional model representing the human skeleton structure and posture, constructed based on three-dimensional skeletal key points.
[0026] In this embodiment, collaborative positioning correction refers to coordinating and correcting the three-dimensional positions and postures of multiple people to eliminate errors and ensure consistency and accuracy in space.
[0027] In this embodiment, the target positioning data refers to the final three-dimensional position and attitude data after collaborative positioning correction, which is used to transmit to the XR rendering engine for rendering processing.
[0028] In this embodiment, the XR rendering engine refers to a software engine used to render extended reality content and process positioning data to generate virtual scenes.
[0029] The beneficial effects of the above technical solution are as follows: multi-person images are captured synchronously by optical sensors deployed in an array, and edge computing nodes are used for real-time preprocessing to reduce data transmission latency and ensure real-time processing. Combined with convolutional neural networks to extract unlabeled skeletal key points, and then multi-view triangulation to construct a three-dimensional human skeleton model, low-latency multi-person motion capture and real-time collaborative positioning are achieved, improving the accuracy of three-dimensional positioning, eliminating positional conflicts in multi-person interaction, ensuring spatial consistency, and improving the overall accuracy of interaction.
[0030] Example 2: Based on Example 1, this example provides a low-latency optical motion capture and markerless real-time cooperative localization method for large-space XR environments with multiple users, such as... Figure 2 As shown, in step 1, multiple-person images are synchronously and in real-time captured by optical sensors deployed in an array in a large space. The obtained multiple-person images are then time-stamp aligned and image enhancement processed based on each edge computing node to obtain pre-processed multiple-person images, including: Step 101: Construct communication links between the optical sensors deployed in the array in a large space and the edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. Step 102: Synchronize the optical sensors deployed in the array based on the periodic trigger signal, and capture the original multi-person image containing multiple human targets based on the synchronization control result; Step 103: Generate the corresponding cycle number based on the trigger signal of each cycle. At the same time, obtain the timestamp when the optical sensor is started, and add the cycle number and timestamp to the original multi-person image. Step 104: Based on the addition results, the original multi-person image is sent to the edge computing node according to the communication link, and the original multi-person image is timestamped and enhanced based on the edge computing node to obtain the preprocessed multi-person image.
[0031] In this embodiment, the global synchronization clock source refers to a time reference device that generates a high-precision, stable periodic signal to send a unified trigger command to all optical sensors, ensuring that all image acquisitions start at the exact same time.
[0032] In this embodiment, the periodic trigger signal refers to an electronic pulse signal with a fixed time interval generated by a global synchronization clock source, which is used to simultaneously trigger all optical sensors to start exposure in order to capture an image.
[0033] In this embodiment, the original multi-person image refers to the initial image data frame that is directly acquired and generated by the optical sensor after receiving the periodic trigger signal, without any processing, which contains multiple human targets in a large space scene.
[0034] The beneficial effects of the above technical solution are as follows: by constructing a unified synchronous acquisition and distributed processing architecture, high-precision temporal alignment and high-quality preprocessing of multi-person scene image data are achieved, effectively eliminating the acquisition time deviation between multiple sensors and ensuring the temporal consistency of subsequent processed data. At the same time, the image quality is improved through edge-side enhancement processing, providing a more reliable input for label-free feature extraction, and ensuring the foundation for low latency and high precision of the entire system from the data source.
[0035] Example 3: Based on Example 2, this example provides a low-latency optical motion capture and markerless real-time collaborative localization method for large-space multi-person XR. It performs timestamp alignment and image enhancement processing on the original multi-person image based on edge computing nodes to obtain a preprocessed multi-person image, including: The edge computing node receives raw multi-person images uploaded by optical sensors and groups the raw multi-person images uploaded by each optical sensor into the same batch based on the period number carried in the raw multi-person images. Extract the timestamps of each original multi-person image in the same batch group. At the same time, access the historical parameters of the global synchronization clock source, extract the time base corresponding to the cycle number, and perform time offset compensation on the timestamps of each original multi-person image based on the time base. Based on the time offset compensation results, the original multi-person image is timestamped and then the human target area and background area in the original multi-person image are separated based on the timestamping alignment results. Based on the separation results, the pixel grayscale of the human target area and the background area are determined respectively. When the pixel grayscale does not meet the preset standard, the original multi-person image is enhanced based on the preset contrast curve. Preprocessed multi-person images are obtained based on timestamp alignment results and image enhancement processing results.
[0036] In this embodiment, grouping in the same batch refers to the operation of grouping all images that belong to the same trigger period and are acquired by different optical sensors into the same set for processing, based on the period number carried in the original multi-person images.
[0037] In this embodiment, time offset compensation refers to the process of using a precise time reference provided by a global synchronization clock source to correct minor system errors or transmission delays in the original image timestamps of each optical sensor, so that they are unified under the same absolute time reference system.
[0038] In this embodiment, the human target region and the background region refer to the connected regions containing human pixels that are distinguished by image segmentation technology in the original multi-person image, and the remaining scene parts that do not contain human pixels.
[0039] In this embodiment, the preset contrast curve refers to a mapping function predefined based on prior knowledge or scene statistical characteristics, used to adjust the distribution of pixel grayscale values in the image to enhance the visual distinction between the target and the background.
[0040] The beneficial effects of the above technical solution are as follows: by performing precise time synchronization calibration and region adaptive quality enhancement on multi-source images, strict temporal consistency of data acquisition under different perspectives is ensured, effectively overcoming alignment errors caused by transmission or processing delays. At the same time, the differentiated enhancement of human body regions and backgrounds improves the recognition and stability of target features, providing a higher quality and more reliable input image foundation for subsequent label-free keypoint detection.
[0041] Example 4: Based on Example 1, this example provides a low-latency optical motion capture and label-free real-time collaborative localization method for large-space multi-person XR. In step 2, feature extraction is performed on the preprocessed multi-person image based on a convolutional neural network to obtain label-free skeletal key point two-dimensional information, including: The preprocessed multi-person images are obtained, and multi-level downsampling convolution processing is performed on the preprocessed multi-person images based on convolutional neural networks to obtain multi-level image feature maps with different spatial resolutions. Based on multi-level image feature maps, deep semantic features and shallow detail features of preprocessed multi-person images are extracted, and the deep semantic features and shallow detail features are fused to obtain an enhanced multi-scale feature map. Convolutional neural networks are used to perform convolution and deconvolution processing on the enhanced multi-scale feature map to determine the existence probability distribution of each human instance in the enhanced multi-scale feature map for each preset skeletal key point type in the two-dimensional space of the image, and a two-dimensional probability heat map corresponding to each preset skeletal key point type is obtained based on the existence probability distribution. Meanwhile, based on the results of convolution and deconvolution, a two-dimensional vector field corresponding to the preprocessed multi-person image is extracted, and the direction information of any pixel in the image space of the preprocessed multi-person image pointing to the adjacent key points inside the human body instance is determined based on the two-dimensional vector field. Based on the two-dimensional probability heat map, the candidate key point positions corresponding to each preset skeletal key point type are determined, and the candidate key points corresponding to the candidate key point positions in the preprocessed multi-person images are clustered based on the direction information. Based on the clustering results, the candidate key point set corresponding to different independent human instances is obtained. Two-dimensional coordinate data of each candidate keypoint in the candidate keypoint set corresponding to different independent human instances are extracted to obtain unlabeled skeletal keypoint two-dimensional information.
[0042] In this embodiment, the local peak position on each heatmap represents the pixel coordinates of the corresponding type of key point.
[0043] In this embodiment, the multi-level image feature map refers to the set of feature tensors with different spatial resolutions (i.e., different lengths and widths) and different semantic abstractions output by different network layers after the input image is downsampled multiple times by a convolutional neural network.
[0044] In this embodiment, the enhanced multi-scale feature map refers to a new feature map that combines rich semantic information and fine spatial details, obtained by fusing deep semantic features and shallow detail features in the multi-level image feature map through a specific network structure (such as a feature pyramid).
[0045] In this embodiment, the two-dimensional probability heatmap refers to a probability value at each location on the two-dimensional grid of the image, which is predicted and output by the convolutional neural network. This probability value represents the likelihood that a corresponding type of skeletal key point (such as the left shoulder or right knee) exists at that location, and is usually presented as a "hotspot" in the form of a Gaussian distribution on the graph.
[0046] In this embodiment, the two-dimensional vector field refers to a two-dimensional vector at each position on the two-dimensional grid of the image, which is predicted and output by the convolutional neural network. This vector encodes directional information from the pixel to an adjacent key point in the human body instance (usually another key point on the same bone), and is used to establish the connection relationship between key points.
[0047] In this embodiment, the candidate key point location refers to a series of pixel coordinate points determined by finding the local probability peak in the two-dimensional probability heatmap corresponding to each type (such as through non-maximum suppression). Each point represents the location of a possible key point of a specific type.
[0048] In this embodiment, the candidate key point set refers to the group of all key points that are assigned to the same independent human body instance after clustering and associating the discrete candidate key point positions belonging to different types by utilizing the directional constraint information provided by the two-dimensional vector field.
[0049] The beneficial effects of the above technical solution are as follows: through multi-scale feature fusion and bi-branch prediction mechanism, it achieves highly robust detection and accurate association of unlabeled skeletal key points of multiple people in complex scenes. It can simultaneously output the position probability distribution of key points and their topological connection relationship, effectively addressing challenges such as human body occlusion, scale changes and pose diversity, ensuring the integrity of key point detection. Through data-driven instance clustering of key points, it significantly improves the accuracy and automation of pose estimation in multi-person scenes, providing reliable two-dimensional input for subsequent 3D reconstruction.
[0050] Example 5: Based on Example 1, this example provides a low-latency optical motion capture and markerless real-time collaborative localization method for large-scale XR with multiple users. In step 3, the two-dimensional information of the skeletal key points corresponding to each edge computing node is triangulated from multiple perspectives to obtain the three-dimensional skeletal key points corresponding to each human body. Based on the three-dimensional skeletal key points, a human skeleton model corresponding to each human body is constructed, including: Obtain the 2D information of the skeletal key points corresponding to each edge computing node, and extract the timestamp, type code, detection confidence and 2D coordinates corresponding to the 2D information of the skeletal key points; Based on the timestamp and the viewpoint corresponding to the optical sensor, the 2D information of the skeletal key points of each edge computing node is time-aligned and grouped, and multi-view 2D data is obtained based on the grouping results. Based on multi-view two-dimensional data, cross-view human instance identity matching is performed on different human instances, and based on the human instance identity matching results, two-dimensional data belonging to the same independent human instance under different views are associated with the same human identity identifier to obtain a multi-view data group for each independent human instance. Based on the intrinsic parameter matrix of the optical sensor under different viewpoints and the extrinsic parameter matrix relative to the global world coordinate system, the projection matrix that projects the three-dimensional world coordinates to the two-dimensional image coordinates under different viewpoints is determined. Based on the type coding corresponding to the two-dimensional information of the skeletal key points, the two-dimensional coordinates and detection confidence of each skeletal key point under each type coding in the multi-view data group are extracted, and the observation angle between the view and each skeletal key point is determined based on the two-dimensional coordinates. Based on the observation angle quantification index and the detection confidence of each skeletal key point under each viewpoint, the target weight of each skeletal key point under different viewpoints is determined, and the three-dimensional spatial coordinates of each skeletal key point are determined based on the projection matrix under each viewpoint and the target weight of each skeletal key point under different viewpoints. Based on the three-dimensional spatial coordinates of each skeletal key point, the corresponding three-dimensional skeletal key points of each human body are obtained, and a human skeletal model corresponding to each human body is constructed based on the three-dimensional skeletal key points.
[0051] In this embodiment, cross-view human instance identity matching is performed on different human instances based on multi-view two-dimensional data, including: First, initial matching is performed based on the consistency between the spatial overlap of the human body bounding boxes and the topological structure of key points in each viewpoint. Subsequently, by utilizing the motion continuity of the two-dimensional trajectories of human body key points in multiple perspectives within a short time window, a graph model based on motion feature similarity is constructed. By solving the graph matching problem, two-dimensional key point data belonging to the same physical human body in different perspectives are associated with the same human body identity, forming a multi-view data group.
[0052] In this embodiment, multi-view two-dimensional data refers to the collection of two-dimensional information of all human skeletal key points acquired from different optical sensor perspectives at the same time after time-stamp alignment and grouping processing.
[0053] In this embodiment, human instance identity matching refers to the process of using features such as the spatial relationship of the human bounding box, the consistency of the key point topology, and the continuity of temporal motion to determine whether the two-dimensional human key point data detected from different optical sensor perspectives belong to the same physical human body.
[0054] In this embodiment, the projection matrix refers to a mathematical transformation matrix composed of the intrinsic parameter matrix (describing its internal imaging geometry properties) and the extrinsic parameter matrix (describing its position and orientation in the global world coordinate system) of the optical sensor, used to project three-dimensional world coordinate points onto the two-dimensional pixel coordinates of the sensor's imaging plane.
[0055] In this embodiment, the observation angle quantification index refers to a parameter used to numerically characterize the angle between the observation line direction of an optical sensor and the surface normal direction at the estimated position of a specific three-dimensional key point to be triangulated. This parameter reflects the observation geometric quality of the key point from this perspective.
[0056] In this embodiment, the target weight refers to the influence coefficient assigned to a specific key point for each optical sensor viewpoint in the weighted triangulation calculation. This coefficient is determined by the detection confidence of the key point under that viewpoint and the aforementioned observation angle quantification index, and is used to balance the reliability of observation values from different viewpoints.
[0057] The beneficial effects of the above technical solution are as follows: By performing precise time alignment and identity matching on multi-view 2D keypoint information, the problem of cross-view human instance confusion in multi-person scenes is effectively solved, ensuring the correctness of data association. By introducing adaptive weight calculation of observation angle and detection confidence, and integrating multi-view geometric constraints in the triangulation process, the accuracy and robustness of 3D keypoint reconstruction are significantly improved, especially mitigating the impact of partial viewpoint occlusion or detection noise. Finally, a personalized skeleton model is constructed based on the optimized 3D keypoints, providing an accurate and structured 3D human representation for subsequent pose estimation and collaborative localization, laying the data foundation for high-precision motion capture.
[0058] Example 6: Building upon Example 5, this example provides a low-latency optical motion capture and markerless real-time collaborative localization method for large-scale XR involving multiple users. It obtains the 3D skeletal keypoints corresponding to each human body based on the 3D spatial coordinates of each skeletal keypoint, and constructs a human skeletal model for each human body based on these 3D skeletal keypoints, including: Extract the three-dimensional spatial coordinates of each three-dimensional bone key point, and determine the instantaneous three-dimensional length of the bone segment formed by adjacent three-dimensional bone key points based on the three-dimensional spatial coordinates. Meanwhile, based on prior knowledge of the inherent length ratio of human skeletons, the range of values for the length of human skeleton segments is determined, and the instantaneous three-dimensional length of the skeleton segment is compared with the range of values. If the instantaneous three-dimensional length of a bone segment is outside the range of values, then it is determined that at least one of the two three-dimensional bone keypoints corresponding to the current bone segment is an abnormal bone keypoint. At the same time, the instantaneous lengths of the associated bone segments formed by two 3D skeletal keypoints and other adjacent 3D skeletal keypoints are determined respectively, and the instantaneous lengths of the associated bone segments are compared with the corresponding standard value ranges. The preprocessed multi-person images in a continuous frame sequence are traversed, and the position trajectory smoothness of two 3D skeletal key points is determined based on the traversal results. The reliability index of each 3D skeletal key point is determined based on the position trajectory smoothness. The key points of the skeleton to be corrected are determined based on the reliability index and the comparison results of the instantaneous length of the associated skeletal segment with the corresponding standard value range. Based on the bone type of the current bone segment, the corresponding historical length average and stable position information of adjacent 3D bone key points are obtained from the database, and the position of the bone key points to be corrected is corrected based on the historical length average and stable position information of adjacent 3D bone key points. The initial human skeleton model is obtained from the model library and then initialized. Based on the position correction results and initialization results, the instantaneous three-dimensional lengths of each human bone segment are mapped to the same position in the initial human bone model to obtain the initial lengths of all bone segments in the initial human bone model. The obtained initial length is proportionally aligned with a predefined standard skeleton model, and the initial scale parameters of the initial human skeleton model are obtained based on the proportional alignment result. The human skeleton model corresponding to each human body is obtained based on the initial scale parameters.
[0059] In this embodiment, the human skeleton model is defined in a tree-like hierarchical structure, with its root node located in the pelvis and containing the degree of freedom constraints for each joint. The initialization process first calculates the mean three-dimensional coordinates of the key points of the pelvis to determine the location of the root node of the model.
[0060] In this embodiment, the instantaneous three-dimensional length refers to the straight-line distance determined by the spatial coordinates of two adjacent three-dimensional skeletal key points, calculated based on the current moment.
[0061] In this embodiment, the historical length mean refers to the arithmetic mean calculated based on the instantaneous three-dimensional length of the specific bone segment (such as the left upper arm) over multiple consecutive time frames in the past.
[0062] In this embodiment, the initial human skeleton model refers to a predefined parametric 3D skeleton template that has a standard human topology and joint hierarchy, but has not yet been scale-aligned with specific user data.
[0063] In this embodiment, scale alignment refers to the process of matching the lengths of all bone segments in the initial human skeleton model with the corresponding bone segment lengths calculated based on the user's 3D key points in an overall scale by calculating a uniform scaling factor.
[0064] In this embodiment, the initial scale parameter refers to the scaling factor calculated during the scaling process, which is used to scale the initial human skeleton model as a whole to match the actual human body size.
[0065] The beneficial effects of the above technical solution are as follows: by combining prior data on the length ratio of human skeleton with historical data, abnormal points appearing in the reconstruction of 3D key points are effectively identified and corrected, significantly improving the physiological rationality and spatial consistency of the data. At the same time, by dynamically adjusting the scale parameters of the skeleton model, accurate adaptation from a general model to personalized human geometry is achieved, ensuring that the constructed skeleton model can not only conform to the constraints of ordinary human structure, but also accurately reflect the individual morphological differences of specific users, providing a more reliable and stable model foundation for subsequent pose estimation.
[0066] Example 7: Based on Example 1, this example provides a low-latency optical motion capture and label-free real-time cooperative localization method for large-scale XR with multiple users. In step 4, the three-dimensional positions and postures of different human bodies are determined based on the human skeleton model, and the three-dimensional positions and postures of different human bodies in the large space are cooperatively localized and corrected to obtain target localization data, including: Obtain human skeleton models corresponding to different human bodies, parse the human skeleton models, extract the three-dimensional spatial coordinates of the root node of the human skeleton model, and use the three-dimensional spatial coordinates of the root node as the three-dimensional position of the current human body. At the same time, the rotation parameters of different joints of the current human body are determined based on the human skeleton model, and the rotation parameters of different joints of the current human body are used as the current posture of the human body. The three-dimensional positions and postures of different human bodies at the same time stamp are summarized to obtain a set of original pose states of multiple people, and the three-dimensional positions and postures of different human bodies in the set of original pose states of multiple people are transformed into a predefined unified world coordinate system. Based on a pre-set human kinematics model, the three-dimensional position and posture of different human bodies in a unified world coordinate system are analyzed to predict the target state of each human body at the next moment. The spatial relative relationship between the original pose data of different human bodies at the same moment is used as a constraint to correct the target state. Based on the correction results, the final three-dimensional position and attitude parameter set of each human body is obtained, and the final three-dimensional position and attitude parameter set of each human body is used as the target positioning data of each human body.
[0067] In this embodiment, the purpose of using the spatial relative relationship between the original pose data of different human bodies at the same time as a constraint to correct the target state is to eliminate the spatial inconsistency between the pose data of different individuals.
[0068] In this embodiment, the root node refers to the specific joint that is defined as the kinematic center or starting point of the human body in the tree-like hierarchical structure of the parameterized human skeleton model, usually corresponding to the spatial location of the pelvis or lower trunk.
[0069] In this embodiment, the set of original pose states of multiple people refers to the set of three-dimensional position and joint posture parameters of all tracked human individuals in the system at a specific timestamp, without co-correction.
[0070] In this embodiment, the unified world coordinate system refers to a globally unique right-handed three-dimensional Cartesian coordinate system that is predefined and fixed in a large three-dimensional space scene. The absolute position and orientation of all sensors, models and targets are described based on this coordinate system.
[0071] In this embodiment, the preset human kinematic model refers to a mathematical model used to describe the motion transmission relationship between various parts of the human body. It predicts or calculates the motion of various parts of the body based on bone length, joint type (such as ball-and-socket joint, hinge joint) and their degree of freedom constraints.
[0072] In this embodiment, spatial relative relationship refers to the measurable geometric relationship between different human individuals at the same time under a unified world coordinate system, such as the straight-line distance between individuals, the direction angle of the connecting line, or the included angle of their respective orientations.
[0073] The beneficial effects of the above technical solution are as follows: by directly extracting root node coordinates and joint parameters from personalized skeletal models, efficient and structured description of the three-dimensional position and posture of multiple people is achieved. By unifying the poses of all individuals to the same world coordinate system and performing kinematic state prediction, a unified spatiotemporal reference framework is established for multi-target collaborative tracking. Furthermore, by using the inherent spatial relative relationships between individuals as constraints for joint correction, the cumulative error and spatial inconsistency between individuals that may exist in single-target tracking are effectively eliminated. Finally, high-precision, low-latency positioning data with global consistency is output, providing reliable state input for large-space multi-person XR interaction.
[0074] Example 8: Based on Example 1, this example provides a low-latency optical motion capture and markerless real-time collaborative localization method for large-scale multi-user XR. In step 5, the target localization data is transmitted to the XR rendering engine in real time for rendering processing, including: Based on the target positioning data, the identification, three-dimensional position and posture parameters of each human body are extracted, and the extracted identification, three-dimensional position and posture parameters are encapsulated into a data packet to be transmitted; The data packets to be transmitted are sent to the data receiving port corresponding to the XR rendering engine based on the low-latency network interface, and the data packets to be transmitted are inserted into the rendering queue after the XR rendering engine receives the data packets to be transmitted. The insertion result-driven XR rendering engine maps the 3D position and pose parameters according to the identity identifier in the data packet to be transmitted and drives the corresponding virtual avatar skeleton model in the scene. Based on the driving results and combined with the virtual scene content, the pre-processed multi-person images are rendered and output using the XR rendering engine.
[0075] In this embodiment, the data packet to be transmitted refers to a data unit that can be transmitted over a network after the identity identifier, three-dimensional position and attitude parameters of each person in the target positioning data are encapsulated according to a predetermined data structure and serialization format.
[0076] In this embodiment, the low-latency network interface refers to a dedicated physical or virtual network communication channel with high bandwidth and deterministic low latency characteristics, adopted to meet the needs of real-time interaction.
[0077] In this embodiment, the data receiving port refers to a specific software port opened by the XR rendering engine program on the host machine for listening to and receiving the data packets to be transmitted from the network.
[0078] In this embodiment, the rendering queue refers to a buffer maintained internally by the XR rendering engine, which is used to cache the received data packets to be transmitted in timestamp order to ensure that the rendering logic processes the positioning data of each frame in the correct timing.
[0079] In this embodiment, the virtual avatar skeleton model refers to a three-dimensional digital character model in an XR virtual scene that is used to visually represent a real user and can be driven by the skeleton to make corresponding movements.
[0080] In this embodiment, virtual scene content refers to a complete three-dimensional digital space created in the XR rendering engine, which includes the environment, objects, and other virtual elements.
[0081] The beneficial effects of the above technical solution are as follows: by encapsulating and transmitting the optimized and corrected positioning data in a standardized format, the efficiency and integrity of the transmission of human motion state information from the capture end to the rendering end are ensured. By utilizing low-latency networks and queue mechanisms, stable synchronization of data stream and rendering frame rate is achieved, effectively avoiding virtual avatar motion stuttering or jumping caused by data transmission delay or jitter. By directly driving the virtual skeleton model through identity mapping, high-fidelity and low-latency mapping between virtual avatar motion and real human body movements is guaranteed, ultimately providing users with a smooth and immersive multi-person collaborative XR experience.
[0082] Example 9: This example provides a low-latency optical motion capture and markerless real-time cooperative positioning system for large-scale multi-person XR applications, such as... Figure 3 As shown, it includes: The image acquisition and preprocessing module is used to synchronously and in real time capture multi-person images based on optical sensors deployed in an array in a large space, and to perform timestamp alignment and image enhancement processing on the obtained multi-person images based on each edge computing node to obtain preprocessed multi-person images. The feature extraction module is used to extract features from preprocessed multi-person images based on convolutional neural networks to obtain unlabeled two-dimensional information of skeletal key points; The model building module is used to perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and to build the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. The positioning and correction module is used to determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and to perform collaborative positioning and correction of the three-dimensional position and posture of different human bodies in a large space to obtain target positioning data. The rendering module is used to transmit target positioning data to the XR rendering engine in real time for rendering processing.
[0083] The beneficial effects of the above technical solution are as follows: multi-person images are captured synchronously by optical sensors deployed in an array, and edge computing nodes are used for real-time preprocessing to reduce data transmission latency and ensure real-time processing. Combined with convolutional neural networks to extract unlabeled skeletal key points, and then multi-view triangulation to construct a three-dimensional human skeleton model, low-latency multi-person motion capture and real-time collaborative positioning are achieved, improving the accuracy of three-dimensional positioning, eliminating positional conflicts in multi-person interaction, ensuring spatial consistency, and improving the overall accuracy of interaction.
[0084] Example 10: Based on Example 9, this example provides a low-latency optical motion capture and markerless real-time collaborative localization system for large-space XR applications involving multiple users. The image acquisition and preprocessing module includes: Image acquisition unit, used for: Establish communication links between optical sensors deployed in arrays in a large space and edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. The optical sensors deployed in the array are synchronously controlled based on periodic trigger signals, and raw multi-person images containing multiple human targets are captured based on the synchronous control results. Image preprocessing unit, used for: Based on the trigger signal of each cycle, a cycle number corresponding to the corresponding cycle is generated. At the same time, the timestamp when the optical sensor is started is obtained, and the cycle number and timestamp are added to the original multi-person image. Based on the added results, the original multi-person image is sent to the edge computing node according to the communication link, and the edge computing node performs timestamp alignment and image enhancement processing on the original multi-person image to obtain a preprocessed multi-person image.
[0085] The beneficial effects of the above technical solution are as follows: by constructing a unified synchronous acquisition and distributed processing architecture, high-precision temporal alignment and high-quality preprocessing of multi-person scene image data are achieved, effectively eliminating the acquisition time deviation between multiple sensors and ensuring the temporal consistency of subsequent processed data. At the same time, the image quality is improved through edge-side enhancement processing, providing a more reliable input for label-free feature extraction, and ensuring the foundation for low latency and high precision of the entire system from the data source.
[0086] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A low-latency optical motion capture and markerless real-time co- localization method for multi-person XR large spaces, characterized in that, include: Step 1: Based on the optical sensors deployed in an array in a large space, multi-person images are captured synchronously in real time, and the obtained multi-person images are time-stamp aligned and image enhancement processed based on each edge computing node to obtain pre-processed multi-person images; Step 2: Extract features from the preprocessed multi-person images using a convolutional neural network to obtain unlabeled 2D information of skeletal key points; Step 3: Perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and construct the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. Step 4: Determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and perform collaborative localization correction on the three-dimensional position and posture of different human bodies in a large space to obtain target localization data; Step 5: Transmit the target location data to the XR rendering engine in real time for rendering.
2. The low-latency optical motion capture and markerless real-time co-located positioning method for multi-person XR large space according to claim 1, wherein, In step 1, multi-person images are synchronously and in real-time captured by optical sensors deployed in an array in a large space. The obtained multi-person images are then time-stamp aligned and image enhancement processed based on each edge computing node to obtain pre-processed multi-person images, including: Establish communication links between optical sensors deployed in arrays in a large space and edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. The optical sensors deployed in the array are synchronously controlled based on periodic trigger signals, and raw multi-person images containing multiple human targets are captured based on the synchronous control results. Based on the trigger signal of each cycle, a cycle number corresponding to the corresponding cycle is generated. At the same time, the timestamp when the optical sensor is started is obtained, and the cycle number and timestamp are added to the original multi-person image. Based on the added results, the original multi-person image is sent to the edge computing node according to the communication link. The edge computing node then performs timestamp alignment and image enhancement processing on the original multi-person image to obtain a preprocessed multi-person image.
3. The low-latency optical motion capture and markerless real-time cooperative localization method for multi-person XR large spaces according to claim 2, characterized in that, Based on edge computing nodes, the original multi-person images are time-stamp aligned and image enhancement is performed to obtain preprocessed multi-person images, including: The edge computing node receives raw multi-person images uploaded by optical sensors and groups the raw multi-person images uploaded by each optical sensor into the same batch based on the period number carried in the raw multi-person images. Extract the timestamps of each original multi-person image in the same batch group. At the same time, access the historical parameters of the global synchronization clock source, extract the time base corresponding to the cycle number, and perform time offset compensation on the timestamps of each original multi-person image based on the time base. Based on the time offset compensation results, the original multi-person image is timestamped and then the human target area and background area in the original multi-person image are separated based on the timestamping alignment results. Based on the separation results, the pixel grayscale of the human target area and the background area are determined respectively. When the pixel grayscale does not meet the preset standard, the original multi-person image is enhanced based on the preset contrast curve. Preprocessed multi-person images are obtained based on timestamp alignment results and image enhancement processing results.
4. The low-latency optical motion capture and markerless real-time collaborative localization method for multi-person XR large spaces according to claim 1, characterized in that, In step 2, feature extraction is performed on the preprocessed multi-person images based on a convolutional neural network to obtain unlabeled two-dimensional information of skeletal key points, including: The preprocessed multi-person images are obtained, and multi-level downsampling convolution processing is performed on the preprocessed multi-person images based on convolutional neural networks to obtain multi-level image feature maps with different spatial resolutions. Based on multi-level image feature maps, deep semantic features and shallow detail features of preprocessed multi-person images are extracted, and the deep semantic features and shallow detail features are fused to obtain an enhanced multi-scale feature map. Convolutional neural networks are used to perform convolution and deconvolution processing on the enhanced multi-scale feature map to determine the existence probability distribution of each human instance in the enhanced multi-scale feature map for each preset skeletal key point type in the two-dimensional space of the image, and a two-dimensional probability heat map corresponding to each preset skeletal key point type is obtained based on the existence probability distribution. Meanwhile, based on the results of convolution and deconvolution, a two-dimensional vector field corresponding to the preprocessed multi-person image is extracted, and the direction information of any pixel in the image space of the preprocessed multi-person image pointing to the adjacent key points inside the human body instance is determined based on the two-dimensional vector field. Based on the two-dimensional probability heat map, the candidate key point positions corresponding to each preset skeletal key point type are determined, and the candidate key points corresponding to the candidate key point positions in the preprocessed multi-person images are clustered based on the direction information. Based on the clustering results, the candidate key point set corresponding to different independent human instances is obtained. Two-dimensional coordinate data of each candidate keypoint in the candidate keypoint set corresponding to different independent human instances are extracted to obtain unlabeled skeletal keypoint two-dimensional information.
5. The low-latency optical motion capture and markerless real-time cooperative localization method for multi-person XR large spaces according to claim 1, characterized in that, In step 3, the 2D information of the skeletal keypoints corresponding to each edge computing node is triangulated from multiple perspectives to obtain the 3D skeletal keypoints corresponding to each human body. Based on these 3D skeletal keypoints, a human skeletal model is constructed for each human body, including: Obtain the 2D information of the skeletal key points corresponding to each edge computing node, and extract the timestamp, type code, detection confidence and 2D coordinates corresponding to the 2D information of the skeletal key points; Based on the timestamp and the viewpoint corresponding to the optical sensor, the 2D information of the skeletal key points of each edge computing node is time-aligned and grouped, and multi-view 2D data is obtained based on the grouping results. Based on multi-view two-dimensional data, cross-view human instance identity matching is performed on different human instances, and based on the human instance identity matching results, two-dimensional data belonging to the same independent human instance under different views are associated with the same human identity identifier to obtain a multi-view data group for each independent human instance. Based on the intrinsic parameter matrix of the optical sensor under different viewpoints and the extrinsic parameter matrix relative to the global world coordinate system, the projection matrix that projects the three-dimensional world coordinates to the two-dimensional image coordinates under different viewpoints is determined. Based on the type coding corresponding to the two-dimensional information of the skeletal key points, the two-dimensional coordinates and detection confidence of each skeletal key point under each type coding in the multi-view data group are extracted, and the observation angle between the view and each skeletal key point is determined based on the two-dimensional coordinates. Based on the observation angle quantification index and the detection confidence of each skeletal key point under each viewpoint, the target weight of each skeletal key point under different viewpoints is determined, and the three-dimensional spatial coordinates of each skeletal key point are determined based on the projection matrix under each viewpoint and the target weight of each skeletal key point under different viewpoints. Based on the three-dimensional spatial coordinates of each skeletal key point, the corresponding three-dimensional skeletal key points of each human body are obtained, and a human skeletal model corresponding to each human body is constructed based on the three-dimensional skeletal key points.
6. The low-latency optical motion capture and markerless real-time cooperative localization method for multi-person XR large spaces according to claim 5, characterized in that, Based on the three-dimensional spatial coordinates of each skeletal key point, the corresponding three-dimensional skeletal key points of each human body are obtained, and a human skeletal model corresponding to each human body is constructed based on the three-dimensional skeletal key points, including: Extract the three-dimensional spatial coordinates of each three-dimensional bone key point, and determine the instantaneous three-dimensional length of the bone segment formed by adjacent three-dimensional bone key points based on the three-dimensional spatial coordinates. Meanwhile, based on prior knowledge of the inherent length ratio of human bones, the range of values for the length of human bone segments is determined, and the instantaneous three-dimensional length of the bone segments is compared with the range of values. If the instantaneous three-dimensional length of a bone segment is outside the range of values, then it is determined that at least one of the two three-dimensional bone keypoints corresponding to the current bone segment is an abnormal bone keypoint. At the same time, the instantaneous lengths of the associated bone segments formed by two 3D skeletal keypoints and other adjacent 3D skeletal keypoints are determined respectively, and the instantaneous lengths of the associated bone segments are compared with the corresponding standard value ranges. The preprocessed multi-person images in a continuous frame sequence are traversed, and the position trajectory smoothness of two 3D skeletal key points is determined based on the traversal results. The reliability index of each 3D skeletal key point is determined based on the position trajectory smoothness. The key points of the skeleton to be corrected are determined based on the reliability index and the comparison results of the instantaneous length of the associated skeletal segment with the corresponding standard value range. Based on the bone type of the current bone segment, the corresponding historical length average and stable position information of adjacent 3D bone key points are obtained from the database, and the position of the bone key points to be corrected is corrected based on the historical length average and stable position information of adjacent 3D bone key points. The initial human skeleton model is obtained from the model library and then initialized. Based on the position correction results and initialization results, the instantaneous three-dimensional lengths of each human bone segment are mapped to the same position in the initial human bone model to obtain the initial lengths of all bone segments in the initial human bone model. The obtained initial length is proportionally aligned with a predefined standard skeleton model, and the initial scale parameters of the initial human skeleton model are obtained based on the proportional alignment result. The human skeleton model corresponding to each human body is obtained based on the initial scale parameters.
7. The method for low-latency optical motion capture and markerless real-time cooperative localization for multi-person XR large spaces according to claim 1, characterized in that, In step 4, the 3D positions and postures of different human bodies are determined based on the human skeletal model, and collaborative localization correction is performed on the 3D positions and postures of different human bodies in a large space to obtain target localization data, including: Obtain human skeleton models corresponding to different human bodies, parse the human skeleton models, extract the three-dimensional spatial coordinates of the root node of the human skeleton model, and use the three-dimensional spatial coordinates of the root node as the three-dimensional position of the current human body. At the same time, the rotation parameters of different joints of the current human body are determined based on the human skeleton model, and the rotation parameters of different joints of the current human body are used as the current posture of the human body. The three-dimensional positions and postures of different human bodies at the same time stamp are summarized to obtain a set of original pose states of multiple people, and the three-dimensional positions and postures of different human bodies in the set of original pose states of multiple people are transformed into a predefined unified world coordinate system. Based on a pre-set human kinematics model, the three-dimensional position and posture of different human bodies in a unified world coordinate system are analyzed to predict the target state of each human body at the next moment. The spatial relative relationship between the original pose data of different human bodies at the same moment is used as a constraint to correct the target state. Based on the correction results, the final three-dimensional position and attitude parameter set of each human body is obtained, and the final three-dimensional position and attitude parameter set of each human body is used as the target positioning data of each human body.
8. The low-latency optical motion capture and markerless real-time collaborative localization method for multi-person XR large spaces according to claim 1, characterized in that, In step 5, the target positioning data is transmitted to the XR rendering engine in real time for rendering processing, including: Based on the target positioning data, the identity identifier, three-dimensional position and posture parameters of each human body are extracted, and the extracted identity identifier, three-dimensional position and posture parameters are encapsulated into a data packet to be transmitted; The data packets to be transmitted are sent to the data receiving port corresponding to the XR rendering engine based on the low-latency network interface, and the data packets to be transmitted are inserted into the rendering queue after the XR rendering engine receives the data packets to be transmitted. The insertion result-driven XR rendering engine maps the 3D position and pose parameters according to the identity identifier in the data packet to be transmitted and drives the corresponding virtual avatar skeleton model in the scene. Based on the driving results and combined with the virtual scene content, the pre-processed multi-person images are rendered and output using the XR rendering engine.
9. A low-latency optical motion capture and markerless real-time cooperative positioning system for multi-person XR large spaces, characterized in that, include: The image acquisition and preprocessing module is used to synchronously and in real time capture multi-person images based on optical sensors deployed in an array in a large space, and to perform timestamp alignment and image enhancement processing on the obtained multi-person images based on each edge computing node to obtain preprocessed multi-person images. The feature extraction module is used to extract features from preprocessed multi-person images based on convolutional neural networks to obtain unlabeled two-dimensional information of skeletal key points; The model building module is used to perform multi-view triangulation on the two-dimensional information of the skeletal key points corresponding to each edge computing node to obtain the three-dimensional skeletal key points corresponding to each human body, and to build the human skeleton model corresponding to each human body based on the three-dimensional skeletal key points. The positioning and correction module is used to determine the three-dimensional position and posture of different human bodies based on the human skeleton model, and to perform collaborative positioning and correction of the three-dimensional position and posture of different human bodies in a large space to obtain target positioning data. The rendering module is used to transmit target positioning data to the XR rendering engine in real time for rendering processing.
10. A low-latency optical motion capture and markerless real-time cooperative positioning system for large-scale multi-user XR applications according to claim 9, characterized in that, The image acquisition and preprocessing module includes: Image acquisition unit, used for: Establish communication links between optical sensors deployed in arrays in a large space and edge computing nodes. At the same time, configure a global synchronization clock source for each optical sensor and send periodic trigger signals to each optical sensor based on the global synchronization clock source. The optical sensors deployed in the array are synchronously controlled based on periodic trigger signals, and raw multi-person images containing multiple human targets are captured based on the synchronous control results. Image preprocessing unit, used for: Based on the trigger signal of each cycle, a cycle number corresponding to the corresponding cycle is generated. At the same time, the timestamp when the optical sensor is started is obtained, and the cycle number and timestamp are added to the original multi-person image. Based on the added results, the original multi-person image is sent to the edge computing node according to the communication link. The edge computing node then performs timestamp alignment and image enhancement processing on the original multi-person image to obtain a preprocessed multi-person image.