Human motion capturing method based on sparse inertial sensor

Through the sparse inertial sensor and text supervision method, combined with uncertainty guidance and spatial relationship modeling, the light occlusion problem of visual methods and the low accuracy and under-constraint problem of inertial sensors are solved, achieving higher accuracy of human motion capture.

CN120406723AActive Publication Date: 2025-08-01UNIV OF SCI & TECH BEIJING
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510335385.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-08-01
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing human motion reconstruction process based on vision methods has problems with light intensity and occlusion, while the human motion reconstruction method based on inertial sensors has problems with low accuracy, under-constraint and ambiguity.

Method used

The method of sparse inertial sensor combined with text supervision is adopted to collect data from key parts of the human body through six inertial measurement units, and the human body motion capture is used using text encoder, sensor encoder and text sensing fusion module, including uncertainty-guided resampling, spatial relationship modeling and time alignment to optimize sensor output and text feature fusion.

Benefits of technology

It improves the naturalness and accuracy of human posture reconstruction, solves the problem of fuzzy sensor output, and achieves higher motion capture accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406723A_ABST
    Figure CN120406723A_ABST
Patent Text Reader

Abstract

The invention discloses a human motion capturing method based on a sparse inertial sensor, and belongs to the technical field of man-machine interaction, and the method comprises the steps: collecting sensor data corresponding to the current motion of a detected object through the sparse inertial sensor; obtaining text description information corresponding to the current action of the tested object; and inputting the sensor data corresponding to the current action of the detected object and the text description information corresponding to the current action of the detected object into a preset human body motion capture model to realize human body motion capture. Through text supervision, the human motion capturing method provided by the invention not only can distinguish undefined actions such as sitting and standing, but also can generate more accurate and natural actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a human motion capture method based on sparse inertial sensors. Background Art

[0002] Human motion reconstruction is a key technology for accurately capturing three-dimensional human kinematics and has important applications in games, sports, healthcare, and film production. One of the popular methods in motion reconstruction is the vision-based motion capture method. Akash Sengupta et al. [Sengupta.A, Budvytis.I, and Cipolla.R. HuManiFlow:Ancestor-Conditioned Normalis-ing Flows on SO(3)Manifolds for Human Pose and ShapeDistribution Estimation.In Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition(CVPR),4779–4789,2023.doi:arxiv-2305.06968] analyzed and studied the images of individuals to determine their respective human postures. However, the vision-based motion capture method has problems of occlusion and illumination.

[0003] With the rapid development of wearable technologies, various sensor devices have also been used to reconstruct human motion. The Xsens system uses 17 densely located IMUs to facilitate the reconstruction of human postures. However, this method has drawbacks, mainly due to the invasive impact on human motion caused by the dense imu placement and its huge cost. With the proposal of the concept of sparse sensor placement, Martin Schepers et al. [Schepers.M, Giuberti.M, Bellusci.G et al. Xsens MVN:Consistent tracking of human motion using inertial sensing.Xsens Technol,1(8):1–8,2018.doi:10.13140 / RG.2.2.22099.07205] achieved human motion reconstruction with only 4 to 5 accelerometers by retrieving pre-recorded postures with similar accelerations from a motion reconstruction database. However, using sparse inertial sensors for motion reconstruction poses an under-constrained problem: different postures can produce the same sensor output. As Figure 2As shown, when a subject is sitting and standing, the sensor produces similar rotation matrices and acceleration outputs, making accurate discrimination between these postures challenging.

[0004] With the advent of the deep learning era, Huang et al. [Huang.Y, Kaufmann.M, Aksan.E, Black.MJ, Hilliges.O, and Pons-Moll.G. Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Transactions on Graphics (TOG), 37(6):1–15, 2018. doi:10.1145 / 3272127.3275108] synthesized inertial data from a broad human motion dataset and trained a deep neural network model based on a bidirectional recurrent neural network that directly maps IMU inputs to body poses. Xu Feng et al. [Yi.X, Zhou.Y, and Xu.F. Transpose: Real-time 3d human translation and pose estimation with six inertial sensors. ACM Transactions on Graphics (TOG), 40(4):1–13, 2021. doi:10.1145 / 3450626.3459786.] decomposed body pose estimation into a multi-stage task to improve the accuracy of pose regression by using joint positions as intermediate representations.

[0005] With the development of the emerging field of multimodal processing, text and its rich semantic information and ease of annotation are increasingly used in the field of human motion. In terms of the instability of IMU readings, relying solely on sensor data is not sufficient to solve the problem of ambiguity. Text has rich motion information, which can help the model identify the human motion state and solve the ambiguity problem. Tevet et al. [Tevet.G, Gordon.B, Hertz.A, Bermano.AH, and Cohen-Or.D, 2022. Motionclip: Exposing human motion generation to clip space. In Proceedings of the European conference on computer vision (ECCV), 358–374, 2022, Springer. doi: arxiv-2203.08063] proposed that high-quality 3D human motion can be generated from text descriptions, assuming that text supervision can eliminate the ambiguity of actions, thereby improving the naturalness and accuracy of the generated actions.

[0006] In summary, the process of human motion reconstruction based on vision methods has problems such as light intensity and occlusion; while the human motion reconstruction method based on inertial sensors has problems such as low accuracy, under-constraint, and ambiguity. Summary of the Invention

[0007] The present invention provides a human motion capture method based on sparse inertial sensors to solve the problems of light intensity and occlusion in the current process of human motion reconstruction based on vision methods; while the human motion reconstruction method based on inertial sensors has technical problems such as low accuracy, under-constraint, and ambiguity.

[0008] To solve the above technical problems, the present invention provides the following technical solutions:

[0009] On the one hand, the present invention provides a human motion capture method based on sparse inertial sensors, including:

[0010] Collecting sensor data corresponding to the current action of the measured object through sparse inertial sensors;

[0011] Obtaining text description information corresponding to the current action of the measured object;

[0012] Inputting the sensor data corresponding to the current action of the measured object and the text description information corresponding to the current action of the measured object into a preset human motion capture model to achieve human motion capture.

[0013] Further, the sparse inertial sensors include six inertial measurement units;

[0014] When collecting sensor data corresponding to the current actions of the object to be measured, six inertial measurement units are respectively placed at two legs, two wrists, the head, and the root of the pelvis of the object to be measured.

[0015] Furthermore, the human motion capture model includes: a text encoder, a sensor encoder, and a text-sensor fusion module;

[0016] The text encoder is used to extract semantic information from the text description information, convert the text description information into an embedding sequence, and use it as text features;

[0017] The sensor encoder is used to capture the relationships between sparse sensors through spatial modeling based on the sensor data, and at the same time introduce the concept of sensor uncertainty to optimize the sensor output and adjust the relative contribution of each sensor, so as to encode the sensor data into an embedding sequence containing in-frame spatial relationships and use it as sensor features;

[0018] The text-sensor fusion module is used to fuse the text features and the sensor features to obtain fusion features, and realize human motion capture based on the fusion features.

[0019] Furthermore, the text encoder is composed of the first 4 layers of the text encoder of the frozen CLIP VIT / B32 model and two additional transformer layers. After the input text description information passes through the first 4 layers of the text encoder of the frozen CLIP VIT / B32 model and two additional transformer layers in sequence, the corresponding embedding sequence is obtained.

[0020] Furthermore, the sensor encoder includes: an uncertainty-guided resampling module and a spatial relationship modeling module;

[0021] The uncertainty-guided resampling module is used to input the original sensor data into an uncertainty regression head to obtain the uncertainty of each channel of the sensor data; during training, after obtaining the uncertainty of each channel, the sensor data is resampled using a Gaussian distribution, where the data with uncertainty lower than the preset threshold remains unchanged, while the data with uncertainty higher than the preset threshold is resampled, so as to optimize the sensor data; during the inference process, a simple regression is performed on the uncertainty of each channel;

[0022] The spatial relationship modeling module is used to calculate the spatial correlation between sensors, and at the same time consider the uncertainty of the data collected by each sensor.

[0023] Furthermore, the spatial relationship modeling module is specifically used for:

[0024] When dealing with the corresponding uncertainty σ tat the t-th frame After sampling the sensor data of map it to an n×c feature embedding Z t where n represents the number of sensors and c represents the dimension of spatial features;

[0025] Perform self-attention on Z t to calculate the spatial correlation between sensors: For calculating the attention of the t-th frame between the j-th sensor and the k-th sensor, by dividing the attention score by to account for the uncertainty of the k-th sensor The formula is:

[0026]

[0027] where P Q , are the query and key projection matrices respectively; T is the total number of frames of the sensor data sequence; and are the feature vectors of the j-th and k-th sensors at the t-th frame respectively.

[0028] Furthermore, the text-sensor fusion module includes a time alignment module, a contrast learning module, and a feature fusion module;

[0029] The time alignment module is used for time alignment between sensor features and text features;

[0030] The contrast learning module is used to apply contrast learning to align sensor features and text features in a high-dimensional space, achieving semantic synchronization of features, that is, in a batch, maximizing the similarity between sensor features and their corresponding text features while minimizing the similarity with other text features;

[0031] The feature fusion module is used to embed the sensor features and text features that have completed time alignment and semantic synchronization into the corresponding time positions, and fuse the sensor features and text features through a cross-attention mechanism.

[0032] Furthermore, the time alignment module restricts the range of attention calculation through window self-attention and sliding window self-attention mechanisms, limiting self-attention within each sub-window; and introduces a concept of locality similar to convolution in this process, and extracts sensor features at different granularity levels through patch merging operations, while reducing the computational complexity from a quadratic relationship with the length of the input data sequence to a linear relationship.

[0033] Furthermore, the processing process of the time alignment module for the input data includes:

[0034] Given a window size I, a sensor feature data sequence of length L is divided into non - overlapping sub - intervals; The window self - attention mechanism is first executed within non - overlapping sub - intervals; To establish an interconnection between these non - overlapping sub - intervals, an inspired sliding window self - attention mechanism is adopted; Among them, the local window attention mechanism and the sliding window self - attention alternate.

[0035] Furthermore, the loss function of the human motion capture model is expressed as:

[0036]

[0037] where is the uncertainty loss; is the contrastive loss; is the reconstruction loss;

[0038] The uncertainty loss has the expression:

[0039]

[0040] In the formula, δ is the loss weight; represents the uncertainty of the j - th sensor at the t - th frame; and respectively represent the squared differences between the predicted value and the true value of the joint rotation angle and the joint position at the t - th frame; T is the total number of frames of the sensor data sequence.

[0041] The contrastive loss has the expression:

[0042]

[0043] In the formula, γ is the loss weight; s i,j represents the similarity calculated by cosine similarity between the i - th sensor sequence and the j - th text; τ is the learnable temperature parameter that controls the distribution concentration; B represents the number of samples in the training batch; e is the natural constant; s i,i represents the cosine similarity between the i - th sensor sequence and the i - th text description; s j,i represents the cosine similarity between the j - th text description and the i - th sensor sequence;

[0044] The reconstruction loss has the expression:

[0045]

[0046] In the formula, λ, α, β are all loss weights; represents the predicted joint rotation q and the true rotation The mean squared error between; denotes the mean squared error between the predicted joint position p and the true position ; denotes the mean squared error between the predicted root node velocity and the true velocity ; T represents the total number of frames in the sensor data sequence; x t represents the model prediction value at the t-th frame; represents the true value at the t-th frame.

[0047] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.

[0048] On another aspect, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.

[0049] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0050] 1. The present invention is enhanced through text supervision, and the rich semantic information contained in the text is used to enhance the naturalness and accuracy of modeling human postures.

[0051] 2. The present invention introduces a spatial relationship representation model, which can calculate the correlation between sensors within a frame and also consider the uncertainty of each IMU.

[0052] 3. The present invention realizes the temporal alignment between sensor features and text semantics by designing an HTT module. And a contrastive learning mechanism is adopted to optimize the alignment between the two modalities in the high-dimensional space. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0054] Figure 1 is a flowchart of a human motion capture method based on sparse inertial sensors provided by an embodiment of the present invention;

[0055] Figure 2 is a schematic diagram of the under-constrained problem of sparse inertial sensors provided by an embodiment of the present invention;

[0056] Figure 3is a framework diagram of a human motion capture model provided by an embodiment of the present invention;

[0057] Figure 4 Schematic diagram of the principles of window self-attention (left) and sliding window self-attention (right) provided by an embodiment of the present invention, where IMU is an inertial sensor;

[0058] Figure 5 Schematic diagram of an effective method for batch computing self-attention in the context of sliding window partitioning provided by an embodiment of the present invention;

[0059] Figure 6 Schematic diagram of mesh error distribution and qualitative comparison between the method of the present invention (with / without text) and Transpose provided by an embodiment of the present invention;

[0060] Figure 7 Schematic diagram of a comparison between the method of the present invention provided by an embodiment of the present invention in a sitting position (with / without text) and Transpose (focusing on the analysis of upper leg rotation error);

[0061] Figure 8 Schematic diagram of the time evolution of uncertainty of six sensors provided by an embodiment of the present invention (each row represents a different sensor, and the color change indicates the change in uncertainty);

[0062] Figure 9 This is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0064] First, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present concepts in a concrete manner. In addition, in the embodiments of the present invention, the meaning of "and / or" can be both or either of the two.

[0065] First embodiment

[0066] In view of the problems of light intensity and occlusion in the existing human motion reconstruction process based on vision methods, and the technical problems of low accuracy, under-constraint, and ambiguity in the human motion reconstruction method based on inertial sensors, this embodiment provides a human motion capture method based on sparse inertial sensors, which reconstructs accurate human postures by using data from six IMUs placed on the legs, wrists, head, and pelvis root and text supervision. The human motion capture method of this embodiment can be implemented by an electronic device, which can be a terminal or a server. The execution process of this method is as Figure 1 shown, including the following steps:

[0067] S1, collecting sensor data corresponding to the current action of the measured object through sparse inertial sensors;

[0068] Specifically, in this embodiment, the sparse inertial sensors include six inertial measurement units IMUs; when collecting sensor data corresponding to the current action of the measured object, the six inertial measurement units IMUs are respectively placed on the two legs, two wrists, head, and pelvis root of the measured object. The six inertial measurement units IMUs capture individual motion data, including three-axis acceleration and rotation matrix.

[0069] S2, obtaining text description information corresponding to the current action of the measured object;

[0070] Among them, it should be noted that the text description information is used to describe the human action.

[0071] S3, inputting the sensor data corresponding to the current action of the measured object and the text description information corresponding to the current action of the measured object into a preset human motion capture model to achieve human motion capture;

[0072] Specifically, the framework of the human motion capture model in this embodiment is as Figure 3 shown, including: a text encoder, a sensor encoder, and a text-sensor fusion module; each module will be described in detail below.

[0073] The text encoder is used to extract semantic information from the text, convert the motion semantic information contained in the text into a series of embeddings as text features, including the embedding of the CLS token and multiple text labels; among them, in this embodiment, the first 4 layers of the text encoder of the frozen CLIP VIT / B32 model and two additional transformer layers are used to construct the text encoder of the present invention, which is responsible for converting the input text into a series of embeddings, extracting the semantic information in the text, and providing support for 3D human action reconstruction through fusion with sensor features.

[0074] Specifically, the text encoder is responsible for converting the input text W, such as ["catching the ball with both hands", "transition"], into an embedding sequence: {Wcls, W1, …, WN}. Given a sequence of text labels W, it is initially tokenized and mapped to a token sequence using CLIP and a randomly initialized tensor is added in front as the [CLS] token. W provides two types of semantic labels: sequence-level labels and frame-level labels, as defined in the dataset configuration. For frame-level labels, although each frame has its own text description, they are largely repetitive. For example, the label "walking" may be applied continuously to a series of frames. To reduce the computational load, only the non-repetitive frame-level texts are sorted in chronological order as the input. For sequence labels, if the total number is less than the threshold M, all sequence labels are used as the input in this embodiment. Otherwise, one-third of the labels are selected based on the time information, specifically the labels that best match the sensor subsequence. To distinguish between sequence-level labels and frame-level labels, two learnable group position embeddings G are developed for each label. In addition, sine position embeddings P are utilized to independently calculate the time information for both the sequence and frame levels to adapt to their unique characteristics.

[0075]

[0076] The processed features and the [CLS] token are input into the self-attention layer to better extract the text semantics.

[0077] The sensor encoder is used to capture the complex relationships between sparse sensors based on sensor data through spatial modeling, while introducing the concept of sensor uncertainty to optimize the sensor output and adjust the relative contribution of each sensor, so as to encode the sensor data into an embedding sequence containing the spatial relationships within the frame as sensor features.

[0078] Furthermore, the sensor encoder of this embodiment includes: an uncertainty-guided resampling module and a spatial relationship modeling module; to implement the resampling strategy and the spatial attention mechanism, both of which are guided by the estimated uncertainty of each IMU. With the input provided in the form of three-axis acceleration and rotation matrix, the motion sequence composed of sensor data frames is encoded into an embedding sequence containing the spatial relationships within the frame. Among them, the uncertainty-guided resampling module is used to input the original sensor data into an uncertainty regression head to obtain the uncertainty of each channel of the sensor data; the spatial relationship modeling module calculates the spatial correlation between sensors by using the uncertainty-guided spatial attention module UGSA, while considering the uncertainty of each IMU. Thus, the model will extract different features of different body regions at the same moment, and the sensors will achieve the effect of adaptive adjustment. The specific data processing process is as follows:

[0079] 1. Uncertainty Estimation

[0080] Estimate the uncertainty of the input IMU data through uncertainty loss to estimate the uncertainty of each IMU reading, where the original IMU readings are represented as Take it as the input of the uncertainty regression head to generate uncertainty for each channel

[0081] 2. Uncertainty-Guided Resampling

[0082] In this embodiment, resampling is performed on the IMU readings represented as from the Gaussian distribution where X t is used as the mean value, and the predicted uncertainty σ t is used as the variance mean. This resampling method ensures that the values with low uncertainty remain basically unchanged, while the values with high uncertainty are resampled, thus optimizing the sensor data. It should be noted that the resampling process is only used during training. During the inference process, simple regression is performed on the uncertainty of each channel, and the original sensor reading X t is used as This stage applies the reparameterization trick to achieve efficient gradient descent by sampling and calculates The formula is:

[0083] 3. Uncertainty-Guided Spatial Attention

[0084] After sampling the IMU readings of the t-th frame with the corresponding uncertainty σ t , map to the 6×c feature embedding Z , where 6 represents the number of sensors of the features, and c represents the dimension of the spatial features. Then perform self-attention on Z t . For calculating the attention of the t-th frame between two sensors, denoted as j and k, consider the uncertainty of sensor k by dividing the attention score by t (summed over its 12 channels), and the formula is: (summed over its 12 channels), the formula is:

[0085]

[0086] where P Q , are the query and key projection matrices respectively; T is the total number of frames of the sensor data sequence; and are the feature vectors of the j-th and k-th sensors at the t-th frame respectively.

[0087] ​This unique change ensures that sensors with high uncertainty contribute less when calculating spatial correlation. Frame t 0 t The output of the UGSA module in matches the input dimension t After flattening the output vectors from T frames are concatenated to form

[0088] The text-sensor fusion module is mainly used to align and fuse two-modal features. Specifically, the text-sensor fusion module obtains spatio-temporal fused sensor features through a hierarchical temporal transformation (HTT) module for temporal synchronization with text features. Subsequently, contrastive learning is applied in the high-dimensional space to align multi-modal features, and finally cross-attention is applied for feature fusion. Through temporal alignment, semantic synchronization between sensor features and text semantics, and by using the cross-attention mechanism to merge the features of the two modalities, fused features are obtained, based on which human motion capture can be achieved. That is: the spatial embeddings obtained above are processed by a hierarchical temporal transformation module (HTT) to extract a unified spatio-temporal fused representation. Before applying cross-attention in the fusion process, text-sensor contrastive learning is strategically implemented to refine the alignment between the unimodal representations of the two modalities. Finally, a simple regression head is used to derive the human pose rotation data of 15 key points.

[0089] Specifically, the text-sensor fusion module includes an HTT module, a contrastive learning module, and a feature fusion module.

[0090] The HTT module is used for temporal alignment between sensor features and text semantics. By obtaining the spatio-temporal fused representation of sensor features, it synchronizes with text features. The HTT module constrains the scope of attention calculation through window self-attention and sliding window self-attention mechanisms, and extracts sensor features at different granularity levels through patch merging operations, while reducing the computational complexity of the transformer from a quadratic relationship with the sequence length to a linear relationship. It should be noted that the present invention assumes that information obtained from adjacent frames is crucial for the estimation of the current frame pose. In response to this assumption, window self-attention (W-SA) and sliding window self-attention (SW-SA) mechanisms are combined to limit the scope of attention calculation, and a concept of convolution-like locality is introduced in this process. In addition, to integrate information from distant frames and thus expand the receptive field, a patch merging operation is implemented. These methods help to extract sensor features at different granularity levels, while reducing the computational complexity of the transformer from a quadratic relationship with the sequence length to a linear relationship.

[0091] Given a window size I, a sensor sequence of length L is divided into non-overlapping sub-intervals. The local window attention calculation is first performed within these sub-intervals. To establish interconnections between these non-overlapping segments, the present invention adopts an inspired sliding window attention module, thus implementing a new partitioning method that can enhance the self-attention across segments. W-SA and SW-SA always alternate, constituting a hierarchical temporal transformation module, as shown in the upper right corner of Figure 3 as shown.

[0092] When applying the sliding window attention and time series, the window count increases from to resulting in some windows being smaller than I. To solve this problem, the present invention introduces batch calculation with left circular shift, as shown in Figure 5 as shown. This transformation can produce windows with discontinuous sub-windows. In this regard, the present invention designs a masking mechanism that restricts the self-attention within each sub-window, maintains the number of batch windows, and ensures computational efficiency. After the calculation is completed, the original sequence order is restored.

[0093] In the patch merging operation, each process merges two adjacent tokens into one, effectively halving the token count and doubling the dimension of each token. Then the transformed tokens are input to the subsequent stage. In the last stage, patch merging is omitted, and the tokens are restored to their original dimension through operations such as linear projection. In the sensor sequence, in this instance, the output feature is mapped to a feature of size 1×C as the [CLS] token. This token forms the cumulative output together with F, containing spatio-temporal features {F cls ,F 1 ,...,F T}.

[0094] Contrastive learning is used to align the features of the sensor data and text semantics in the high-dimensional space, that is, in a batch, maximizing the similarity between the sensor sequence and its corresponding text, while minimizing the similarity with other texts.

[0095] Feature fusion, for the text and sensor features, embeds the corresponding time positions, and fuses the sensor features with the text features through cross-attention, thereby improving the accuracy of 3D human motion reconstruction. By combining the sensor data and text information, the model can better solve the ambiguity problem in motion reconstruction, thus generating more natural and accurate human postures. Specifically, given the sensor feature set {F cls ,F 1 ,...,F T} and the text feature set {W cls ,W 1 ,…,W N}, the present invention applies contrastive learning to align these features in a high-dimensional joint space, using the [CLS] token as an anchor point. Subsequently, the sensor features and text features are fused through cross-attention. The corresponding group embeddings and temporal position embeddings are designed for text and sensor features.

[0096] Based on the above, there are three training objectives for the human motion capture model of this embodiment: uncertainty learning of the sensor encoder, text-sensor contrastive learning of the unimodal encoder, and reconstruction loss of the text-sensor fusion module. The relevant equations are shown below. The parameters δ, γ, λ, α, β are used to balance different loss weights.

[0097] Uncertainty loss: The objective of the present invention is to estimate the uncertainty of the input IMU data. The present invention sets the uncertainty estimation loss as follows:

[0098]

[0099] In the formula, represents the uncertainty of the j-th sensor at the t-th frame; and respectively represent the squared differences between the predicted value and the true value of the joint rotation angle and the joint position at the t-th frame; T is the total number of frames of the sensor data sequence;

[0100] Contrastive loss: The present invention uses text-sensor contrastive learning to perform better unimodal representation before feature fusion. Given a batch of B text-sensor pairs, the model learns to maximize the similarity between the sensor sequence and its corresponding text, while minimizing the similarity with the other B-1 texts in the batch, and vice versa.

[0101] The contrastive loss is expressed as:

[0102]

[0103] In the formula, γ is the loss weight; s i,j represents the similarity calculated by the cosine similarity between the i-th sensor sequence and the j-th text; τ is a learnable temperature parameter that controls the concentration of the distribution; B represents the number of samples in the training batch; e is the natural constant; s i,i represents the cosine similarity between the i-th sensor sequence and the cosine similarity of the i-th text description; s j,i represents the cosine similarity between the j-th text description and the i-th sensor sequence;

[0104] Reconstruction loss: The model of the present invention is optimized to capture motion features by minimizing the L2 loss on the joint orientation q, joint position p, and root velocity s, as shown in the following equation:

[0105]

[0106] where λ, α, and β are all loss weights; represents the mean squared error between the predicted joint rotation q and the true rotation ; represents the mean squared error between the predicted joint position p and the true position ; represents the mean squared error between the predicted root node velocity and the true velocity ; T represents the total number of frames in the sensor data sequence; x t represents the predicted value of the model at the t-th frame; represents the true value at the t-th frame.

[0107] Calculate the average difference between the predicted value x of the model and the true value . The complete objective of the model is as follows:

[0108]

[0109] where, is the uncertainty loss; is the contrast loss; is the reconstruction loss.

[0110] For the above model, this embodiment uses two types of datasets: sensor data captured during human movement and corresponding text annotations for training and validation, including:

[0111] (1) Babel dataset: For semantic annotation, this dataset provides two-level text labels for nearly 43 hours of AMASS motion capture sequences. The sequence labels describe the overall action, while the frame labels detail each action in each frame. For the DIP-IMU dataset lacking Babel semantic annotation, the present invention manually added sequence-level labels.

[0112] (2) Motion Datasets: Considering the scarcity of real datasets and the extensive data requirements inherent in deep learning, the present invention synthesized a subset from the extensive AMASS dataset, which includes CMU, Eyes Japan, KIT, ACCAD, DFaust 67, HumanEva, MPI Limits, MPI mosh, and SFU datasets. Secondly, the DIP-IMU dataset contains IMU readings and pose parameters of approximately 90 minutes of activities of 10 subjects. The present invention reserved testers 9 and 10 specifically for evaluation, and the rest were used for training. The Totalcapture dataset contains 50 minutes of actions captured from 5 objects. The present invention used real IMU data for evaluation, but the ground truth and synthetic IMU readings were still integrated into the training set. Due to the lack of semantic annotations of Babel in some sequences, only 27 fully annotated sequences were used.

[0113] The present invention uses five metrics for pose evaluation: 1) SIP Err, which measures the average global rotation error when measuring the limbs, in degrees; 2) Angle Err, the average global rotation error of all body joints, in degrees; 3) Pos Err, the average Euclidean distance error of all joints when the spine is aligned, in centimeters; 4) Mesh Err, the average Euclidean distance error between the body mesh vertices and the aligned spine, in centimeters; 5) Jitter, the average jerk of all body joints in the predicted motion, in 100m / s 3 , reflecting the smoothness of the motion.

[0114] The training and evaluation scheme of the present invention was carried out on a system equipped with 1 Intel(R) Xeon(R) Silver 4110 CPU and 1 NVIDIA GeForce RTX 2080Ti GPU. The model of the present invention was developed using PyTorch 1.13.0 and further accelerated through CUDA 11.6. The model configuration of the present invention set the input sequence length T to 80 frames, the window and shift sizes to 20 frames and 10 frames respectively, and the threshold M to 15. The training process used a batch size of 40 and was initialized with an Adam optimizer with a learning rate of 0.00002. To balance the magnitude of the loss, γ and α were set to 1, β was set to 10, δ was set to 0.1, and λ was set to 0.01. The rotation information of 15 main joints defined in the SMPL model was focused on instead of all joints. In addition, the present invention applied a moving average with a window size of 15 to the output of the model, enhancing the smoothness of the predicted pose. [[ID=!0]]

[0115] In this embodiment, a comparative experiment is set up. In this embodiment, quantitative and qualitative comparisons are made with methods such as SIP, DIP, Transpose, PIP, and TIP on the Totalcapture and DIP-IMU datasets. During this process, the model with the best performance released by the author is used. For TIP, the author adopted a different human body format from this embodiment. Therefore, it is necessary to convert the output of TIP into the format of the present invention before comparison. The results of Totalcapture and DIP in the offline mode dataset settings are shown in Table 1. Different from previous methods, the method of the present invention does not consider all IMU readings when estimating the current pose. However, the present invention has achieved satisfactory results after integrating semantic information. Due to the few and not very detailed semantic annotations in the DIP dataset, the performance of the present invention on the DIP dataset is lower than that on the Totalcapture dataset, as Figure 6 shown, the present invention is good at handling ambiguous actions such as standing and sitting, and is good at capturing finer details.

[0116] Table 1 Method evaluation in offline mode

[0117]

[0118]

[0119] Due to the requirements of semantic annotation, the complete model of the present invention cannot reconstruct human motion in real time. Therefore, only the sensor encoder and the HTT module are used for real-time mode evaluation. The present invention accesses 70 past frames, 5 current frames, and 5 future frames through a sliding window method, and the tolerable delay is 83 milliseconds. As shown in Table 2, although lacking semantic information, the present invention still achieves superiority in multiple metrics, thus verifying the effectiveness of the network design of the present invention.

[0120] Table 2 Method evaluation in real-time mode

[0121]

[0122] This embodiment performs three ablations to verify the key design choices of the present invention:

[0123] (1) Without text semantic information;

[0124] (2) Without the Uncertainty-Guided Spatial Attention module (UGSA);

[0125] (3) Without the Hierarchical Time Transformer module (HTT).

[0126] Table 3 summarizes the results of the Totalcapture dataset (offline). The ablation experiments highlight the effectiveness of the proposed method, where the integration of semantic information is the most significant contribution, followed by the implementation of the UGSA and HTT modules.

[0127] Table 3 Evaluation of the ablation models on the Totalcapture dataset

[0128]

[0129] Without semantic information, the predictions of the model fluctuate in ambiguity, Figure 7 and this phenomenon is illustrated by the unstable alternation between sitting and standing postures in []. By incorporating simple semantic annotations like "sitting", the model of the present invention can effectively maintain the desired sitting posture. The research results show that the lack of uncertainty-guided spatial attention affects the accuracy of the results. Figure 8 It illustrates how uncertainty fluctuates over time. During complex movements such as squatting and crawling, the uncertainty of all sensors increases, especially in the hand area. Conversely, transitioning to a standing posture results in a significant reduction in uncertainty, with the leg sensors showing the lowest levels. The use of window attention and patch merging instead of global attention in the HTT module not only reduces the computational requirements but also improves the performance of almost all metrics (except jitter). These ablation findings confirm the excellent ability of the present invention to model sensor information and the ability to utilize semantic cues to generate more accurate and natural movements.

[0130] Second Embodiment

[0131] This embodiment provides an electronic device, as Figure 9 shown, the electronic device includes: a processor and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the above first embodiment. In addition, the electronic device may further include a transceiver, and the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.

[0132] Next, in combination with Figure 9 the various components of this electronic device will be specifically introduced:

[0133] Among them, the processor is the control center of the electronic device. The electronic device may include multiple processors, and each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a single processor or a collective term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or it can also be other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0134] In a specific implementation, as an embodiment, the processor may include one or more CPUs. For example Figure 9 CPU0 and CPU1 shown in [reference], of course, this is only an illustrative example.

[0135] The memory is used to store the software program for executing the solution of the present invention and is controlled by the processor for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.

[0136] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 9 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0137] The transceiver may include a receiver and a transmitter ( Figure 9 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 9 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0138] In addition, it should be noted that Figure 9 the structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have a different component layout. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above may refer to the technical effects described in the first embodiment above, so they will not be elaborated here.

[0139] Third Embodiment

[0140] This embodiment provides a computer-readable storage medium in which at least one instruction is stored. The instruction is loaded and executed by the processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be ROM, random access memory, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.

[0141] In addition, it should be noted that the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented using software, embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center containing one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0142] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes. Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.

[0144] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element. In addition, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, in this article, the character " / " generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship, which can be understood specifically with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural items (items). For example, at least one of a, b or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0145] In addition, it can be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not mean the sequence of execution order. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0146] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0147] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0148] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0149] Finally, it should be noted that the above are only the preferred embodiments of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once they know the basic creative concept of the present invention, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A human motion capture method based on sparse inertial sensors, characterized in that, Including: Collecting sensor data corresponding to the current action of the object to be measured through sparse inertial sensors; Obtaining text description information corresponding to the current action of the object to be measured; Inputting the sensor data corresponding to the current action of the object to be measured and the text description information corresponding to the current action of the object to be measured into a preset human motion capture model to achieve human motion capture.

2. The human motion capture method based on a sparse inertial sensor according to claim 1, characterized in that The sparse inertial sensor includes six inertial measurement units; When collecting sensor data corresponding to the current action of the object to be measured, the six inertial measurement units are respectively placed at two legs, two wrists, the head, and the root of the pelvis of the object to be measured.

3. The human motion capture method based on a sparse inertial sensor according to claim 1, wherein, The human motion capture model includes: a text encoder, a sensor encoder, and a text-sensor fusion module; The text encoder is used to extract semantic information in the text description information, convert the text description information into an embedding sequence, and use it as text features; The sensor encoder is used to capture the relationship between sparse sensors through spatial modeling based on the sensor data, and at the same time introduce the concept of sensor uncertainty to optimize the sensor output and adjust the relative contribution of each sensor, so as to encode the sensor data into an embedding sequence containing intra-frame spatial relationships as sensor features; The text-sensor fusion module is used to fuse the text features and the sensor features to obtain fusion features, and achieve human motion capture based on the fusion features.

4. The human motion capture method based on a sparse inertial sensor according to claim 3, characterized in that, The text encoder is composed of the first 4 layers of the text encoder of the frozen CLIP VIT / B32 model and two additional transformer layers. After the input text description information passes through the first 4 layers of the text encoder of the frozen CLIP VIT / B32 model and two additional transformer layers in sequence, the corresponding embedding sequence is obtained.

5. The method for human motion capture based on a sparse inertial sensor according to claim 3, wherein The sensor encoder includes: an uncertainty-guided resampling module and a spatial relationship modeling module; The uncertainty-guided resampling module is used to input the original sensor data into an uncertainty regression head to obtain the uncertainty of each channel of the sensor data; during training, after obtaining the uncertainty of each channel, the sensor data is resampled using a Gaussian distribution, where the data with uncertainty lower than the preset threshold remains unchanged, while the data with uncertainty higher than the preset threshold is resampled to optimize the sensor data; during the inference process, a simple regression is performed on the uncertainty of each channel; The spatial relationship modeling module is used to calculate the spatial correlation between sensors, and at the same time consider the uncertainty of the data collected by each sensor.

6. The human motion capture method based on a sparse inertial sensor according to claim 5, characterized in that Specifically, the spatial relationship modeling module is used for: After sampling the sensor data of the t-th frame with the corresponding uncertainty σ t the is mapped to an n×c feature embedding Z where n represents the number of sensors and c represents the dimension of the spatial features; t ​ On Z t Perform self-attention on it to calculate the spatial correlation between sensors: For calculating the attention of the j-th sensor and the k-th sensor at the t-th frame, consider the uncertainty of the k-th sensor by dividing the attention score by The formula is:​ Among them, are the query and key projection matrices respectively; T is the total number of frames of the sensor data sequence; and are the feature vectors of the j-th and k-th sensors at the t-th frame respectively.

7. The human motion capture method based on a sparse inertial sensor according to claim 3, characterized in that The text-sensor fusion module includes a time alignment module, a contrast learning module, and a feature fusion module; The time alignment module is used for time alignment between sensor features and text features; The contrast learning module is used to apply contrast learning to align sensor features and text features in a high-dimensional space to achieve semantic synchronization of features, that is, in a batch, maximize the similarity between sensor features and their corresponding text features, while minimizing the similarity with other text features; The feature fusion module is used to embed the sensor features and text features that have completed time alignment and semantic synchronization into the corresponding time positions, and fuse the sensor features and text features through the cross-attention mechanism.

8. The human motion capture method based on a sparse inertial sensor according to claim 7, characterized in that, The time alignment module constrains the scope of attention calculation through the window self-attention and sliding window self-attention mechanisms, restricting self-attention within each sub-window; and introduces the concept of locality similar to convolution in this process, and extracts sensor features at different granularity levels through patch merging operations, while reducing the computational complexity from a quadratic relationship with the length of the input data sequence to a linear relationship.

9. The human motion capture method based on a sparse inertial sensor according to claim 8, wherein The processing process of the time alignment module for the input data includes: Given the window size I, the sensor feature data sequence of length L is divided into non-overlapping sub-intervals; the window self-attention mechanism is first executed within non-overlapping sub-intervals; to establish an interconnection between these non-overlapping sub-intervals, an inspired sliding window self-attention mechanism is adopted; among them, the local window attention mechanism and the sliding window self-attention appear alternately.

10. The method for human motion capture based on a sparse inertial sensor according to claim 3, wherein, The loss function of the human motion capture model is expressed as: Among them, is the uncertainty loss; is the contrast loss; is the reconstruction loss; Uncertainty loss The expression is as follows: where δ is the loss weight; represents the uncertainty of the j-th sensor at the t-th frame; and respectively represent the squared differences between the predicted and true values of the joint rotation angle and the joint position at the t-th frame; T is the total number of frames in the sensor data sequence; Contrastive loss The expression is as follows: where γ is the loss weight; s i,j represents the similarity calculated by cosine similarity between the i-th sensor sequence and the j-th text; τ is a learnable temperature parameter for controlling the distribution concentration; B represents the number of samples in the training batch; e is the natural constant; s i,i represents the cosine similarity between the i-th sensor sequence and the i-th text description; s j,i represents the cosine similarity between the j-th text description and the i-th sensor sequence; Reconstruction loss The expression is as follows: where λ, α, and β are all loss weights; denotes the mean squared error between the predicted joint rotation q and the true rotation ; denotes the mean squared error between the predicted joint position p and the true position ; denotes the mean squared error between the predicted root node velocity and the true velocity ; T denotes the total number of frames in the sensor data sequence; x t represents the model prediction value at the t-th frame; represents the true value at the t-th frame.

Citation Information

Patent Citations

  • Human motion capturing method

    CN117170496A

  • Method and system for realizing three-dimensional human body posture estimation by fusing image and sparse IMU (Inertial Measurement Unit)

    CN117351564A

  • Human motion generation method and medium

    CN118823511A

  • Zero-sample multi-mode first-view-angle behavior recognition method introduced based on visual language knowledge

    CN119203019A

  • Limb nerve rehabilitation state evaluation method based on multi-modal sensor data fusion

    CN119318470A