Multi-modal fusion human motion representation and quality evaluation method

By employing a multimodal fusion method for human motion representation and quality assessment, and utilizing a teacher-student network framework and state-space model, the problem of insufficient modeling of modal differences and sensor characteristics is solved, achieving low-cost and high-precision motion assessment that is applicable to rehabilitation training and sports assessment.

CN122290880APending Publication Date: 2026-06-26HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-05-22
Publication Date
2026-06-26

Smart Images

  • Figure CN122290880A_ABST
    Figure CN122290880A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal fusion method for human motion representation and quality assessment, belonging to the field of human motion perception and assessment technology. It addresses the problems of large modal differences, insufficient sensor characteristic modeling, lack of physical constraints, and insufficient motion quality assessment indicators in existing traditional human motion representation and quality assessment methods. This invention collects multimodal data, preprocesses the data, and inputs the preprocessed multimodal data into a teacher-student network framework, outputting high-precision kinematic features and spatiotemporal relationships. A physical heuristic constraint loss is introduced to obtain a trained teacher-student network framework. Motion phase curves are extracted, and coordination and rhythmic stability indicators are calculated based on the phase curves, outputting the motion quality assessment results. This invention effectively improves the accuracy of human motion assessment and can be applied to sports rehabilitation training, sports movement assessment, ergonomics analysis, and intelligent human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human motion perception and assessment technology, specifically to a multimodal fusion method for human motion representation and quality assessment. Background Technology

[0002] With the accelerating aging of the population and the increasing awareness of health, the demand for sports rehabilitation and health management is growing. Statistics show that the number of people worldwide suffering from motor dysfunction due to neurological diseases, sports injuries, and age-related degenerative changes continues to expand, creating an urgent need for accurate and convenient methods of motor function assessment. At the same time, the pursuit of refined movement techniques in competitive sports and the widespread application of ergonomic analysis in intelligent manufacturing are further driving the development of quantitative assessment technologies for human movement.

[0003] Traditional human motion analysis primarily relies on optical motion capture systems, such as Vicon and OptiTrack. While offering high accuracy, these systems suffer from drawbacks such as high cost, limited space, and complex operation, hindering their widespread adoption in clinical rehabilitation and home training scenarios. In recent years, with the rapid development of microelectromechanical systems (MEMS) technology, flexible electronics, and computer vision algorithms, low-cost motion sensing solutions based on wearable and visual sensors have become a research hotspot. Among these, flexible stretch sensors, due to their excellent wearability, lightweight characteristics, and insensitivity to lighting conditions, have shown broad application prospects in the field of human motion monitoring. However, single-modal sensors cannot fully characterize the complexity of human motion. Flexible stretch sensors can accurately measure deformation information on the limb surface, but their signals are affected by nonlinear characteristics such as sensor hysteresis and creep, and they cannot directly obtain deep kinematic parameters such as joint angles. Visual sensors can provide rich spatial structural information, but are susceptible to interference from factors such as occlusion and changes in lighting. Inertial sensors can measure acceleration and angular velocity, but suffer from integral drift and noise accumulation problems. Therefore, integrating complementary information from multiple modal sensors to construct a comprehensive and accurate method for representing human motion has become an important development direction in this field.

[0004] Based on multimodal fusion, achieving quantitative assessment of exercise quality is the core objective at the application level. Whether it is functional recovery assessment in rehabilitation training or movement standardization judgment in physical education, it is necessary to establish a scientific, objective, and quantifiable evaluation index system. Existing exercise quality assessments mostly rely on expert experience or subjective scoring, lacking standardized and automated assessment methods. Therefore, researching exercise representation and quality assessment methods that can integrate multimodal sensor information, embed physical prior knowledge, and output interpretable indicators has important theoretical value and practical significance.

[0005] Existing methods for human motion characterization and quality assessment are mainly divided into three categories: optical motion capture-based methods, wearable sensor-based methods, and multimodal fusion-based methods. Each method has limitations to varying degrees, making it difficult to meet practical application needs. Specifically: 1) Optical motion capture-based methods: Optical motion capture systems use multiple high-speed cameras to capture reflective markers attached to key points on the human body, reconstructing three-dimensional motion trajectories. As the most accurate motion measurement method currently available, it is widely used in film and television special effects, sports biomechanics research, and other fields. However, its limitations are also quite obvious. First, the equipment cost is high; commercial optical motion capture systems are expensive, making them difficult to apply in clinical rehabilitation, home training, and other similar settings. Firstly, the system is complex to deploy, requiring specialized facilities and operators, and has strict requirements regarding ambient lighting and marker occlusion. Secondly, the process of marking and calibrating markers is cumbersome, requiring subjects to wear specific clothing, which affects their natural movement. Furthermore, this type of method can only provide location information and cannot directly measure physiological signals such as muscle activity and skin deformation, making it difficult to comprehensively characterize the movement state. 2) Methods based on wearable sensors: Wearable sensors mainly include two types: inertial measurement units (IMUs) and flexible stretch sensors. IMUs measure acceleration and angular velocity during movement using accelerometers and gyroscopes, and the attitude angle can be estimated through integration algorithms. This type of sensor has relatively low cost and is not affected by ambient lighting. While IMUs are suitable for outdoor and everyday applications, they suffer from significant integral drift, leading to accumulated attitude estimation errors over time. They are also sensitive to magnetic field interference. Flexible stretch sensors, which measure limb surface deformation to infer joint angles or muscle activity, offer advantages such as comfortable wear, lightweight design, and low power consumption. However, flexible sensors generally exhibit nonlinear characteristics such as hysteresis and creep, resulting in complex mappings between measurement results and actual motion states, directly affecting the accuracy of motion representation. Existing research lacks sufficient modeling of sensor characteristics, often employing simple linear correction or black-box fitting, which fails to fundamentally address signal distortion. 3) Multimodal fusion-based methods aim to compensate for the shortcomings of single-modal methods. To address the shortcomings of existing methods, researchers have attempted to fuse data from multiple modalities, including vision, IMU, and stretch sensors. Current fusion methods primarily employ feature stitching, decision-level weighting, or attention-based fusion strategies. These methods perform simple fusion at the data or feature layers, lacking in-depth modeling of the physical relationships between modalities. Firstly, the physical principles of different sensor modalities differ significantly, making direct fusion difficult to obtain a unified motion representation. Secondly, existing methods generally neglect modeling the dynamic characteristics of flexible sensors, such as hysteresis and creep, leading to deviations between the fusion results and real motion. Thirdly, deep learning methods struggle to embed human kinematics and biomechanical constraints during the fusion process, often resulting in outputs that do not conform to physical laws and lack interpretability.4) Deficiencies in exercise quality assessment methods: Existing methods for assessing exercise quality largely rely on manual scoring or simple statistical analysis of exercise parameters. While assessment tools widely used in clinical rehabilitation, such as the Fugl-Meyer and Berg balance scales, have good reliability and validity, they depend on therapist experience and suffer from high subjectivity, time-consuming and labor-intensive processes, and the inability to perform continuous monitoring. Automated assessment methods based on sensor data often extract statistical features such as mean, variance, and peak value, making it difficult to capture the dynamic coordination and rhythmic characteristics during exercise. The assessment results do not correlate well with clinical standards. Therefore, existing technologies still have significant shortcomings in multimodal fusion characterization, sensor characteristic modeling, physical constraint embedding, and quantitative assessment of exercise quality.

[0006] In summary, traditional technologies lack systematic solutions to existing problems, necessitating multimodal fusion methods for human motion representation and quality assessment. These methods aim to overcome the aforementioned technical bottlenecks and provide solutions for human motion representation and quality assessment. Summary of the Invention

[0007] A brief overview of the invention is given below to provide a basic understanding of certain aspects of it. It should be understood that this overview is not an exhaustive summary of the invention. It is not intended to identify key or essential parts of the invention, nor is it intended to limit the scope of the invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0008] In view of this, in order to solve the problems of large modal differences in multimodal fusion, insufficient sensor characteristic modeling, lack of physical constraints and lack of motion quality assessment indicators in the traditional human motion characterization and quality assessment methods in the prior art, this invention provides a multimodal fusion method for human motion characterization and quality assessment.

[0009] The technical solution is as follows: a multimodal fusion method for human motion representation and quality assessment, including the following steps:

[0010] S1. Multimodal data is synchronously collected through a unified clock source and hardware triggering mechanism to ensure the accurate correspondence of multimodal data on the time axis. The multimodal data is preprocessed to obtain preprocessed multimodal data.

[0011] Specifically: multimodal data includes three-dimensional skeletal data and signals from flexible stretching sensors;

[0012] S2. Construct a teacher-student network framework that includes a teacher network and a student network, input preprocessed multimodal data into it, and output high-precision kinematic features and spatiotemporal relationships;

[0013] S3. Physically inspired constraint loss is introduced into the high-precision kinematic features and spatiotemporal model, including topological consistency loss, waveform consistency loss and unidirectional motion constraint loss, to obtain the trained teacher-student network framework;

[0014] S4. Based on the high-precision kinematic features output by the student network, extract the motion phase curve, calculate the coordination index and rhythm stability index based on the phase curve, and finally output the motion quality assessment result by the trained teacher-student network framework.

[0015] Furthermore, step S2 includes the following steps:

[0016] S21. The teacher network extracts kinematic features based on preprocessed 3D skeleton data and generates a one-dimensional physical representation through kinematic projection constraints to obtain attention features after physical projection.

[0017] In step S21, the preprocessed 3D skeletal data is input into the teacher network, and the Euclidean distance between joints is calculated.

[0018] The Euclidean distance between joints is expressed as:

[0019]

[0020] in, For the first Human joints at all times With joints The Euclidean distance between them and Representing joints and joints Coordinate vectors in three-dimensional space Let L2 be the norm of the vector;

[0021] The rate of change of length is calculated based on the Euclidean distance between joints at time intervals:

[0022] The rate of change of length is expressed as:

[0023]

[0024] in, Indicates the first The change in inter-joint distance at a given moment relative to the previous moment. For the first Human joints at all times With joints The Euclidean distance between them;

[0025] Constructing a one-dimensional physical mask It is then applied to a high-dimensional feature tensor to obtain attention features after physical projection. ;

[0026] Attention characteristics after physical projection Represented as:

[0027]

[0028] in, The attentional characteristics after physical projection. This is a high-dimensional feature tensor extracted from 3D skeletal data. For a one-dimensional physical mask, This represents the element-wise multiplication operation;

[0029] S22. Student network spatiotemporal decoupling modeling based on flexible sensor signals.

[0030] In step S22, the preprocessed flexible sensor signal Input the student network and perform spatiotemporal decoupling modeling through the state-space model. The resulting spatiotemporal model models spatial and temporal dependencies through bidirectional scanning.

[0031] The spatiotemporal decoupling modeling process is represented as follows:

[0032]

[0033]

[0034] in, For the first The hidden state vector at time step 1. For the first The hidden state vector at time -1 For the first The input at any given time, i.e., the preprocessed flexible sensor signal or its characteristic representation. For the first The output at any given time, i.e., the motion characteristics reconstructed or predicted by the student network. is the learnable parameter matrix of the state-space model, which controls state transition, input mapping, state-to-output mapping, and direct jump connection, respectively.

[0035] Furthermore, step S3 includes the following steps:

[0036] S31. Use topological consistency to measure the differences in joint coordination between teacher networks and student networks;

[0037] In S31, topology consistency loss is used. The joint topology of the teacher network and the student network outputs must be consistent.

[0038] Topology consistency loss Represented as:

[0039]

[0040] in, Number of sensor channels Joints for Teachers' Network Output With joints The strength of the correlation between them The corresponding correlation strength output for the student network;

[0041] S32. Waveform consistency is used to measure the similarity of the output waveforms of the teacher network and the student network;

[0042] In S32, waveform consistency loss is used. The waveforms of the student network output are similar in shape to those of the teacher network output.

[0043] Waveform consistency loss Represented as:

[0044]

[0045] in, The Pearson correlation coefficient is used. The motion feature sequence output by the teacher network. The corresponding sequence output by the student network;

[0046] S33. Use unidirectional motion constraints to penalize situations where the direction of change of the flexible sensor signal is inconsistent with the direction of change of the physical length;

[0047] In S33, unidirectional motion constraint loss is used. Constrain the direction of motion to be consistent with the rate of change of physical length, and suppress the prediction of non-physical motion direction;

[0048]

[0049] in, The rate of change of physical length calculated by the teacher network. The change in the signal from the flexible sensor. To obtain The negative part, To obtain The positive part.

[0050] Furthermore, in S4, the coordination index is the time delay of the phase peak between symmetrical limbs, which reflects the degree of limb synchronization, and the rhythm stability index is the coefficient of variation of the movement cycle, which reflects the consistency and stability of the repetition of the movement.

[0051] The coordination index is expressed as follows:

[0052]

[0053] in, The absolute time difference of the phase peak of the left and right symmetrical limb movements. This refers to the time point when the left limb movement phase curve reaches its peak. The time point at which the right limb movement phase curve reaches its peak;

[0054] The rhythm stability index is expressed as:

[0055]

[0056] in, The coefficient of variation is 1. It is a time length sequence of adjacent motion cycles. The standard deviation of the periodic series, It is the mean of the periodic sequence.

[0057] The beneficial effects of this invention are as follows: The core technical challenges overcome by this invention include: 1) Bridging the physical differences between modalities: There are significant differences between three-dimensional skeletal data and flexible tensile sensor signals in terms of physical dimension, physical meaning, and sampling frequency. Skeletal data provides joint position information in three-dimensional space, reflecting deep skeletal movement; while flexible sensors measure skin surface deformation, which is affected by various factors such as muscle contraction, skin sliding, and sensor attachment state. How to establish a physical mapping relationship between these two modalities so that high-precision skeletal data can effectively guide the representation learning of flexible sensor signals is the primary problem solved by this invention; 2) Modeling the nonlinear characteristics of flexible sensors: Flexible tensile sensors exhibit obvious nonlinear characteristics such as hysteresis, creep, and rate correlation during operation. Hysteresis is manifested as inconsistent response paths during loading and unloading; creep is manifested as a slow change in resistance value over time under constant tension; rate correlation is manifested as the response characteristics being affected by the stretching rate. The above nonlinear characteristics make the sensor output and the actual deformation present a complex dynamic relationship, and simple linear correction is difficult to eliminate its influence. This invention designs a network structure that can effectively model nonlinear characteristics and improve the accuracy of motion representation; 3) Physical constraints Effective embedding of bundles: While pure data-driven deep learning methods have strong fitting capabilities, they often ignore the inherent physical laws of human movement, easily leading to unreasonable prediction results, such as sudden changes in joint angles and movement directions violating biomechanical constraints. This invention effectively embeds human kinematics and biomechanics knowledge into the network training process in the form of constraint loss, guiding the model to learn movement characteristics that conform to physical laws and ensuring the rationality and interpretability of the output results; 4) Construction of interpretable movement quality indicators: Movement quality assessment not only needs to output scores or grades, but also needs to provide physically meaningful and interpretable quantitative indicators so that the assessment results can guide rehabilitation training or technical improvement. This invention accurately extracts abstract concepts such as coordination and rhythm from movement time-series data and establishes a correspondence between them and clinical assessment standards, enabling movement quality assessment methods to move towards practical application; 5) Balance between low cost and high performance: This invention is geared towards wearable and deployable practical application scenarios, requiring that the inference stage rely solely on low-cost sensors (flexible stretch sensors) for input, while maintaining high representation accuracy and assessment accuracy. It fully utilizes the guiding role of high-precision skeletal data in the training stage and achieves independent deployment of low-cost sensors in the inference stage.

[0058] This invention aims to provide a multimodal fusion method for human motion representation and quality assessment. By constructing a teacher-student network framework and fusing high-precision 3D skeletal data with flexible stretch sensor signals, it achieves high-quality representation and quantitative assessment of human motion quality under the condition of embedded physical priors. To address the problem that a single modality cannot comprehensively characterize human motion, this paper constructs a multimodal dataset including Kinect (skeleton, depth, infrared) and flexible stretch sensors, and proposes a motion representation and assessment method that integrates physical priors. At the data level, a self-designed synchronous acquisition system achieves time alignment and unified annotation of multimodal data. The dataset contains 24 typical motion categories, providing a foundation for subsequent model training. Methodologically, the teacher network utilizes high-precision... Kinematic features are extracted from 3D skeletal data, and the changes in 3D joint distances are mapped to a one-dimensional physical representation consistent with the sensor data through a kinematic projection mechanism, thereby reducing intermodal differences. The student network, based on flexible sensor signals, uses a spatiotemporal decoupling structure to model human motion, where the spatial branch is used to capture the coordination relationship between joints, and the temporal branch is used to model the hysteresis and drift characteristics of the sensor. During training, various physical heuristic constraints are introduced, including topological consistency constraints, waveform consistency constraints, and unidirectional motion constraints, to guide the model to learn motion features that conform to actual physical meaning. Based on the motion phase curves output by the model, human coordination and rhythmicity indicators are further extracted, such as the time delay and periodic variation coefficient between symmetrical limbs, to achieve a quantitative assessment of motion quality.

[0059] Specifically, 1) Constructing a teacher-student network framework including a teacher network and a student network: Existing multimodal fusion methods mostly adopt simple fusion strategies such as feature splicing or decision layer weighting, ignoring the essential differences in physical dimensions between different modalities. This invention is the first to construct a teacher-student network framework. The teacher network extracts the Euclidean distance and length change rate between joints based on high-precision three-dimensional skeletal data, and constructs a one-dimensional physical mask through a kinematic projection mechanism to project the kinematic features in three-dimensional space onto a one-dimensional physical representation space consistent with the flexible sensor. This mechanism enables the teacher network to "teach" the student network to learn the mapping relationship from the flexible sensor signal to the real motion state, realizing the transfer of physical knowledge from high-precision modality to low-precision modality; 2) Proposing a spatiotemporal decoupled state-space modeling method to accurately characterize the hysteresis and creep nonlinear characteristics of the flexible sensor: Existing flexible sensor signal processing methods mostly adopt linear correction or simple filtering, which makes it difficult to accurately model the inherent nonlinear dynamic characteristics of the sensor, such as hysteresis and creep.This invention innovatively employs a state-space model to perform spatiotemporal decoupling modeling of the student network: spatial and temporal dependencies are modeled separately through a bidirectional scanning mechanism. The spatial branch utilizes a graph neural network to capture the cooperative motion relationships between joints, implicitly encoding human biomechanical constraints; the temporal branch uses state-space equations to model the dynamic response process of sensor outputs, effectively characterizing the asymmetry of hysteresis loops, inconsistent loading and unloading paths, and creep effects. 3) Multiple physics-inspired constraint losses are introduced to guide the model to learn physically reasonable representations that conform to the laws of human movement: Purely data-driven deep learning methods excel in fitting training data, but are prone to learning representations that do not conform to physical laws, such as abrupt changes in joint angles and movement directions that violate biological laws. Mechanical constraints, etc., severely affect the rationality and interpretability of the output results. This invention innovatively designs a triple physical heuristic constraint loss: topology consistency loss: constrains the joint topology of the teacher network and student network outputs to be consistent, ensuring that the joint coordination relationship learned by the student network conforms to the human anatomical structure; waveform consistency loss: constrains the output of the student network and the output of the teacher network to be similar in temporal form through waveform correlation coefficient, ensuring that the rhythm and phase information of motion are accurately preserved; unidirectional motion constraint loss: combines the physical rate of change and the direction of change of sensor signals to punish the prediction of motion direction that contradicts the physical law, avoiding non-physical results such as joint reverse movement; 4) proposes a quantification of coordination and rhythm based on motion phase. Evaluation Index System: Existing methods for assessing motor quality often rely on manual scoring or simple statistical characteristics (such as mean, variance, and peak value), making it difficult to capture the dynamic coordination and rhythmic characteristics during movement. The evaluation results are highly subjective and poorly interpretable. This invention innovatively extracts the motor phase curve based on the motor representation output by the student network. On this basis, two types of core quality assessment indicators are constructed: Coordination Index: This calculates the time delay of the phase peak between symmetrical limbs to quantify motor coordination. This index directly reflects the degree of synchronization between the left and right limbs during movement execution and has clinical significance for assessing motor coordination on the hemiplegic and non-hemiplegic sides of stroke patients. Rhythmic Stability Index: This calculates the coefficient of variation of the motor cycle to quantify the stability of the motor rhythm. This indicator reflects the consistency of repetitive movements and has a sensitive ability to identify gait rhythm abnormalities and movement variations caused by exercise fatigue in Parkinson's disease patients; 5) Achieving high-quality motion assessment under low-cost wearable conditions, with both theoretical innovation and engineering practical value: In terms of technical architecture, this invention achieves the design goal of "high-precision training and low-cost inference": During the training phase, high-precision three-dimensional skeletal data and physical constraints are used to guide the representation learning of flexible sensor signals; During the inference phase, only flexible stretching sensor input is required to output high-precision motion representation and quality assessment results. This design frees the system from dependence on expensive hardware such as optical motion capture systems and large computing devices, and achieves high-quality motion assessment under low-cost, wearable conditions.

[0060] In summary, this invention has achieved technological innovations in multiple dimensions, including physical knowledge transfer, sensor nonlinear modeling, physical constraint embedding, quantitative evaluation index construction, and low-cost deployment. It has formed a complete system of human motion characterization and quality assessment methods that combine theoretical depth and engineering practical value, filling the gaps in existing technologies in terms of multimodal fusion and physical interpretability. Attached Figure Description

[0061] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0062] Figure 1 This is a flowchart illustrating a multimodal fusion method for human motion representation and quality assessment. Detailed Implementation

[0063] To make the technical solutions and advantages of the embodiments of the present invention clearer, the exemplary embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0064] Example 1: Reference Figure 1 This embodiment describes a multimodal fusion method for human motion characterization and quality assessment, specifically including the following steps:

[0065] S1. Multimodal data is synchronously collected through a unified clock source and hardware triggering mechanism to ensure the accurate correspondence of multimodal data on the time axis. The multimodal data is preprocessed to obtain preprocessed multimodal data.

[0066] Specifically: multimodal data includes three-dimensional skeletal data and signals from flexible stretching sensors;

[0067] S2. Construct a teacher-student network framework that includes a teacher network and a student network, input preprocessed multimodal data into it, and output high-precision kinematic features and spatiotemporal relationships;

[0068] S3. Physically inspired constraint loss is introduced into the high-precision kinematic features and spatiotemporal model, including topological consistency loss, waveform consistency loss and unidirectional motion constraint loss, to obtain the trained teacher-student network framework;

[0069] S4. Based on the high-precision kinematic features output by the student network, extract the motion phase curve, calculate the coordination index and rhythm stability index based on the phase curve, and finally output the motion quality assessment result by the trained teacher-student network framework.

[0070] This invention achieves high-precision alignment of multimodal data through a self-designed synchronous acquisition system. The system structurally includes: a main control computer for sending synchronization trigger signals, a synchronization signal distribution / forwarding module (or branch circuit), a three-dimensional vision sensor (including depth and infrared modules), and a multi-channel flexible tensile sensor array;

[0071] In terms of connectivity, the main control computer sends a synchronization trigger signal to the synchronization signal distribution module via the digital I / O interface. This distribution module simultaneously routes the synchronization signal to the external trigger input of the 3D vision sensor and the trigger input of the flexible tensile sensor acquisition card, thereby ensuring that the two types of sensors start data acquisition in strict synchronization on the time axis. Each sensor and the acquisition card transmits the acquired data back to the main control computer via independent data transmission channels.

[0072] The signal flow is as follows: The main control computer sends a synchronization trigger pulse → The synchronization signal distribution module splits the pulse in two, simultaneously driving the 3D vision sensor to acquire the depth / infrared image of the current frame, and driving the flexible tensile sensor to acquire the corresponding strain / resistance change signal → After each sensor completes a single acquisition, it sends the original data with timestamps back to the main control computer through its respective data channel → In the preprocessing stage, the skeletal data is sequentially normalized for joint coordinates and filtered for noise reduction, and the flexible sensor signals are sequentially baseline corrected and outlier removed, finally outputting time-aligned and quality-enhanced multimodal data to provide high-quality input for subsequent fusion modeling.

[0073] Furthermore, step S2 includes the following steps:

[0074] S21. The teacher network extracts kinematic features based on preprocessed 3D skeleton data and generates a one-dimensional physical representation through kinematic projection constraints to obtain attention features after physical projection.

[0075] In step S21, the preprocessed 3D skeletal data is input into the teacher network, and the Euclidean distance between joints is calculated.

[0076] The Euclidean distance between joints is expressed as:

[0077]

[0078] in, For the first Human joints at all times With joints The Euclidean distance between them and Representing joints and joints Coordinate vectors in three-dimensional space Let L2 be the norm of the vector;

[0079] The rate of change of length is calculated based on the Euclidean distance between joints at time intervals:

[0080] The rate of change of length is expressed as:

[0081]

[0082] in, Indicates the first The change in inter-joint distance at a given moment relative to the previous moment. For the first Human joints at all times With joints The Euclidean distance between them;

[0083] Constructing a one-dimensional physical mask It is then applied to a high-dimensional feature tensor to obtain attention features after physical projection. ;

[0084] Attention characteristics after physical projection Represented as:

[0085]

[0086] in, These are attention features after physical projection, which are used for subsequent network processing. This is a high-dimensional feature tensor extracted from 3D skeletal data. A one-dimensional physical mask used to map three-dimensional features to a dimension consistent with the sensor signal. This indicates the element-wise multiplication (Hadamard product) operation;

[0087] S22. Student network spatiotemporal decoupling modeling based on flexible sensor signals.

[0088] In step S22, the preprocessed flexible sensor signal Input the student network and perform spatiotemporal decoupling modeling through the state space model. The resulting spatiotemporal model models spatial dependence (inter-joint coordination) and temporal dependence (sensor hysteresis and drift) through bidirectional scanning.

[0089] The spatiotemporal decoupling modeling process is represented as follows:

[0090]

[0091]

[0092] in, For the first The hidden state vector at time t, which encodes the dynamic information of the flexible sensor signal, For the first The hidden state vector at time -1 For the first The input at any given time, i.e., the preprocessed flexible sensor signal or its characteristic representation. For the first The output at any given time, i.e., the motion characteristics reconstructed or predicted by the student network. is the learnable parameter matrix of the state-space model, which controls state transition, input mapping, state-to-output mapping, and direct jump connection, respectively.

[0093] Specifically, this invention innovatively constructs a teacher-student network architecture to achieve knowledge transfer from high-precision modalities to low-precision modalities. The teacher network, based on 3D skeletal data, uses a graph convolutional network to extract Euclidean distances and length change rates between joints, outputting high-precision kinematic features. The student network takes flexible sensor signals as input and uses a state-space model for spatiotemporal decoupling modeling. The teacher network guides the student network training, enabling the student network to output high-precision motion representations during the inference phase by relying solely on low-cost flexible sensors.

[0094] The teacher network achieves physical mapping from three-dimensional space to sensor perception space through kinematic projection mechanism, calculates Euclidean distance and its rate of change between joints, constructs a one-dimensional physical mask, and projects three-dimensional features onto a one-dimensional representation space consistent with the physical meaning of the flexible sensor. This effectively reduces the difference in physical dimensions between modalities, enabling the teacher network to effectively transfer three-dimensional kinematic knowledge to the student network.

[0095] The kinematic projection mechanism includes calculating the Euclidean distance and length change rate between joints, constructing a one-dimensional physical mask, and multiplying the three-dimensional features with the one-dimensional mask element by element. Step S2 realizes the physical projection from three-dimensional space to the sensor perception space, reducing the difference in physical dimensions between modes.

[0096] Furthermore, step S3 includes the following steps:

[0097] S31. Use topological consistency to measure the differences in joint coordination between teacher networks and student networks;

[0098] In S31, topology consistency loss is used. The joint topology of the teacher network and the student network outputs must be consistent.

[0099] Topology consistency loss Represented as:

[0100]

[0101] in, Number of sensor channels Joints for Teachers' Network Output With joints The strength of the association between them (such as correlation or attention weight). The corresponding correlation strength output for the student network;

[0102] S32. Waveform consistency is used to measure the similarity of the output waveforms of the teacher network and the student network;

[0103] In S32, waveform consistency loss is used. The waveforms of the student network output are similar in shape to those of the teacher network output.

[0104] Waveform consistency loss Represented as:

[0105]

[0106] in, The smaller the value, the more similar the two waveforms are. The Pearson correlation coefficient is used. The sequence of motion features (such as joint distance change curves) output by the teacher network. The corresponding sequence output by the student network;

[0107] S33. Use unidirectional motion constraints to penalize situations where the direction of change of the flexible sensor signal is inconsistent with the direction of change of the physical length;

[0108] In S33, unidirectional motion constraint loss is used. Constrain the direction of motion to be consistent with the rate of change of physical length, and suppress the prediction of non-physical motion direction;

[0109]

[0110] in, The rate of change of physical length (which can be positive or negative) is calculated by the teacher's network. This represents the change in the signal of the flexible sensor (which can be positive or negative). To obtain The negative part (non-zero only when the physical length is shortened). To obtain The positive part (non-zero only when the flexible sensor signal increases) is positive only when the physical length decreases and the sensor signal increases, at which point the loss is greater than 0, forcing the model to avoid such predictions that violate physical common sense.

[0111] Specifically, the teacher-student network adopts a two-stage training strategy. In the first stage, the teacher network is pre-trained using three-dimensional skeleton data. In the second stage, the teacher network parameters are fixed, and the student network is trained by jointly optimizing the physical heuristic constraint loss and the student network's self-supervised loss to achieve knowledge transfer.

[0112] The spatiotemporal decoupling modeling adopts a state-space model, which models spatial dependence and temporal dependence respectively through a bidirectional scanning mechanism. Spatial dependence uses graph neural networks to capture the cooperative relationship between joints, while temporal dependence uses state-space equations to model the hysteresis and creep nonlinear characteristics of the sensor.

[0113] Furthermore, in S4, the coordination index is the time delay of the phase peak between symmetrical limbs, which reflects the degree of limb synchronization, and the rhythm stability index is the coefficient of variation of the movement cycle, which reflects the consistency and stability of the repetition of the movement.

[0114] The coordination index is expressed as follows:

[0115]

[0116] in, This refers to the absolute time difference between the peak phases of movements of symmetrical limbs (such as the left and right arms or legs). It reflects the coordination of movements; the smaller the difference, the better the coordination. This refers to the time point when the left limb movement phase curve reaches its peak. The time point at which the right limb movement phase curve reaches its peak;

[0117] The rhythm stability index is expressed as:

[0118]

[0119] in, The coefficient of variation measures the stability of the movement cycle; the smaller the value, the more stable the rhythm. It is a time length sequence of adjacent movement cycles (e.g., time values ​​of multiple gait cycles). The standard deviation of the periodic series, It is the mean of the periodic sequence.

[0120] Specifically, this invention can effectively reflect human motion characteristics under low-cost sensing conditions and has good application potential. In the inference stage, this invention only requires the input of flexible stretch sensor signals and does not require three-dimensional skeletal data, so as to realize high-quality motion representation and quality assessment under low-cost wearable conditions. This invention can be deployed in sports rehabilitation training systems, sports movement assessment systems, ergonomic analysis systems or intelligent human-computer interaction systems, and supports real-time motion feedback and quality assessment.

[0121] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the invention is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.

Claims

1. A multimodal fusion method for human motion representation and quality assessment, characterized in that, Includes the following steps: S1. Multimodal data is synchronously collected through a unified clock source and hardware triggering mechanism to ensure the accurate correspondence of multimodal data on the time axis. The multimodal data is preprocessed to obtain preprocessed multimodal data. Specifically: multimodal data includes three-dimensional skeletal data and signals from flexible stretching sensors; S2. Construct a teacher-student network framework that includes a teacher network and a student network, input preprocessed multimodal data into it, and output high-precision kinematic features and spatiotemporal relationships; S3. Physically inspired constraint loss is introduced into the high-precision kinematic features and spatiotemporal model, including topological consistency loss, waveform consistency loss and unidirectional motion constraint loss, to obtain the trained teacher-student network framework; S4. Based on the high-precision kinematic features output by the student network, extract the motion phase curve, calculate the coordination index and rhythm stability index based on the phase curve, and finally output the motion quality assessment result by the trained teacher-student network framework.

2. The multimodal fusion method for human motion characterization and quality assessment according to claim 1, characterized in that, S2 includes the following steps: S21. The teacher network extracts kinematic features based on preprocessed 3D skeleton data and generates a one-dimensional physical representation through kinematic projection constraints to obtain attention features after physical projection. In step S21, the preprocessed 3D skeletal data is input into the teacher network, and the Euclidean distance between joints is calculated. The Euclidean distance between joints is expressed as: in, For the first Human joints at all times With joints The Euclidean distance between them and Representing joints and joints Coordinate vectors in three-dimensional space Let L2 be the norm of the vector; The rate of change of length is calculated based on the Euclidean distance between joints at time intervals: The rate of change of length is expressed as: in, Indicates the first The change in inter-joint distance at a given moment relative to the previous moment. For the first Human joints at all times With joints The Euclidean distance between them; Constructing a one-dimensional physical mask It is then applied to a high-dimensional feature tensor to obtain attention features after physical projection. ; Attention characteristics after physical projection Represented as: in, The attentional characteristics after physical projection. This is a high-dimensional feature tensor extracted from 3D skeletal data. For a one-dimensional physical mask, This represents the element-wise multiplication operation; S22. Student network spatiotemporal decoupling modeling based on flexible sensor signals. In step S22, the preprocessed flexible sensor signal Input the student network and perform spatiotemporal decoupling modeling through the state-space model. The resulting spatiotemporal model models spatial and temporal dependencies through bidirectional scanning. The spatiotemporal decoupling modeling process is represented as follows: in, For the first The hidden state vector at time step 1. For the first The hidden state vector at time -1 For the first The input at any given time, i.e., the preprocessed flexible sensor signal or its characteristic representation. For the first The output at any given time, i.e., the motion characteristics reconstructed or predicted by the student network. is the learnable parameter matrix of the state-space model, which controls state transition, input mapping, state-to-output mapping, and direct jump connection, respectively.

3. The multimodal fusion method for human motion characterization and quality assessment according to claim 2, characterized in that, S3 includes the following steps: S31. Use topological consistency to measure the differences in joint coordination between teacher networks and student networks; In S31, topology consistency loss is used. The joint topology of the teacher network and the student network outputs must be consistent. Topology consistency loss Represented as: in, Number of sensor channels Joints for Teachers' Network Output With joints The strength of the correlation between them The corresponding correlation strength output for the student network; S32. Waveform consistency is used to measure the similarity of the output waveforms of the teacher network and the student network; In S32, waveform consistency loss is used. The waveforms of the student network output are similar in shape to those of the teacher network output. Waveform consistency loss Represented as: in, The Pearson correlation coefficient is used. The motion feature sequence output by the teacher network. The corresponding sequence output by the student network; S33. Use unidirectional motion constraints to penalize situations where the direction of change of the flexible sensor signal is inconsistent with the direction of change of the physical length; In S33, unidirectional motion constraint loss is used. Constrain the direction of motion to be consistent with the rate of change of physical length, and suppress the prediction of non-physical motion direction; in, The rate of change of physical length calculated by the teacher network. The change in the signal from the flexible sensor. To obtain The negative part, To obtain The positive part.

4. The multimodal fusion method for human motion characterization and quality assessment according to claim 3, characterized in that, In S4, the coordination index is the time delay of the phase peak between symmetrical limbs, which reflects the degree of limb synchronization, and the rhythm stability index is the coefficient of variation of the movement cycle, which reflects the consistency and stability of the repetition of the movement. The coordination index is expressed as follows: in, The absolute time difference of the phase peak of the left and right symmetrical limb movements. This refers to the time point when the left limb movement phase curve reaches its peak. The time point at which the right limb movement phase curve reaches its peak; The rhythm stability index is expressed as: in, The coefficient of variation is 1. It is a time length sequence of adjacent motion cycles. The standard deviation of the periodic series, It is the mean of the periodic sequence.