Multi-stage collaboration based multi-modal cross-subject rapid adaptive assessment method and system
Patent Information
- Application Number
- CN202611151463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-29
AI Technical Summary
然而,脑卒中上肢康复评估中的sEMG偏移更接近协方差结构畸变,IMU偏移更接近刚体旋转引起的几何偏移,EEG偏移则呈现高维非线性漂移特征,若采用统一对齐方式,容易发生过度修正并破坏模态间原有生理拓扑关系
[0158]与现有技术相比,本发明的显著进步在于:1)第一,本发明通过多阶段协同框架将无标签表征预训练、源域双主任务监督学习、目标域差异化适配和少样本个体化校准串联为完整闭环,可充分利用临床历史无标签多模态数据,并降低新患者部署时对大量人工量表标注的依赖;2)本发明依据EEG、sEMG和IMU各自的跨患者偏移机理分别设计适配路径,并通过跨模态图结构保持约束维持脑肌耦合拓扑,能够在提高跨患者泛化能力的同时避免统一全局对齐造成的生理结构失真;3)本发明将FMA评分与MBI量表评分设为双主任务输出,并以代偿识别、肌力分级和CMCI等辅助任务增强解释能力,使评估结果同时覆盖运动控制恢复、生活功能结局和代偿模式识别,较单一量表或单一传感器评估更完整;4)本发明通过MAML-Rehab少样本快速适配机制,使目标患者仅需少量标准化动作样本即可完成个体化校准,减少临床重标定时间和患者负担,适用于脑卒中康复评估、疗效随访、过程监测和个体化康复建议生成等真实部署场景。
Smart Images

Figure CN122842944A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical health and intelligent rehabilitation assessment technology, and in particular relates to a multimodal cross-individual rapid adaptive modeling method and system for assessing limb motor function in stroke patients. Specifically, it includes collaborative representation of EEG, sEMG, and IMU multimodal signals, differential domain adaptation implemented according to modal shift mechanism, preservation of cross-modal physiological structure, and individualized calibration for small sample patients. Background Technology
[0002] Stroke is a leading cause of long-term disability in adults worldwide. Upper limb motor dysfunction is present in more than half of stroke survivors, profoundly impacting their Activities of Daily Living (ADL). Assessment of limb motor function in stroke patients is crucial for developing rehabilitation plans, monitoring treatment effectiveness, and predicting prognosis. The Fugl-Meyer Assessment Upper Extremity (FMA-UE), commonly used in clinical practice, characterizes upper limb motor impairment and recovery levels, while the Modified Barthel Index (MBI) characterizes ADL and functional independence. While related, they are not equivalent; the former leans more towards motor control, while the latter focuses more on functional outcomes. Therefore, the output of a single scale is insufficient to fully support clinical decision-making.
[0003] Traditional manual assessment relies on therapist observation and scale scoring, which suffers from problems such as being time-consuming, highly subjective, having repeatability greatly influenced by the assessor's experience, and being difficult to quantify continuously at high frequencies. In recent years, sensors such as IMU, sEMG, and EEG have been used for motor assessment and rehabilitation training in stroke, providing objective information such as kinematics, muscle activation, and cortical activity. However, existing solutions still face challenges in real-world clinical deployment, including poor cross-patient generalization, high annotation costs, large requirements for calibration samples from target patients, susceptibility of cross-modal structures to disruption by standardized alignment strategies, and incomplete assessment dimensions.
[0004] In existing technologies, there are already schemes for automatic assessment of upper limb function based on multi-sensor fusion. For example, CN109875565A discloses a method for automatic assessment of upper limb motor function in stroke based on deep learning. It mainly uses IMU and 7-channel sEMG as inputs and extracts motor features through convolutional neural networks to achieve Brunnstrom staging assessment. Such schemes demonstrate that multi-source motor signals have a realistic basis for automatic assessment of limb function in stroke. However, these methods mainly focus on the assessment mode of "direct prediction after training," lacking a dedicated mechanism to handle cross-individual distribution bias when deploying to new patients. They also do not introduce EEG modalities to characterize cortical motor-related information, and do not involve unsupervised pre-training and FMA+MBI dual-main-task joint output mechanism.
[0005] Existing technologies also include rehabilitation assessment methods based on the coupling relationship between EEG and EMG. For example, CN112494054A discloses a method for lower limb motor rehabilitation assessment after stroke based on the coherence of multichannel EMG and EEG, which focuses on assessing rehabilitation status through EEG and EMG preprocessing, coherence analysis, and muscle synergistic decomposition. These methods demonstrate that EEG and EMG coupling indicators can serve as important biomarkers for rehabilitation assessment; however, they primarily focus on lower limb rehabilitation scenarios and coherence analysis indicators, without addressing the kinematic supplementary information of IMU modalities, and without establishing a deep learning inference and cross-patient rapid calibration mechanism for multi-task assessment of upper limb injuries after stroke. In addition, CN119655773A relates to the use of electromyography signals for upper limb rehabilitation assessment in stroke patients, and CN116966056A relates to an upper limb rehabilitation assessment system for stroke patients. However, none of the above solutions have achieved a complete technical combination of EEG+sEMG+IMU trimodal synergy with ST-HGNN backbone, CDDA differential domain adaptation, graph structure preservation, MAML-Rehab, and FMA+MBI dual master task output.
[0006] Furthermore, publicly available technologies also include approaches such as quantifying motor function using single-modal EEG or handling cross-session EEG recognition using domain adaptation. This indicates that intelligent analysis of physiological signals and domain adaptation methods are not entirely unexplored areas. However, related solutions typically only address single-modal or recognition tasks and have not yet resolved the heterogeneous bias problem that exists simultaneously across EEG, sEMG, and IMU modalities in stroke upper limb rehabilitation assessment across different patient scenarios. Specifically, in the actual deployment of stroke rehabilitation assessment, the same pre-trained model can achieve good results on known patient data, but when the model is transferred to new, unseen patients, significant performance degradation often occurs due to factors such as differences in electrode impedance, subcutaneous fat thickness, muscle coordination patterns, individual variations in scalp rhythms, and IMU wearing posture deviations. This leads to increased FMA-UE regression error, confusion in compensatory movement classification boundaries, and decreased stability of cognitive-motor coupling indices.
[0007] Furthermore, traditional domain adaptation methods typically assume that the shifts in different modalities have similar statistical properties, thus tending to use a uniform global alignment strategy. However, sEMG shifts in upper limb rehabilitation assessment after stroke are closer to covariance structure distortions, IMU shifts are closer to geometric shifts caused by rigid body rotation, and EEG shifts exhibit high-dimensional nonlinear drift characteristics. If a uniform alignment method is used, overcorrection can easily occur, disrupting the original physiological topological relationships between modalities. General algorithms such as Deep CORAL, DANN, MK-MMD, and MAML have disclosed algorithmic ideas such as second-order statistical alignment, gradient inversion domain adversarial, multi-kernel mean difference, and fast adaptation with few samples. However, these algorithms are general machine learning methods and have not made differentiated combinations for the three-modal shift mechanisms in upper limb rehabilitation after stroke, nor have they been combined with cross-modal physiological structure maintenance and FMA+MBI clinical dual-master task output.
[0008] Furthermore, in clinical settings, target patients often cannot tolerate repeated acquisition of a large number of standardized actions. Traditional retraining or recalibration methods often require 20 to 30 action samples, which significantly increases assessment time and patient burden, making it difficult to meet the requirements of low cost, speed, and high stability for real-world clinical deployment. At the same time, a large amount of historical EEG, sEMG, and IMU data is in a state of "signal with few labels" or "signal with no labels." Most existing supervised learning methods cannot fully utilize this unlabeled data, resulting in limited initial model representation quality and poor starting point for small sample fit.
[0009] Therefore, there is an urgent need for a new limb rehabilitation assessment technology for stroke patients that can fully utilize unlabeled historical multimodal data for pre-training, implement differentiated domain adaptation for the shift mechanisms of different signal modalities, and complete rapid individualized calibration for new patients with a very small number of labeled samples. At the same time, it can meet the complete clinical assessment needs by jointly outputting the FMA score and MBI scale score as dual main tasks, thereby improving the assessment accuracy and clinical deployability while maintaining the stability of the cross-modal coupling structure. Summary of the Invention
[0010] The purpose of this invention is to address the problems mentioned in the background art by proposing a multimodal cross-individual rapid adaptive assessment method and system based on multi-stage collaboration. Its core lies in integrating multimodal unlabeled representation learning, clinical dual-master task supervised modeling, differential domain adaptation designed according to modality shift mechanism, cross-modal physiological structure preservation, and small-sample individualized calibration tissue as a continuous and collaborative technical link, enabling the model to migrate from population data to new patients and quickly generate patient-specific assessment results.
[0011] To achieve the objectives of this invention, a multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration is disclosed, comprising the following steps:
[0012] Step S1: Acquire and preprocess multimodal data of EEG, sEMG, and IMU signals; the patient performs a standardized upper limb movement task under a unified assessment process, and EEG, sEMG, and inertial motion signals are acquired simultaneously, and filtering, artifact suppression, rectification, normalization, posture initialization, time synchronization, and windowing segmentation are completed.
[0013] Step S2: Unsupervised pre-training is performed using unlabeled historical multimodal data; through intramodal reconstruction, cross-modal contrast consistency and temporal structure learning tasks, the shared encoder learns general neuro-muscle-motor representations without human scale labels.
[0014] Step S3: Construct a multimodal spatiotemporal heterogeneous graph neural network ST-HGNN;
[0015] Step S4: Perform supervised training of FMA score and MBI scale score on labeled source domain data, and simultaneously train compensation recognition, muscle strength classification and CMCI auxiliary tasks.
[0016] Step S5: Perform covariance alignment and off-diagonal preservation constraints on the sEMG channels to handle second-order statistical structure drift caused by differences in electrode location, subcutaneous fat thickness, and muscle synergy patterns.
[0017] Step S6: First, perform a gravity-direction-based geometric rotation calibration on the IMU channel, and then perform a lightweight adversarial domain adaptation to handle wearing posture deviation and residual statistical offset.
[0018] Step S7: Perform frequency band weighted multi-core maximum mean difference alignment on the EEG channel to handle high-dimensional nonlinear offsets caused by scalp impedance, individual variations in cortical rhythms, and non-steady-state drift.
[0019] Step S8: Introduce cross-modal graph Laplacian structure preservation constraints to prevent independent adaptation from disrupting the EEG-sEMG brain-muscle coupling topology and PLV-related phase synchronization structure;
[0020] Step S9: Use MAML-Rehab to perform rapid individualized adaptation of target patients with a small number of samples. Define a single patient as a meta-task and use a 5-shot support set to complete a small number of gradient step updates.
[0021] Step S10: Output patient-specific assessment results, including FMA score, MBI scale score, compensatory movement category, muscle strength grade, and optional clinical interpretation information.
[0022] Furthermore, in step S1, a source domain sample set and a target domain sample set are constructed for the rehabilitation assessment scenario of stroke patients. The source domain sample set consists of fully calibrated stroke patients and is used to train the general model for the population. The target domain sample set corresponds to newly enrolled patients and is used to simulate the cross-patient generalization and rapid individualized adaptation process of the model in real clinical deployment scenarios.
[0023] The source and target domain sample sets are represented as follows:
[0024]
[0025] in, This represents the source domain patient dataset, where the subscript 's' indicates 'source', i.e., the source domain. This represents the target domain patient dataset, where the subscript t indicates target, i.e., the target domain; This represents the multimodal input of the i-th sample in the source domain. This represents the clinical assessment label corresponding to the i-th sample in the source domain; This represents the multimodal input of the j-th sample in the target domain. This represents the clinical assessment label corresponding to the j-th sample in the target domain; the superscripts s and t represent the source domain and target domain respectively, and the subscripts i and j represent the sample indices in the source domain and target domain respectively; x is the lowercase input sample vector, y is the lowercase label vector; D is the uppercase dataset symbol; n s and n t These represent the number of samples in the source domain and the target domain, respectively.
[0026] Multimodal joint input is defined as:
[0027]
[0028] in, Indicates EEG characteristics, Indicates sEMG features, Indicate IMU features;
[0029] After obtaining the raw multimodal time series data, targeted preprocessing was performed on the three types of signals: EEG, sEMG, and IMU. Furthermore, cross-modal time synchronization, sample windowing, and input format unification were completed to form the standardized input for the subsequent multimodal heterogeneous graph model.
[0030] For EEG signals, bandpass filtering and power frequency interference suppression are first performed to reduce low-frequency drift and power supply noise; then artifact suppression and baseline correction are performed to reduce the impact of eye movement, electromyography crosstalk and electrode contact instability; then the EEG signals are segmented according to the task time window, and frequency band related features are extracted according to the predetermined frequency band, or the segmented time segments are directly used as input to the depth model.
[0031] The features after EEG preprocessing are represented as follows: :
[0032]
[0033] in, Represents raw brain electrical signals. This indicates filtering, artifact suppression, and baseline correction operations. This indicates a time window segmentation operation. This represents a brainwave frequency band feature mapping or deep coding mapping;
[0034] For sEMG signals, bandpass filtering and power frequency notch filtering are first performed, followed by rectification, envelope extraction and normalization, and time-series segmentation is performed according to a fixed sliding window or fixed event window aligned with the action task.
[0035] The sEMG feature construction process is as follows :
[0036]
[0037] in, Indicates primitive muscle, This indicates bandpass and notch filtering. Indicates rectification or envelope extraction. Indicates normalization, This indicates the division of the time window. This represents electromyographic feature mapping or deep coding mapping;
[0038] For IMU signals, zero-bias calibration and sampling synchronization are performed first. Then, the gravity direction is estimated based on the static sitting posture stage or other reference motion stage to obtain the rotation relationship of the current coordinate system of each IMU unit relative to the human anatomical coordinate system. Next, the motion time sequence signal is uniformly segmented and normalized. Static gravity reference information is preserved in the preprocessing stage.
[0039] IMU input is represented as :
[0040]
[0041] in, Represents the original IMU signal. This indicates the attitude initialization, zero-bias calibration, and coordinate adjustment operations. This indicates the division of the time window. This represents the IMU feature map or deep coding map;
[0042] The mean acceleration vector within the target patient's stationary segment is defined as :
[0043]
[0044] in, This represents the IMU acceleration vector within a stationary segment. Indicates the length of a static segment. Used for subsequent IMU gravity orientation geometry calibration; t is the time index within the static segment;
[0045] After completing the single-modal preprocessing, cross-modal time synchronization and sample windowing are further performed. Using the unified action start time, trigger event, or predetermined time anchor point as the alignment benchmark, EEG, sEMG, and IMU are resampled or indexed and aligned in a unified time axis so that the three types of signals have a one-to-one corresponding time segment within the same action window. Subsequently, the multimodal signals within the corresponding time window are combined into the same sample unit and action labels, scale scores, or other evaluation labels are attached to it.
[0046] No. A joint sample is represented as :
[0047]
[0048] in, and Other means the first The start and end positions of the time window for each action sample.
[0049] Furthermore, in step S2, a multimodal unsupervised pre-training stage is added before supervised learning to learn stable shared encoder parameters using clinical history unlabeled EEG, sEMG, and IMU data;
[0050] Unsupervised pre-training includes intra-modal reconstruction tasks, cross-modal contrastive consistency tasks, and temporal structure learning tasks;
[0051] Intramodal reconstruction is used to recover the original signal structure from the perturbed modal representation, enabling the encoder to learn the temporal patterns and local structural information of each mode. The loss is defined as... :
[0052]
[0053] Where m represents the mode number. Indicates the first The original input or feature representation of a modality. This indicates the reconstruction result of the corresponding mode; Represents the square of the L2 norm;
[0054] The cross-modal contrast consistency task aims to enhance the consistency between different modal representations within the same time window, making EEG, sEMG, and IMU features of the same action segment close to each other in a shared representation space, while distinguishing non-corresponding segments from different actions, different time windows, or different patients. The loss is defined as... :
[0055]
[0056] in, Denotes the set of positive sample pairs. Representation and Sample The corresponding candidate sample set, , , These represent the encoded sample representations, Represents the similarity function. Indicates the temperature coefficient; , , Used to distinguish different sample representations; exp is an exponential function;
[0057] The temporal structure learning task is used to preserve the sequential relationships between actions during execution, enabling the model to learn the dynamic changes between the initiation, execution, peak, and recovery phases of an action. The loss is defined as... :
[0058]
[0059] in, This represents a time-series relationship prediction function. and Encoded representation of adjacent time segments, A marker indicating the target temporal relationship between adjacent time segments; Represents the square of the L2 norm;
[0060] The total loss of multimodal unsupervised pre-training is :
[0061]
[0062]
[0063] in, , , These are the weight coefficients for intra-modal reconstruction loss, cross-modal contrast consistency loss, and temporal structure learning loss, respectively.
[0064] By minimizing It can learn a stable multimodal shared encoder; the pre-trained encoder parameters are used as initialization parameters for subsequent dual-master task supervised learning of FMA score and MBI scale score, CDDA differential domain adaptation and MAML-Rehab few-shot fast adaptation.
[0065] Furthermore, in step S3, a multimodal spatiotemporal heterogeneous graph neural network ST-HGNN is used as a shared backbone model to uniformly model the intramodal dependencies and cross-modal coupling relationships among EEG, sEMG, and IMU, and output deep feature representations suitable for rehabilitation assessment and domain adaptation.
[0066] The node set of the backbone model consists of three types of heterogeneous nodes: EEG node set, sEMG node set, and IMU node set. Among them, EEG nodes correspond to dual-channel EEG acquisition positions, sEMG nodes correspond to 14 surface electromyography channels, and IMU nodes correspond to 3 inertial measurement units. Thus, the entire heterogeneous map contains a total of 19 nodes, which can simultaneously carry information from three levels: cortex, muscle group, and limb segment.
[0067] In terms of edge relationship construction, both intramodal edges and intermodal edges are established simultaneously. Intramodal edges are used to represent physiological proximity, temporal continuity, or statistical correlation within the same modality. Intermodal edges are used to represent functional connectivity between different modalities. Through the graph structure, the backbone model can jointly encode cortical activity, muscle activation patterns, and action execution results within the same framework.
[0068] The node is initially represented as:
[0069]
[0070] in, These represent the initial feature representations of the EEG, sEMG, and IMU nodes, respectively.
[0071] The graph convolution or message passing update process is as follows:
[0072]
[0073] in, Represents the normalized adjacency matrix. Indicates the first Layer node representation, This represents the learnable weight matrix of the l-th layer. This represents a non-linear activation function.
[0074] Furthermore, in step S4, FMA score prediction and MBI scale score prediction are set as parallel dual main tasks, and compensation identification, muscle strength grading and CMCI are set as auxiliary tasks.
[0075] Assume the shared encoder output is The output of the dual main task and auxiliary task is then represented as follows:
[0076]
[0077] in, This indicates the FMA prediction result. This indicates the MBI prediction result. This indicates the result of the compensation action recognition. This indicates the results of the muscle strength grading prediction. This represents the prediction results of the cognitive-motor coupling index.
[0078] The corresponding multi-task prediction head is:
[0079]
[0080]
[0081] in, , , , and These represent the FMA score, MBI scale score, compensatory movement recognition, muscle strength grading, and CMCI prediction head, respectively.
[0082] A joint loss mechanism is adopted, prioritizing the primary task and enhancing the secondary task. :
[0083]
[0084]
[0085] in, and These represent the prediction losses of FMA and MBI, respectively, using mean squared error, mean absolute error, or Huber loss. The loss for the compensatory action classification is represented by cross-entropy loss; Indicates muscle strength grade loss; This represents the predictive loss of the cognitive-motor coupling index. This indicates the cross-modal graph structure preservation loss; , , , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade.
[0086] Furthermore, in steps S5, S6, and S7, a channel-specific differentiated domain adaptation framework (CDDA) is proposed, and adaptation paths are designed according to the different offset mechanisms of EEG, sEMG, and IMU modalities in cross-patient scenarios.
[0087] For sEMG modalities, covariance alignment constraints are applied to the deep sEMG features of the source and target domains to make their second-order statistics consistent; at the same time, off-diagonal preservation constraints are introduced to maintain the channel-related structures corresponding to the muscle co-activation mode.
[0088] The deep sEMG features of the source and target domains are respectively , :
[0089]
[0090] in, and These represent the number of samples in the source and target domains, respectively. Indicates the dimension of deep features in sEMG;
[0091] The source domain sEMG eigencovariance matrix is estimated as follows: :
[0092]
[0093] The target domain sEMG feature covariance matrix is estimated as follows: :
[0094]
[0095] Covariance alignment loss is defined as :
[0096]
[0097] in, Let T denote the Frobenius norm, and T denote the transpose. Represents a unit vector;
[0098] To maintain muscle synergy, off-diagonal retention constraints are introduced. :
[0099]
[0100] in, The off-diagonal part of the target domain sEMG covariance matrix is represented. This represents the off-diagonal portion of the reference covariance matrix. To constrain the weights, the values range from [0.1, 10]. The values are determined by optimizing the comprehensive index of muscle synergistic structure similarity and downstream task performance on the validation set through grid search, including FMA score, MBI scale score of compensatory movement category, and muscle strength grade; the off-diagonal elements reflect the correlation between different electromyographic channels, so this constraint helps to maintain the muscle synergistic activation mode.
[0101] The total adaptation loss of the sEMG branch is :
[0102]
[0103] For IMU modes, the rotational alignment relationship between the IMU coordinate system and the human anatomical coordinate system is obtained based on the static gravity direction, and then lightweight adversarial adaptation is performed on the residual statistical offset after geometric calibration.
[0104] Let the direction of gravity estimated from the static segment of the target patient be... The source region reference gravity direction is The optimal rotation matrix is then solved as follows: :
[0105]
[0106] in, Represents a set of three-dimensional rotation matrices; Represents the square of the L2 norm;
[0107] The calibrated target domain IMU acceleration or characteristic is represented as :
[0108]
[0109] in, This represents the original target domain IMU acceleration vector or feature vector. This represents the IMU representation after geometric calibration;
[0110] After geometric calibration, a domain discriminator D and a feature extractor G are introduced, and a lightweight adversarial loss is used to handle the residual statistical offset. :
[0111]
[0112] For the EEG modality, the maximum mean difference of multiple kernels was used to align the spatial distribution of the kernels, and different kernel components were weighted according to the importance of rehabilitation assessment indicators in different frequency bands; D s and D t These represent the source and target domain datasets, respectively. / This indicates that sample x originates from D.s / D t Distribution; E[ϕ(x)] represents the average feature of the data after mapping to a high-dimensional space;
[0113] MMD distance is defined as:
[0114]
[0115] in, The source domain D is represented in the reproducing kernel Hilbert space corresponding to kernel function k. s With target domain D t The maximum mean square difference, Represents the kernel mapping function. Represents the regenerated nucleus Hilbert space;
[0116] Frequency band weighted multi-core MMD loss is :
[0117]
[0118] in, These represent the EEG frequency bands associated with exercise preparation and exercise execution, respectively. and Represent the source and target domains in the frequency band respectively. Deep EEG features Indicates the weight of the corresponding frequency band. Indicates the corresponding kernel function; when emphasizing movement intention and movement preparation information, set... Frequency band weight greater than Frequency band weighting; when emphasizing cortical rhythms related to action execution, the weight distribution is adjusted according to the clinical task.
[0119] Furthermore, in step S8, in order to prevent the original cross-modal physiological connection structure from being destroyed after the three modal branches of EEG, sEMG and IMU adapt independently, a cross-modal graph structure preservation mechanism is introduced.
[0120] In the domain adaptation process, a cross-modal graph Laplacian regularization term is introduced to constrain the relative feature distance or connectivity of key cross-modal node pairs to remain stable before and after adaptation:
[0121]
[0122] Among them, L graph This indicates the cross-modal graph structure preservation loss. Denotes the set of cross-modal edges. This represents the cross-modal connectivity weights obtained from PLV, CMCI, or other physiological connectivity metrics. and These represent the adapted node features. This represents the square of the L2 norm.
[0123] Furthermore, in step S9, after completing differential domain adaptation and cross-modal structure preservation, the MAML-Rehab few-sample rapid individualized adaptation mechanism is adopted, enabling the model to complete patient-specific calibration when only a very small number of labeled samples are provided by the target patient.
[0124] Define a single patient as a meta-task :
[0125]
[0126] in, Indicates the first Support set for each patient Indicates the first The query set for each patient; the support set is used for rapid adaptation of the inner loop, and the query set is used for meta-optimization of the outer loop and verification of generalization performance.
[0127] Label vectors are uniformly defined as :
[0128]
[0129] Among them, y i,k This represents the label vector of the i-th patient and the k-th sample. Indicates FMA rating labels, This indicates the MBI scale rating label. Indicates a compensatory action label. Indicates muscle strength grading labels, This indicates the label for the cognitive-motor coupling index; T indicates transpose.
[0130] The inner loop update rule of MAML-Rehab is:
[0131]
[0132] in, Indicates the first The patient in Step parameters, Indicates the first The patient in Parameters updated during the inner loop step This represents the learning rate of the inner loop. Indicates support set loss; the inner loop performs 5 steps of gradient update; Indicates to Find the gradient;
[0133] Support set loss adopts a task uncertainty weighted form :
[0134]
[0135] in, Indicates the first The loss of each task on the support set Indicates the uncertainty parameter of the task;
[0136] The outer loop is updated as follows:
[0137]
[0138] in, This represents the outer loop learning rate. Indicates the batch size of the meta-task. Indicates based on initialization parameters Patient-specific parameters obtained after support set circulation update, Indicates the query set loss; Indicates to Find the gradient.
[0139] Furthermore, in step S10, unsupervised pre-training, dual-master task supervised learning, differential domain adaptation, cross-modal structure preservation, and rapid adaptation with few samples constitute a multi-stage collaborative optimization framework.
[0140] The total loss during the adaptation phase of the differentiated domain is defined as follows: :
[0141]
[0142]
[0143] in, This indicates the sEMG branch covariance alignment and off-diagonal preservation loss. This represents the adversarial adaptation loss of the residual statistical offset in the IMU branch. This indicates the weighted multi-core MMD loss in the EEG tributary band. This indicates the cross-modal graph structure preservation loss. , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade.
[0144] The overall joint optimization objective is expressed as:
[0145]
[0146] in, Used for unsupervised representation learning This is used for joint supervised learning of FMA scoring and MBI scale scoring in a dual-master task. Used for query set optimization after individualized and rapid adaptation. Used to maintain cross-modal physiological connectivity structures. Indicates the model initialization parameters; For different tasks T i Seeking expectations;
[0147] Through the aforementioned joint optimization objectives, an integrated technical loop is achieved, encompassing unlabeled data representation learning, labeled source domain supervised evaluation, target domain differential alignment, and small-sample individualized calibration of target patients.
[0148] To achieve the objectives of this invention, this invention also discloses a multimodal cross-individual fast adaptive evaluation system based on multi-stage collaboration, comprising the following modules:
[0149] Data acquisition module: used to acquire EEG, sEMG and IMU signals, and record action events, scale labels and patient basic information;
[0150] Preprocessing module: used to perform filtering, artifact suppression, rectification, normalization, attitude initialization, time synchronization, and windowed segmentation;
[0151] Unsupervised pre-training module: used to train a shared encoder based on historical unlabeled multimodal data;
[0152] Shared encoder module: used to construct the spatiotemporal heterogeneous graph backbone of ST-HGNN and extract node-level and graph-level deep features;
[0153] Dual-master-task supervised learning module: used to perform joint training of FMA score and MBI scale score, and simultaneously train auxiliary task head;
[0154] Differential Domain Adaptation Module: Used to perform sEMG covariance alignment, IMU geometric calibration and adversarial adaptation, and EEG multi-core MMD alignment respectively;
[0155] Cross-modal structure preservation module: used to maintain the cross-modal physiological connectivity topology of EEG-sEMG based on graph Laplacian regularization;
[0156] Small-sample rapid adaptation module: used to perform MAML-Rehab individualized calibration based on the target patient's 5-shot support set;
[0157] Assessment output module: Used to output FMA score, MBI scale score, compensation category, muscle strength grade, CMCI, PLV retention rate, reliability, and optional rehabilitation suggestions.
[0158] Compared with existing technologies, the significant advancements of this invention are as follows: 1) First, this invention connects unlabeled representation pre-training, source domain dual-master task supervised learning, target domain differential adaptation, and few-sample individualized calibration into a complete closed loop through a multi-stage collaborative framework. This can fully utilize unlabeled multimodal data from clinical history and reduce the reliance on a large number of manual scale annotations when deploying new patients; 2) This invention designs adaptation paths based on the cross-patient offset mechanisms of EEG, sEMG, and IMU, and maintains the brain-muscle coupling topology through cross-modal graph structure-maintaining constraints. This can improve cross-patient generalization ability while avoiding physiological structure caused by uniform global alignment. Distortion; 3) This invention sets the FMA score and MBI scale score as dual main task outputs, and enhances the interpretability with auxiliary tasks such as compensation recognition, muscle strength grading and CMCI, so that the assessment results simultaneously cover motor control recovery, life function outcome and compensation pattern recognition, which is more complete than single scale or single sensor assessment; 4) This invention uses the MAML-Rehab small sample rapid adaptation mechanism, so that the target patient only needs a small number of standardized movement samples to complete individualized calibration, reducing clinical recalibration time and patient burden, and is suitable for real deployment scenarios such as stroke rehabilitation assessment, efficacy follow-up, process monitoring and individualized rehabilitation suggestion generation.
[0159] To more clearly illustrate the functional characteristics and structural parameters of the present invention, further explanation is provided below in conjunction with the accompanying drawings and specific embodiments. Attached Figure Description
[0160] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0161] Figure 1 This is the overall system architecture diagram;
[0162] Figure 2 It is a multimodal acquisition and heterogeneous graph node mapping graph;
[0163] Figure 3 This is a flowchart of multimodal signal preprocessing and unsupervised pretraining;
[0164] Figure 4 This is a diagram of the ST-HGNN backbone model and the dual-master task output structure.
[0165] Figure 5 This is a diagram of the CDDA three-channel differentiated domain adaptive architecture;
[0166] Figure 6 This is a flowchart of the MAML-Rehab few-sample rapid individualized adaptation process;
[0167] Figure 7It is a flowchart of two-level domain adaptation and clinical deployment. Detailed Implementation
[0168] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0169] This invention discloses a multimodal cross-individual rapid adaptive assessment method and system based on multi-stage collaboration, which is aimed at the rehabilitation assessment of limb motor function in stroke patients. It solves the problems of poor cross-individual generalization of existing models when deployed to new patients, high requirements for labeled samples of target patients, and the disruption of cross-modal physiological structure by uniform domain alignment. This method simultaneously acquires EEG, sEMG, and IMU signals, and performs filtering, artifact suppression, normalization, posture initialization, time synchronization, and windowing segmentation. A shared neuromuscular-kinematic representation pedestal is established using historical unlabeled multimodal data. Dual-task supervised learning of FMA and MBI scores is performed on labeled source-domain patient data, with auxiliary constraints such as compensation identification, muscle strength grading, and the cognitive-motor coupling index (CMCI). Subsequently, collaborative adaptation is performed based on cross-patient shift mechanisms across different modalities. For the sEMG side, covariance alignment and muscle co-structural preservation are used; for the IMU side, gravity-direction geometric calibration and residual adversarial adaptation are used; and for the EEG side, frequency-weighted multi-core maximum mean difference alignment is used. Simultaneously, cross-modal graph structure preservation constraints maintain the topological stability of the EEG-sEMG neuromuscular coupling. Finally, MAML-Rehab rapid individualized calibration is performed based on a small number of labeled samples from the target patient, outputting patient-specific FMA scores, MBI scores, compensatory movement categories, muscle strength grading, and interpretable assessment indicators. This invention achieves rapid cross-patient deployment through a multi-stage collaborative closed loop: label-free representation pre-training, source domain clinical task learning, heterogeneous bias differential adaptation, and small-sample individualized calibration. This improves the accuracy, stability, and generalizability of clinical assessment. Figure 1 As shown, the overall system architecture of this invention includes modules such as multimodal acquisition and preprocessing, unsupervised pre-training, ST-HGNN shared encoder, dual master task supervised learning, differential domain adaptation, cross-modal structure preservation, MAML-Rehab fast adaptation, and evaluation output.
[0170] Specifically, the present invention includes the following steps:
[0171] Step S1: Acquire and preprocess EEG, sEMG, and IMU multimodal data. The patient performs a standardized upper limb movement task under a unified assessment process, while EEG, surface electromyography, and inertial motion signals are acquired simultaneously, and filtering, artifact suppression, rectification, normalization, posture initialization, time synchronization, and windowing segmentation are performed.
[0172] Step S2 involves unsupervised pre-training using unlabeled historical multimodal data. Through intramodal reconstruction, cross-modal contrastive consistency, and temporal structure learning tasks, the shared encoder learns general neuromuscular-motor representations without human-labeled scales. Figure 2 As shown, the three types of acquired signals, EEG, sEMG, and IMU, are mapped to EEG nodes, EMG nodes, and inertial nodes in the heterogeneous graph, respectively, for subsequent cross-modal graph modeling.
[0173] Step S3: Construct a multimodal spatiotemporal heterogeneous graph neural network (ST-HGNN). Organize 2 EEG nodes, 14 sEMG nodes, and 3 IMU nodes into a total of 19 heterogeneous nodes, establish intramodal dependency edges and cross-modal physiological coupling edges, and output the optimal 128-dimensional deep node features. For example... Figure 3 As shown, after filtering, artifact suppression, rectification, normalization, attitude initialization, time synchronization and windowing segmentation, the multimodal signal enters the unsupervised pre-training process, including intramodal reconstruction, cross-modal contrast consistency and temporal structure learning.
[0174] Step S4 involves performing supervised training on labeled source domain data using both FMA and MBI scales, and simultaneously training compensation recognition, muscle strength grading, and CMCI-assisted tasks. Figure 4 As shown, the ST-HGNN backbone model performs temporal encoding and graph convolution aggregation on heterogeneous EEG, sEMG, and IMU nodes, and connects the FMA score, MBI scale score, and auxiliary task output head.
[0175] Step S5 involves applying covariance alignment and off-diagonal preservation constraints to the sEMG channels to handle second-order statistical structure drift caused by differences in electrode location, subcutaneous fat thickness, and muscle synergistic patterns. Figure 5 As shown, the CDDA three-channel differential domain adaptation architecture performs covariance alignment, geometric calibration and adversarial adaptation, and frequency band weighted multi-core MMD alignment for cross-patient offset mechanisms of sEMG, IMU, and EEG.
[0176] Step S6: First, perform a gravity-direction-based geometric rotation calibration on the IMU channel, and then perform lightweight adversarial domain adaptation to handle wearing posture deviation and residual statistical offset.
[0177] Step S7: Perform frequency band weighted multi-core maximum mean difference alignment on the EEG channel to handle high-dimensional nonlinear offsets caused by scalp impedance, individual variations in cortical rhythms, and non-steady-state drift.
[0178] Step S8 introduces a cross-modal graph Laplacian structure preservation constraint to prevent independent adaptation from disrupting the EEG-sEMG brain-muscle coupling topology and PLV-related phase synchronization structure.
[0179] Step S9 involves using MAML-Rehab for rapid, individualized adaptation of the target patient with a small sample size. Each patient is defined as a meta-task, and a small number of gradient steps are updated using a 5-shot support set. For example... Figure 6 As shown, MAML-Rehab defines each target patient as a meta-task, enabling rapid individualized adaptation with few samples by supporting in-set loop updates and out-of-set loop optimization.
[0180] Step S10 outputs patient-specific assessment results, including FMA score, MBI scale score, compensatory movement category, muscle strength grade, and optional clinical interpretation information. For example... Figure 7 As shown, this invention forms a two-level domain adaptation process between source domain model training, target domain differential adaptation, and clinical deployment, enabling new patients to obtain patient-specific assessment results with a small number of labeled samples.
[0181] Example
[0182] To address the issues of insufficient cross-patient generalization ability, high cost of target patient recalibration, uniform handling of heterogeneous modal shifts, and incomplete clinical output dimensions in existing technologies, this embodiment is based on the joint input of EEG, sEMG, and IMU three modalities. First, a shared representation is established using historical unlabeled data. Then, dual-master task learning of FMA score and MBI scale score is completed on the source domain patients. Subsequently, differential adaptation is performed according to sEMG covariance shift, IMU geometric shift, and EEG nonlinear drift, respectively. The brain-muscle coupling topology is maintained through a cross-modal graph structure preservation mechanism. Finally, rapid individualized calibration is completed using a small number of labeled samples from the target patients.
[0183] (I) Data Acquisition and Task Design
[0184] This invention constructs a source domain sample set and a target domain sample set for the rehabilitation assessment of stroke patients. The source domain sample set consists of fully calibrated stroke patients and is used to train a general model for the population; the target domain sample set corresponds to newly enrolled patients and is used to simulate the cross-patient generalization and rapid individualized adaptation process of the model in real clinical deployment scenarios.
[0185] In a preferred embodiment, a source domain dataset can be established based on data from 50 stroke patients collected from the same clinical center, and a leave-one-out cross-validation method can be used to simulate the real cross-patient deployment process. In each round of the experiment, one patient is designated as the target patient who has not been seen at all, while the remaining patients are used as source domain training or meta-training data. After multiple rounds, the overall performance is statistically evaluated.
[0186] Patients performed standardized upper limb movement tasks under a unified assessment procedure. These tasks preferably included movements related to key FMA-UE assessment dimensions, such as shoulder flexion, shoulder abduction, elbow flexion and extension, forearm supination, and hand-to-lumbar spine touch. For each movement test, EEG, sEMG, and IMU signals were simultaneously acquired, and the corresponding clinical labels or functional assessment results were recorded.
[0187] EEG acquisition preferably uses dual-channel EEG signals to characterize cortical motor-related activities during patient preparation and execution of movements; sEMG acquisition preferably uses 14-channel surface electromyography signals to characterize activation patterns and muscle synergy of relevant upper limb muscle groups; IMU acquisition preferably uses three inertial measurement units, respectively positioned at corresponding locations in key upper limb motor pathways, to characterize limb posture changes, angular velocity, acceleration, and trajectory information. This forms a trimodal joint input including cortical activity, muscle activation, and kinematic output.
[0188] To balance offline population modeling with rapid online individualized calibration, this invention further designs the data acquisition process for new patients in the target domain in a task-oriented manner. In actual deployment, target patients first complete a small number of standardized support set movements. The support set preferably consists of five samples, corresponding to five key movements: T1 shoulder flexion, T2 shoulder abduction, T3 elbow flexion and extension, T4 forearm supination, and T5 hand-to-lumbar spine touch, thereby covering the main upper limb functional assessment dimensions with relatively low annotation costs.
[0189] For each target domain patient, the support set is used for rapid online adaptation, while the remaining unseen action samples of the same patient can be used as a query set or test set to verify the generalization performance after adaptation. In a preferred embodiment, each patient is considered a meta-task unit, the support set contains 5 complete assessment action samples, and the query set contains several additional unseen samples of the same patient, to simulate the actual scenario in clinical practice where new patients have only a small number of labeled action samples and need to quickly complete individualized model calibration.
[0190] The data collection labels include at least the FMA score or FMA-UE score, and may also include the MBI scale score, compensatory movement classification labels, muscle strength grading labels, and the cognitive-motor coupling correlation index (CMCI). Through this design, the present invention establishes a labeling system for multi-task assessment during the data collection phase, providing a unified data foundation for subsequent regression tasks, classification tasks, and cross-modal coupling index estimation tasks.
[0191] The source and target domain sample sets can be represented as follows:
[0192]
[0193] in, This represents the source domain patient dataset, where the subscript 's' indicates 'source', i.e., the source domain. This represents the target domain patient dataset, where the subscript t indicates target, i.e., the target domain; This represents the multimodal input of the i-th sample in the source domain. This represents the clinical assessment label corresponding to the i-th sample in the source domain; This represents the multimodal input of the j-th sample in the target domain. This represents the clinical assessment label corresponding to the j-th sample in the target domain; the superscripts s and t represent the source domain and target domain respectively, and the subscripts i and j represent the sample indices in the source domain and target domain respectively; x is the lowercase input sample vector, y is the lowercase label vector; D is the uppercase dataset symbol; n s and n t These represent the number of samples in the source and target domains, respectively.
[0194] Multimodal joint input is defined as:
[0195]
[0196] in, Indicates EEG characteristics, Indicates sEMG features, This represents IMU features.
[0197] (II) Multimodal signal preprocessing
[0198] After obtaining the raw multimodal time-series data, this invention performs targeted preprocessing on EEG, sEMG, and IMU signals respectively, and further completes cross-modal time synchronization, sample windowing, and input format unification to form the standardized input for subsequent multimodal heterogeneous graph models. The purpose of this preprocessing is to reduce the impact of noise, artifacts, drift, and temporal misalignment on the subsequent graph modeling and domain adaptation processes while preserving the physiological significance of the three types of signals as much as possible.
[0199] For EEG signals, bandpass filtering and power line interference suppression are first performed to reduce low-frequency drift and power supply noise. Artifact suppression and baseline correction are then performed to reduce the impact of eye movement, electromyography crosstalk, and electrode contact instability. The EEG signals are then segmented according to the task time window, and frequency band-related features can be extracted according to predetermined frequency bands, or the segmented time segments can be directly used as input to the depth model. Considering that EEG cross-patient offset mainly originates from scalp impedance fluctuations, individual variations in cortical rhythms, and non-steady-state drift, this type of preprocessing helps improve the stability of EEG features and provides a more suitable input distribution for subsequent frequency band-weighted multi-nucleus MMD alignment.
[0200] The feature representation after EEG preprocessing can be written as:
[0201]
[0202] in, Represents raw brain electrical signals. This indicates filtering, artifact suppression, and baseline correction operations. This indicates a time window segmentation operation. This represents the brainwave frequency band feature mapping or deep coding mapping.
[0203] For sEMG signals, bandpass filtering and power frequency notch filtering are first performed, followed by rectification, envelope extraction, and normalization. The signals are then time-series segmented according to a fixed sliding window or fixed event window aligned with the action task. Since the differences in sEMG across patients mainly manifest as electrode position offset, subcutaneous fat thickness differences, and the resulting distortion of the inter-channel covariance structure, sEMG preprocessing preferably preserves the relative synergistic relationships between channels, rather than focusing solely on the amplitude of individual channels. This provides a data foundation for subsequent domain adaptation based on covariance alignment and off-diagonal preservation constraints.
[0204] The sEMG feature construction process can be written as:
[0205]
[0206] in, Indicates primitive muscle, This indicates bandpass and notch filtering. Indicates rectification or envelope extraction. Indicates normalization, This indicates the division of the time window. This represents electromyographic feature mapping or deep coding mapping.
[0207] For IMU signals, zero-bias calibration and sampling synchronization are performed first. Then, the gravity direction is estimated based on the static sitting posture stage or other reference motion stages to obtain the rotation relationship of the current coordinate system of each IMU unit relative to the human anatomical coordinate system. Next, the motion time-series signal is uniformly segmented and normalized. Since IMU cross-patient offset is primarily manifested as geometric offset caused by differences in wearing posture, the preprocessing stage retains static gravity reference information, which directly supports the subsequent geometric calibration steps.
[0208] The IMU input representation can be written as:
[0209]
[0210] in, Represents the original IMU signal. This indicates the attitude initialization, zero-bias calibration, and coordinate adjustment operations. This indicates the division of the time window. This represents the IMU feature map or deep coding map.
[0211] The mean acceleration vector within the target patient's stationary segment can be defined as:
[0212]
[0213] in, This represents the IMU acceleration vector within a stationary segment. Indicates the length of a static segment. Used for subsequent IMU gravity orientation geometry calibration.
[0214] After completing single-modal preprocessing, this invention further performs cross-modal time synchronization and sample windowing processing. Preferably, using a unified action start time, triggering event, or predetermined time anchor point as an alignment benchmark, EEG, sEMG, and IMU are resampled or indexed for a unified time axis, ensuring that the three types of signals have one-to-one corresponding time segments within the same action window. Subsequently, the multimodal signals within the corresponding time window are combined into the same sample unit, and action labels, scale scores, or other evaluation labels are attached to it.
[0215] No. A joint sample can be represented as:
[0216]
[0217] in, and Other means the first The start and end positions of the time window for each action sample.
[0218] (III) Multimodal Unsupervised Pre-training Module
[0219] This invention adds a multimodal unsupervised pre-training stage before supervised learning to learn stable shared encoder parameters using unlabeled clinical historical EEG, sEMG, and IMU data. This module is introduced because real-world clinical scenarios often involve a large amount of historical data with or without labels, and high-quality labels for scales such as FMA and MBI rely on manual assessment by therapists or clinicians, which is costly to obtain. Through unsupervised pre-training, this invention can learn general physiological and motor representations without relying on manual scale labels, providing a more stable initialization for subsequent supervised assessment, differential domain adaptation, and rapid adaptation with few samples.
[0220] The unsupervised pre-training preferably includes intra-modal reconstruction tasks, cross-modal contrast consistency tasks, and temporal structure learning tasks.
[0221] Intramodal reconstruction is used to recover the original signal structure from the perturbed modal representation, enabling the encoder to learn the temporal patterns and local structural information of each mode. Its loss can be defined as:
[0222]
[0223] in, Indicates the first The original input or feature representation of a modality. This indicates the reconstruction result of the corresponding mode.
[0224] The cross-modal contrast consistency task aims to enhance the consistency between different modal representations within the same time window, ensuring that EEG, sEMG, and IMU features of the same action segment are close to each other in a shared representation space, while non-corresponding segments from different actions, different time windows, or different patients are distinguished from each other in the shared representation space. Its loss can be defined as:
[0225]
[0226] in, Denotes the set of positive sample pairs. Representation and Sample The corresponding candidate sample set, , , These represent the encoded sample representations, Represents the similarity function. Indicates the temperature coefficient; , , Used to distinguish different sample representations; exp is an exponential function.
[0227] Temporal structure learning tasks are used to preserve the temporal relationships between actions during execution, enabling the model to learn the dynamic changes between stages such as initiation, execution, peak, and recovery. Its loss can be defined as:
[0228]
[0229] in, This represents a time-series relationship prediction function. and Encoded representation of adjacent time segments, A marker indicating the target temporal relationship between adjacent time segments.
[0230] The total loss for multimodal unsupervised pre-training is:
[0231]
[0232]
[0233] in, , , These are the weight coefficients for intra-modal reconstruction loss, cross-modal contrastive consistency loss, and temporal structure learning loss, respectively. The weight coefficients are empirically set based on the importance of the three pre-training tasks, with the cross-modal contrastive consistency task, used to learn a unified shared representation, thus receiving a higher weight. =0.4; The intramodal reconstruction task and the temporal structure learning task are used to preserve the modality's own structural information and temporal dynamic features, therefore they are given the same weight. ;
[0234] By minimizing This invention enables the learning of a stable multimodal shared encoder. The pre-trained encoder parameters serve as initialization parameters for subsequent dual-main-task supervised learning of FMA and MBI scale scores, CDDA differential domain adaptation, and MAML-Rehab few-shot fast adaptation, thereby reducing label dependence and improving model stability across patients and in low-sample scenarios.
[0235] (iv) Backbone Model and Spatiotemporal Heterogeneous Graph Modeling
[0236] This invention employs a multimodal spatiotemporal heterogeneous graph neural network (ST-HGNN) as a shared backbone model to uniformly model intramodal dependencies and cross-modal coupling relationships among EEG, sEMG, and IMU, and outputs deep feature representations suitable for rehabilitation assessment and domain adaptation. The role of this backbone model is to transform traditional multimodal parallel inputs into heterogeneous graph inputs with explicit structural relationships, thereby simultaneously utilizing complementary information from brain activity, muscle activation, and kinematic trajectory.
[0237] The backbone model's node set consists of three types of heterogeneous nodes: EEG nodes, sEMG nodes, and IMU nodes. Specifically, EEG nodes preferably correspond to dual-channel EEG acquisition locations, sEMG nodes preferably correspond to 14 surface electromyography (EMG) channels, and IMU nodes preferably correspond to 3 inertial measurement units (IMUs). Thus, the entire heterogeneous map contains 19 nodes, capable of simultaneously carrying information from the cortex, muscle groups, and limb segments.
[0238] In terms of edge relationship construction, this invention preferably establishes both intramodal and intermodal edges simultaneously. Intramodal edges represent physiological proximity, temporal continuity, or statistical correlation within the same modality, such as muscle synergy between different sEMG channels, brain region synergy between different EEG nodes, and kinetic chain coupling between IMU nodes. Intermodal edges represent functional connectivity between different modalities, particularly the brain-muscle coupling between EEG and sEMG, and the muscle activation-kinematic response relationship between sEMG and IMU. Through the above graph structure, the backbone model can jointly encode cortical activity, muscle activation patterns, and action execution results within the same framework.
[0239] The node is initially represented as:
[0240]
[0241] in, These represent the initial feature representations of the EEG, sEMG, and IMU nodes, respectively.
[0242] The graph convolution or message passing update process can be written as:
[0243]
[0244] in, Represents the normalized adjacency matrix. Indicates the first Layer node representation, This represents the learnable weight matrix of the lll-th layer. This represents a non-linear activation function.
[0245] In a preferred embodiment, the ST-HGNN may include temporal coding units and graph convolutional units. The temporal coding units are used to extract the local dynamic representation of each node within a time window, while the graph convolutional units are used for information propagation and feature aggregation on a heterogeneous graph topology. After multi-layer stacking, the model outputs deep node embedding features of a uniform dimension and optional graph-level feature representations. The backbone network outputs 128-dimensional deep node features for use by the subsequent dual-master task evaluation head and differential domain adaptation module.
[0246] To balance population modeling capabilities with rapid clinical deployment efficiency, this invention preferably trains the ST-HGNN backbone model extensively during the source domain training phase, enabling it to learn a relatively stable neuromuscular-kinematic coupling topology within the stroke patient population. In the subsequent cross-patient domain adaptation and individualized rapid calibration phases, most parameters of the backbone network are preferably frozen, and lightweight adaptation layers are only superimposed on its output or mid-to-high-level feature spaces for differentiated domain adaptation and rapid meta-learning updates, in order to avoid disrupting the neuromuscular conduction topology already learned in the source domain.
[0247] (v) Combined output of FMA score and MBI scale score dual main task
[0248] This invention sets FMA score prediction and MBI scale score prediction as parallel dual-primary tasks, and sets compensation identification, muscle strength grading, and CMCI as auxiliary tasks. The FMA score primarily reflects the degree of motor impairment and motor function recovery in stroke patients, while the MBI scale score primarily reflects the patient's activities of daily living and functional independence. The two are related but not equivalent: the FMA score focuses more on the quality of movement completion and motor control, while the MBI score focuses more on whether the patient can translate motor recovery into practical daily living abilities. Therefore, setting FMA and MBI as dual-primary tasks with joint output allows the model to simultaneously assess the patient's motor performance and the extent of functional recovery, improving the clinical completeness of the assessment results.
[0249] Assume the shared encoder output is The output of the dual main task and auxiliary task is then represented as follows:
[0250]
[0251] in, This indicates the FMA prediction result. This indicates the MBI prediction result. This indicates the result of the compensation action recognition. This indicates the results of the muscle strength grading prediction. This represents the prediction results of the cognitive-motor coupling index.
[0252] The corresponding multi-task prediction head can be written as:
[0253]
[0254]
[0255] in, , , , and These represent the FMA score, MBI scale score, compensatory movement recognition, muscle strength grading, and CMCI prediction head, respectively.
[0256] To highlight the importance of the FMA score and MBI scale score as dual main tasks, this invention adopts a joint loss form that prioritizes the main tasks and enhances the auxiliary tasks:
[0257]
[0258]
[0259] in, and These represent the prediction losses of FMA and MBI, respectively, with mean squared error, mean absolute error, or Huber loss being preferred. The classification loss for compensatory actions is represented, with cross-entropy loss being the preferred method. Indicates muscle strength grade loss; This represents the predictive loss of the cognitive-motor coupling index. This indicates the cross-modal graph structure preservation loss; , , , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade.
[0260] In the preferred embodiment, the following condition is met:
[0261]
[0262] The primary output targets are FMA and MBI scores, while compensation identification, muscle strength grading, and CMCI are used to enhance representation learning and clinical interpretability. In another implementation, an adaptive weighting mechanism based on task uncertainty can be used to achieve a dynamic balance between the dual primary tasks and auxiliary tasks.
[0263] (vi) Differentiated Domain Adaptation Module (CDDA)
[0264] This invention proposes a channel-specific differentiated domain adaptation framework (CDDA). Instead of directly applying a uniform global domain adaptation algorithm to the fused multimodal features, this framework designs adaptation paths based on the different offset mechanisms of EEG, sEMG, and IMU modalities across patient scenarios. Specifically, sEMG offset is mainly manifested as channel covariance structure drift, IMU offset is mainly manifested as geometric rotational offset caused by wearing posture differences, and EEG offset is mainly manifested as high-dimensional nonlinear non-steady-state distribution drift. Applying the same global alignment strategy to all three can easily lead to over-correction and disrupt cross-modal physiological connectivity structures. Therefore, this invention employs a modality- and mechanism-specific differentiated adaptation strategy.
[0265] sEMG channel: covariance alignment and off-diagonal preservation.
[0266] For sEMG modalities, cross-patient differences are mainly manifested as channel covariance structure drift caused by electrode position offset, subcutaneous fat thickness differences, and differences in muscle synergistic patterns. Therefore, this invention applies covariance alignment constraints to the deep sEMG features of the source and target domains to make their second-order statistics more consistent; at the same time, off-diagonal preservation constraints are introduced to maintain the channel-related structure corresponding to the muscle synergistic activation patterns.
[0267] Let the deep sEMG features of the source and target domains be as follows:
[0268]
[0269] in, and These represent the number of samples in the source and target domains, respectively. This indicates the dimension of deep features in sEMG.
[0270] The source domain sEMG eigencovariance matrix is estimated as follows:
[0271]
[0272] The target domain sEMG feature covariance matrix is estimated as follows:
[0273]
[0274] Covariance alignment loss is defined as:
[0275]
[0276] in, Let T denote the Frobenius norm, and T denote the transpose. This represents a unit vector.
[0277] To maintain the synergistic muscle structure, off-diagonal conservation constraints are further introduced:
[0278]
[0279] in, The off-diagonal part of the target domain sEMG covariance matrix is represented. This represents the off-diagonal portion of the reference covariance matrix. The weights are set to a range of [0.1, 10]. The values are determined by optimizing the similarity of muscle synergistic structures and the performance of downstream tasks on the validation set through grid search. These values include FMA score, MBI scale score, compensatory movement category, and muscle strength grade. The off-diagonal elements reflect the correlation between different electromyographic channels, so this constraint helps to maintain the muscle synergistic activation mode.
[0280] The total adaptation loss of the sEMG tributary is:
[0281]
[0282] IMU Channels: Geometric Calibration and Residual Adversarial Adaptation.
[0283] For IMU modalities, cross-patient offset primarily manifests as geometric rotation caused by differences in sensor wearing posture. Directly aligning the statistical distribution of IMU features may overlook the physical causes of IMU signal offset. Therefore, this invention first determines the rotational alignment relationship between the IMU coordinate system and the human anatomical coordinate system based on the static gravity direction, and then performs lightweight adversarial adaptation on the residual statistical offset after geometric calibration.
[0284] Let the direction of gravity estimated from the static segment of the target patient be... The source region reference gravity direction is The optimal rotation matrix can be solved as follows:
[0285]
[0286] in, Represents a set of three-dimensional rotation matrices. This represents the square of the L2 norm.
[0287] The calibrated target domain IMU acceleration or characteristic is expressed as follows:
[0288]
[0289] in, This represents the original target domain IMU acceleration vector or feature vector. This represents the IMU representation after geometric calibration.
[0290] After geometric calibration, this invention further introduces a domain discriminator D and a feature extractor G, and employs a lightweight adversarial loss to handle the residual statistical offset:
[0291]
[0292] Through the above design, the IMU branch first eliminates the main geometric offset caused by wearing posture, and then processes patient individual differences, movement amplitude differences and residual statistical drift through adversarial adaptation, thereby obtaining more stable inertial motion characteristics.
[0293] EEG Channel: Band-weighted multi-core MMD distribution alignment.
[0294] For EEG modalities, cross-patient shifts primarily originate from scalp impedance fluctuations, individual variability in cortical rhythms, and non-steady-state drift, with the shift pattern more closely resembling high-dimensional nonlinear distribution drift. Therefore, this invention employs multi-nuclear maximum mean difference for kernel spatial distribution alignment and assigns weights to different kernel components according to the importance of rehabilitation assessment indicators across different frequency bands; D s and D t These represent the source and target domain datasets, respectively. / This indicates that sample x originates from D. s / D t Distribution; E[ϕ(x)] represents the average feature of the data after mapping to a high-dimensional space;
[0295] MMD distance is defined as:
[0296]
[0297] in, The source domain D is represented in the reproducing kernel Hilbert space corresponding to kernel function k. s With target domain D t The maximum mean square difference, Represents the kernel mapping function. This represents the regenerated nucleus Hilbert space.
[0298] The frequency band weighted multi-core MMD loss is:
[0299]
[0300] in, These represent the EEG frequency bands associated with exercise preparation and exercise execution, respectively. and Represent the source and target domains in the frequency band respectively. Deep EEG features Indicates the weight of the corresponding frequency band. This indicates the corresponding kernel function. When emphasizing movement intention and preparation information, this can be set... Frequency band weight greater than Frequency band weighting; when emphasizing cortical rhythms related to action execution, the weight distribution can be adjusted according to the clinical task.
[0301] Through the above methods, the present invention can reduce the difference in deep EEG feature distribution between the source domain and the target domain in the kernel space, while retaining key frequency band information related to rehabilitation assessment.
[0302] (vii) Preservation of cross-modal graph structure
[0303] To prevent the disruption of the original cross-modal physiological connectivity structure after the independent adaptation of the three modal branches of EEG, sEMG, and IMU, this invention introduces a cross-modal graph structure preservation mechanism. This mechanism focuses on constraining the brain-muscle coupling topology between EEG and sEMG, as well as the phase synchronization and cognitive-motor coupling relationships characterized by indicators such as PLV and CMCI.
[0304] Specifically, this invention introduces a cross-modal graph Laplacian regularization term during the domain adaptation process to constrain the relative feature distances or connectivity relationships of key cross-modal node pairs to remain stable before and after adaptation:
[0305]
[0306] Among them, L graph This indicates the cross-modal graph structure preservation loss. The set of cross-modal edges is preferably a neuromuscular coupling edge between EEG nodes and sEMG nodes, and may also include a muscle activation-motor response edge between sEMG nodes and IMU nodes. This represents the cross-modal connectivity weights obtained from PLV, CMCI, or other physiological connectivity metrics. and These represent the adapted node features.
[0307] This constraint allows the model to reduce the distributional differences between the source and target domains while maintaining the interpretability of the neuromuscular coupling physiological topology and phase synchronization structure. Compared to unified global alignment methods, this invention can reduce cross-modal structural distortion caused by over-alignment, resulting in more stable CMCI estimation and PLV retention rate.
[0308] (viii) MAML-Rehab rapid individualized adaptation with few samples
[0309] After completing differential domain adaptation and cross-modal structure preservation, this invention further employs the MAML-Rehab few-sample rapid individualized adaptation mechanism, enabling the model to complete patient-specific calibration when only a very small number of labeled samples are provided by the target patient.
[0310] This invention defines a single patient as a meta-task:
[0311]
[0312] in, Indicates the first Support set for each patient Indicates the first The query set is used for each patient. The support set is used for fast adaptation in the inner loop, and the query set is used for meta-optimization of the outer loop and verification of generalization performance.
[0313] In a preferred embodiment, the support set contains 5 complete evaluation action samples, and the query set contains 10 unseen samples:
[0314]
[0315] The label vector is uniformly defined as:
[0316]
[0317] Among them, y i,k This represents the label vector of the i-th patient and the k-th sample. Indicates FMA rating labels, This indicates the MBI scale rating label. Indicates a compensatory action label. Indicates muscle strength grading labels, This indicates the labeling of cognitive-motor coupling indicators.
[0318] The inner loop update rule of MAML-Rehab is:
[0319]
[0320] in, Indicates the first The patient in Step parameters, Indicates the first The patient in Parameters updated during the inner loop step This represents the learning rate of the inner loop. This indicates a support set loss. Preferably, the inner loop performs 5 steps of gradient updates; Indicates to Find the gradient.
[0321] The support set loss can be expressed as a task uncertainty-weighted form:
[0322]
[0323] in, Indicates the first The loss of each task on the support set This represents the task uncertainty parameter. This method allows for adaptive adjustment of the loss weights for each task under different patient conditions, task requirements, and annotated noise levels.
[0324] The outer loop is updated as follows:
[0325]
[0326] in, This represents the outer loop learning rate. Indicates the batch size of the meta-task. Indicates based on initialization parameters Patient-specific parameters obtained after support set circulation update, Indicates the query set loss; Indicates to Find the gradient.
[0327] Through the aforementioned two-layer optimization, the model can learn an initialization parameter that is easily adapted quickly with a small number of samples, enabling the target patient to obtain a patient-specific assessment model with only a few labeled action samples. Compared with traditional retraining or long-term recalibration methods, this invention significantly reduces annotation costs, data collection time, and patient burden in clinical deployment.
[0328] (ix) Joint optimization objectives
[0329] In this invention, unsupervised pre-training, dual-master task supervised learning, differential domain adaptation, cross-modal structure preservation, and rapid adaptation with few samples constitute a multi-stage collaborative optimization framework.
[0330] The total loss during the differentiation domain adaptation phase is defined as:
[0331]
[0332]
[0333] in, This indicates the sEMG branch covariance alignment and off-diagonal preservation loss. This represents the adversarial adaptation loss of the residual statistical offset in the IMU branch. This indicates the weighted multi-core MMD loss in the EEG tributary band. This indicates the cross-modal graph structure preservation loss. , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade.
[0334] In a preferred embodiment, the following can be configured:
[0335]
[0336] However, the above parameters are only preferred examples and can be adjusted according to different equipment, different center data and different clinical tasks.
[0337] The overall joint optimization objective can be expressed as:
[0338]
[0339] in, Used for unsupervised representation learning This is used for joint supervised learning of FMA scoring and MBI scale scoring in a dual-master task. Used for query set optimization after individualized and rapid adaptation. Used to maintain cross-modal physiological connectivity structures. Indicates the model initialization parameters. Indicates different tasks T i Seeking expectations.
[0340] Through the aforementioned joint optimization objectives, this invention can achieve an integrated technical closed loop from unlabeled data representation learning, labeled source domain supervised evaluation, target domain differential alignment to small-sample individualized calibration of target patients.
[0341] This invention can also be implemented as a rehabilitation assessment system for limb motor function in stroke patients. This system includes at least the following modules:
[0342] Data acquisition module: used to acquire EEG, sEMG and IMU signals, and record action events, scale labels and patient basic information.
[0343] Preprocessing module: used to perform filtering, artifact suppression, rectification, normalization, attitude initialization, time synchronization, and windowed segmentation.
[0344] Unsupervised pre-training module: used to train a shared encoder based on historical unlabeled multimodal data.
[0345] Shared encoder module: used to construct the spatiotemporal heterogeneous graph backbone of ST-HGNN and extract node-level and graph-level deep features.
[0346] Dual-master-task supervised learning module: used to perform joint training of FMA score and MBI scale score, and simultaneously train auxiliary task head.
[0347] Differentiated domain adaptation module: used to perform sEMG covariance alignment, IMU geometric calibration and adversarial adaptation, and EEG multi-core MMD alignment respectively.
[0348] Cross-modal structure preservation module: used to maintain the cross-modal physiological connection topology of EEG-sEMG based on graph Laplacian regularization.
[0349] Small sample rapid adaptation module: used to perform MAML-Rehab individualized calibration based on the target patient 5-shot support set.
[0350] Assessment output module: Used to output FMA score, MBI scale score, compensation category, muscle strength grade, CMCI, PLV retention rate, reliability, and optional rehabilitation suggestions.
[0351] This invention proposes a multi-stage collaborative rapid adaptive evaluation framework consisting of "label-free pre-training - dual-primary task supervision - differential domain adaptation - small-sample individualized calibration". This framework uses shared representations learned from historical label-free EEG / sEMG / IMU data as the starting point for cross-patient transfer, and uses FMA scores and MBI scale scores as clinical dual-primary tasks to form source domain supervision constraints. Then, it eliminates heterogeneous bias in new patients through target domain differential adaptation, and finally uses a very small number of labeled samples from target patients to correct individual residuals, forming a closed-loop technical path for rapid deployment to new clinical patients. It proposes a joint mechanism of three-channel differential domain adaptation and cross-modal structure preservation based on modal shift mechanisms. For the sEMG channel, which is mainly characterized by second-order statistical drift caused by differences in electrode position, subcutaneous fat, and muscle coordination, covariance alignment is used while maintaining off-diagonal muscle coordination structures. For the IMU channel, which is mainly characterized by geometric rotational shift caused by wearing posture, physical geometric calibration is first performed in the direction of gravity, and then lightweight adversarial adaptation is applied to the residual shift. To address the high-dimensional nonlinear drift characteristics of EEG channels, primarily caused by individual differences in scalp impedance, cortical rhythms, and non-stationary changes, a frequency-weighted multi-core MMD alignment is employed. Three adaptation paths are simultaneously constrained by cross-modal graph Laplace structure preservation to maintain the EEG-sEMG brain-muscle coupling topology and PLV / CMCI-related phase synchronization structure, avoiding over-correction and physiological connectivity distortion caused by traditional unified global alignment. A dual-scale main output and 5-shot individualized meta-adaptation mechanism for rapid deployment to new patients are proposed. The system defines a single patient as a meta-task. Even with only a small number of standardized action samples provided by the target patient, patient-specific parameters are updated via MAML-Rehab, simultaneously outputting FMA and MBI scale scores. This allows the model to assess both the degree of motor control recovery and the outcome of activities of daily living. Compensation identification, muscle strength grading, and CMCI serve as auxiliary outputs in the constraints, ensuring that the results after small-sample calibration possess scale integrity, cross-modal interpretability, and clinical usability.
[0352] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0353] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration, characterized in that, Includes the following steps: Step S1: Acquire and preprocess multimodal data including electroencephalogram (EEG), surface electromyography (sEMG), and inertial measurement unit (IMU) signals; Patients performed standardized upper limb movement tasks under a unified assessment process, while EEG, surface electromyography, and inertial motion signals were collected simultaneously. Filtering, artifact suppression, rectification, normalization, posture initialization, time synchronization, and windowing segmentation were also performed. Step S2: Unsupervised pre-training is performed using unlabeled historical multimodal data; through intramodal reconstruction, cross-modal contrast consistency and temporal structure learning tasks, the shared encoder learns general neuro-muscle-motor representations without human scale labels. Step S3: Construct a multimodal spatiotemporal heterogeneous graph neural network ST-HGNN; Step S4: Perform supervised training of FMA score and MBI scale score on labeled source domain data, and simultaneously train compensation recognition, muscle strength classification and CMCI auxiliary tasks. Step S5: Perform covariance alignment and off-diagonal preservation constraints on the sEMG channels to handle second-order statistical structure drift caused by differences in electrode location, subcutaneous fat thickness, and muscle synergy patterns. Step S6: First, perform a gravity-direction-based geometric rotation calibration on the IMU channel, and then perform a lightweight adversarial domain adaptation to handle wearing posture deviation and residual statistical offset. Step S7: Perform frequency band weighted multi-core maximum mean difference alignment on the EEG channel to handle high-dimensional nonlinear offsets caused by scalp impedance, individual variations in cortical rhythms, and non-steady-state drift. Step S8: Introduce cross-modal graph Laplacian structure preservation constraints to prevent independent adaptation from disrupting the EEG-sEMG brain-muscle coupling topology and PLV-related phase synchronization structure; Step S9: Use MAML-Rehab to perform rapid individualized adaptation of target patients with a small number of samples. Define a single patient as a meta-task and use a 5-shot support set to complete a small number of gradient step updates. Step S10: Output patient-specific assessment results, including FMA score, MBI scale score, compensatory movement category, muscle strength grade, and optional clinical interpretation information.
2. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S1, a source domain sample set and a target domain sample set are constructed for the rehabilitation assessment scenario of stroke patients. The source domain sample set consists of fully calibrated stroke patients and is used to train the general model for the population. The target domain sample set corresponds to newly enrolled patients and is used to simulate the cross-patient generalization and rapid individualized adaptation process of the model in real clinical deployment scenarios. The source and target domain sample sets are represented as follows: in, This represents the source domain patient dataset, where the subscript 's' indicates 'source', i.e., the source domain. This represents the target domain patient dataset, where the subscript t indicates target, i.e., the target domain; This represents the multimodal input of the i-th sample in the source domain. This represents the clinical assessment label corresponding to the i-th sample in the source domain; This represents the multimodal input of the j-th sample in the target domain. This represents the clinical assessment label corresponding to the j-th sample in the target domain; the superscripts s and t represent the source domain and target domain respectively, and the subscripts i and j represent the sample indices in the source domain and target domain respectively; x is the lowercase input sample vector, y is the lowercase label vector; D is the uppercase dataset symbol; n s and n t These represent the number of samples in the source and target domains, respectively. Multimodal joint input is defined as: in, Indicates EEG characteristics, Indicates sEMG features, Indicate IMU features; After obtaining the raw multimodal time series data, targeted preprocessing was performed on the three types of signals: EEG, sEMG, and IMU. Furthermore, cross-modal time synchronization, sample windowing, and input format unification were completed to form the standardized input for the subsequent multimodal heterogeneous graph model. For EEG signals, bandpass filtering and power frequency interference suppression are first performed to reduce low-frequency drift and power supply noise; then artifact suppression and baseline correction are performed to reduce the impact of eye movement, electromyography crosstalk and electrode contact instability; then the EEG signals are segmented according to the task time window, and frequency band related features are extracted according to the predetermined frequency band, or the segmented time segments are directly used as input to the depth model. The features after EEG preprocessing are represented as follows: : in, Represents raw brain electrical signals. This indicates filtering, artifact suppression, and baseline correction operations. This indicates a time window segmentation operation. This represents a brainwave frequency band feature mapping or deep coding mapping; For sEMG signals, bandpass filtering and power frequency notch filtering are first performed, followed by rectification, envelope extraction and normalization, and time-series segmentation is performed according to a fixed sliding window or fixed event window aligned with the action task. The sEMG feature construction process is as follows : in, Indicates primitive muscle, This indicates bandpass and notch filtering. Indicates rectification or envelope extraction. Indicates normalization, This indicates the division of the time window. This represents electromyographic feature mapping or deep coding mapping; For IMU signals, zero-bias calibration and sampling synchronization are performed first. Then, the gravity direction is estimated based on the static sitting posture stage or other reference motion stage to obtain the rotation relationship of the current coordinate system of each IMU unit relative to the human anatomical coordinate system. Next, the motion time sequence signal is uniformly segmented and normalized. Static gravity reference information is preserved in the preprocessing stage. IMU input is represented as : in, Represents the original IMU signal. This indicates the attitude initialization, zero-bias calibration, and coordinate adjustment operations. This indicates the division of the time window. This represents the IMU feature map or deep coding map; The mean acceleration vector within the target patient's stationary segment is defined as : in, This represents the IMU acceleration vector within a stationary segment. Indicates the length of a static segment. Used for subsequent IMU gravity orientation geometry calibration; t is the time index within the static segment; After completing the single-modal preprocessing, cross-modal time synchronization and sample windowing are further performed. Using the unified action start time, trigger event, or predetermined time anchor point as the alignment benchmark, EEG, sEMG, and IMU are resampled or indexed and aligned in a unified time axis so that the three types of signals have a one-to-one corresponding time segment within the same action window. Subsequently, the multimodal signals within the corresponding time window are combined into the same sample unit and action labels, scale scores, or other evaluation labels are attached to it. No. A joint sample is represented as : in, and Other means the first The start and end positions of the time window for each action sample.
3. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S2, a multimodal unsupervised pre-training stage is added before supervised learning to learn stable shared encoder parameters using clinical history unlabeled EEG, sEMG, and IMU data; Unsupervised pre-training includes intra-modal reconstruction tasks, cross-modal contrastive consistency tasks, and temporal structure learning tasks; Intramodal reconstruction is used to recover the original signal structure from the perturbed modal representation, enabling the encoder to learn the temporal patterns and local structural information of each mode. The loss is defined as... : Where m represents the mode number. Indicates the first The original input or feature representation of a modality. This indicates the reconstruction result of the corresponding mode; Represents the square of the L2 norm; The cross-modal contrast consistency task aims to enhance the consistency between different modal representations within the same time window, making EEG, sEMG, and IMU features of the same action segment close to each other in a shared representation space, while distinguishing non-corresponding segments from different actions, different time windows, or different patients. The loss is defined as... : in, Denotes the set of positive sample pairs. Representation and Sample The corresponding candidate sample set, , , These represent the encoded sample representations, Represents the similarity function. Indicates the temperature coefficient; , , Used to distinguish different sample representations; exp is an exponential function; The temporal structure learning task is used to preserve the sequential relationships between actions during execution, enabling the model to learn the dynamic changes between the initiation, execution, peak, and recovery phases of an action. The loss is defined as... : in, This represents a time-series relationship prediction function. and Encoded representation of adjacent time segments, A marker indicating the target temporal relationship between adjacent time segments; Represents the square of the L2 norm; The total loss of multimodal unsupervised pre-training is : in, , , These are the weight coefficients for intra-modal reconstruction loss, cross-modal contrast consistency loss, and temporal structure learning loss, respectively. By minimizing It can learn a stable multimodal shared encoder; the pre-trained encoder parameters are used as initialization parameters for subsequent dual-master task supervised learning of FMA score and MBI scale score, CDDA differential domain adaptation and MAML-Rehab few-shot fast adaptation.
4. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S3, the multimodal spatiotemporal heterogeneous graph neural network ST-HGNN is used as a shared backbone model to uniformly model the intramodal dependencies and cross-modal coupling relationships among EEG, sEMG and IMU, and output deep feature representations suitable for rehabilitation assessment and domain adaptation. The node set of the backbone model consists of three types of heterogeneous nodes: EEG node set, sEMG node set, and IMU node set. Among them, EEG nodes correspond to dual-channel EEG acquisition positions, sEMG nodes correspond to 14 surface electromyography channels, and IMU nodes correspond to 3 inertial measurement units. Thus, the entire heterogeneous map contains a total of 19 nodes, which can simultaneously carry information from three levels: cortex, muscle group, and limb segment. In terms of edge relationship construction, both intramodal edges and intermodal edges are established simultaneously. Intramodal edges are used to represent physiological proximity, temporal continuity, or statistical correlation within the same modality. Intermodal edges are used to represent functional connectivity between different modalities. Through the graph structure, the backbone model can jointly encode cortical activity, muscle activation patterns, and action execution results within the same framework. The node is initially represented as: in, These represent the initial feature representations of the EEG, sEMG, and IMU nodes, respectively. The graph convolution or message passing update process is as follows: in, Represents the normalized adjacency matrix. Indicates the first Layer node representation, This represents the learnable weight matrix of the l-th layer. This represents a non-linear activation function.
5. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S4, FMA score prediction and MBI scale score prediction are set as parallel dual main tasks, and compensation identification, muscle strength grading and CMCI are set as auxiliary tasks. Assume the shared encoder output is The output of the dual main task and auxiliary task is then represented as follows: in, This indicates the FMA prediction result. This indicates the MBI prediction result. This indicates the result of the compensation action recognition. This indicates the results of the muscle strength grading prediction. This represents the prediction results of the cognitive-motor coupling index. The corresponding multi-task prediction head is: in, , , , and These represent the FMA score, MBI scale score, compensatory movement recognition, muscle strength grading, and CMCI prediction head, respectively. A joint loss mechanism is adopted, prioritizing the primary task and enhancing the secondary task. : in, and These represent the prediction losses of FMA and MBI, respectively, using mean squared error, mean absolute error, or Huber loss. The loss for the compensatory action classification is represented by cross-entropy loss; Indicates muscle strength grade loss; This represents the predictive loss of the cognitive-motor coupling index. This indicates the cross-modal graph structure preservation loss; , , , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade.
6. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In steps S5, S6, and S7, a channel-specific differentiated domain adaptation framework (CDDA) is proposed, and adaptation paths are designed according to the different offset mechanisms of EEG, sEMG, and IMU modalities in cross-patient scenarios. For sEMG modalities, covariance alignment constraints are applied to the deep sEMG features of the source and target domains to make their second-order statistics consistent; at the same time, off-diagonal preservation constraints are introduced to maintain the channel-related structures corresponding to the muscle co-activation mode. The deep sEMG features of the source and target domains are respectively , : in, and These represent the number of samples in the source and target domains, respectively. Indicates the dimension of deep features in sEMG; The source domain sEMG eigencovariance matrix is estimated as follows: : The target domain sEMG feature covariance matrix is estimated as follows: : Covariance alignment loss is defined as : in, Let T denote the Frobenius norm, and T denote the transpose. Represents a unit vector; To maintain muscle synergy, off-diagonal retention constraints are introduced. : in, The off-diagonal part of the target domain sEMG covariance matrix is represented. This represents the off-diagonal portion of the reference covariance matrix. To constrain the weights, the values range from [0.1, 10]. The values are determined by optimizing the comprehensive index of muscle synergistic structure similarity and downstream task performance on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade; off-diagonal elements reflect the correlation between different electromyographic channels. The total adaptation loss of the sEMG branch is : For IMU modes, the rotational alignment relationship between the IMU coordinate system and the human anatomical coordinate system is obtained based on the static gravity direction, and then lightweight adversarial adaptation is performed on the residual statistical offset after geometric calibration. Let the direction of gravity estimated from the static segment of the target patient be... The source region reference gravity direction is The optimal rotation matrix is then solved as follows: : in, Represents a set of three-dimensional rotation matrices; Represents the square of the L2 norm; The calibrated target domain IMU acceleration or characteristic is represented as : in, This represents the original target domain IMU acceleration vector or feature vector. This represents the IMU representation after geometric calibration; After geometric calibration, a domain discriminator D and a feature extractor G are introduced, and a lightweight adversarial loss is used to handle the residual statistical offset. : For the EEG modality, the maximum mean difference of multiple kernels was used to align the spatial distribution of the kernels, and different kernel components were weighted according to the importance of rehabilitation assessment indicators in different frequency bands; D s and D t These represent the source and target domain datasets, respectively. / This indicates that sample x originates from D. s / D t Distribution; E[ϕ(x)] represents the average feature of the data after mapping to a high-dimensional space; MMD distance is defined as: in, The source domain D is represented in the reproducing kernel Hilbert space corresponding to kernel function k. s With target domain D t The maximum mean square difference, Represents the kernel mapping function. Represents the regenerated nucleus Hilbert space; Frequency band weighted multi-core MMD loss is : in, These represent the EEG frequency bands associated with exercise preparation and exercise execution, respectively. and Represent the source and target domains in the frequency band respectively. Deep EEG features Indicates the weight of the corresponding frequency band. This indicates the corresponding kernel function; when emphasizing movement intention and preparation information, it is set... Frequency band weight greater than Frequency band weighting; when emphasizing cortical rhythms related to action execution, the weight distribution is adjusted according to the clinical task.
7. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S8, to prevent the original cross-modal physiological connectivity structure from being destroyed after the three modal branches of EEG, sEMG and IMU adapt independently, a cross-modal graph structure preservation mechanism is introduced. In the domain adaptation process, a cross-modal graph Laplacian regularization term is introduced to constrain the relative feature distance or connectivity of key cross-modal node pairs to remain stable before and after adaptation: Among them, L graph This indicates the cross-modal graph structure preservation loss. Denotes the set of cross-modal edges. This represents the cross-modal connectivity weights obtained from PLV, CMCI, or other physiological connectivity metrics. and These represent the adapted node features. This represents the square of the L2 norm.
8. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S9, after completing differential domain adaptation and cross-modal structure preservation, the MAML-Rehab few-sample rapid individualized adaptation mechanism is adopted, enabling the model to complete patient-specific calibration when only a very small number of labeled samples are provided by the target patient. Define a single patient as a meta-task : in, Indicates the first Support set for each patient Indicates the first The query set for each patient; the support set is used for rapid adaptation of the inner loop, and the query set is used for meta-optimization of the outer loop and verification of generalization performance. Label vectors are uniformly defined as : Among them, y i,k This represents the label vector of the i-th patient and the k-th sample. Indicates FMA rating labels, This indicates the MBI scale rating label. Indicates a compensatory action label. Indicates muscle strength grading labels, This indicates a cognitive-motor coupling index label; T indicates transpose. The inner loop update rule of MAML-Rehab is: in, Indicates the first The patient in Step parameters, Indicates the first The patient in Parameters updated during the inner loop step This represents the learning rate of the inner loop. Indicates support set loss; the inner loop performs 5 steps of gradient update; Indicates to Find the gradient; Support set loss adopts a task uncertainty weighted form : in, Indicates the first The loss of each task on the support set Indicates the uncertainty parameter of the task; The outer loop is updated as follows: in, This represents the outer loop learning rate. Indicates the batch size of the meta-task. Indicates based on initialization parameters Patient-specific parameters obtained after support set circulation update, Indicates the query set loss; Indicates to Find the gradient.
9. The multimodal cross-individual fast adaptive evaluation method based on multi-stage collaboration according to claim 1, characterized in that, In step S10, unsupervised pre-training, dual-master task supervised learning, differential domain adaptation, cross-modal structure preservation, and rapid adaptation with few samples constitute a multi-stage collaborative optimization framework. The total loss during the adaptation phase of the differentiated domain is defined as follows: : in, This indicates the sEMG branch covariance alignment and off-diagonal preservation loss. This represents the adversarial adaptation loss of the residual statistical offset in the IMU branch. This indicates the weighted multi-core MMD loss in the EEG tributary band. This indicates the cross-modal graph structure preservation loss. , , and The weighting coefficients are determined by optimizing the comprehensive performance indicators of downstream tasks on the validation set through grid search, including FMA score, MBI scale score, compensatory movement category, and muscle strength grade. The overall joint optimization objective is expressed as: in, Used for unsupervised representation learning This is used for joint supervised learning of FMA scoring and MBI scale scoring in a dual-master task. Used for query set optimization after individualized and rapid adaptation. Used to maintain cross-modal physiological connectivity structures. Indicates the model initialization parameters; Indicates different tasks T i Seeking expectations.
10. A multimodal cross-individual rapid adaptive evaluation system based on multi-stage collaboration, the system being based on the method of any one of claims 1-9, characterized in that, Includes the following modules: Data acquisition module: used to acquire EEG, sEMG and IMU signals, and record action events, scale labels and patient basic information; Preprocessing module: used to perform filtering, artifact suppression, rectification, normalization, attitude initialization, time synchronization, and windowed segmentation; Unsupervised pre-training module: used to train a shared encoder based on historical unlabeled multimodal data; Shared encoder module: used to construct the spatiotemporal heterogeneous graph backbone of ST-HGNN and extract node-level and graph-level deep features; Dual-master-task supervised learning module: used to perform joint training of FMA score and MBI scale score, and simultaneously train auxiliary task head; Differential Domain Adaptation Module: Used to perform sEMG covariance alignment, IMU geometric calibration and adversarial adaptation, and EEG multi-core MMD alignment respectively; Cross-modal structure preservation module: used to maintain the cross-modal physiological connectivity topology of EEG-sEMG based on graph Laplacian regularization; Small-sample rapid adaptation module: used to perform MAML-Rehab individualized calibration based on the target patient's 5-shot support set; Assessment output module: Used to output FMA score, MBI scale score, compensation category, muscle strength grade, CMCI, PLV retention rate, reliability, and optional rehabilitation suggestions.
Citation Information
Patent Citations
Automatic evaluation method of upper limb motion function in stroke based on deep learning
CN109875565A