Femoral head necrosis risk prediction method and system based on multi-modal gait analysis

CN122604355APending Publication Date: 2026-08-21TIANJIN HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610974284.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]然而,目前的此类评估方法存在明显的局限性

Benefits of technology

[0064]上述基于多模态步态分析的股骨头坏死风险预测方法及系统,通过同步获取涵盖关节运动学、足底压力动力学及下肢肌肉电生理的原始时序数据及相应的临床诊断信息,构建初始多模态时序数据集;进而对该数据集进行步态周期分割与时间归一化以统一时序基准,并对其中模态缺失样本进行跨模态数据补全以确保数据完整性;在此基础上,从补全后的数据中提取多维度特征并计算其间的跨模态关联特征,生成深度融合的特征集;随后,利用专为多源数据设计的多分支深度学习模型,以最小化多任务损失函数为目标对该特征集进行训练,生成具备多模态信息融合与理解能力的风险预测模型;最终,通过该模型对受试者的多模态时序数据进行计算,直接输出风险预测结果。该方法实现了对股骨头坏死风险从多模态数据协同感知、到特征深度融合、再到智能模型判别的一体化、自动化评估,有效提升了风险预测的全面性、鲁棒性与临床实用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122604355A_ABST
    Figure CN122604355A_ABST
Patent Text Reader

Abstract

The application relates to a femoral head necrosis risk prediction method and system based on multi-modal gait analysis. The method comprises the following steps: acquiring an initial multi-modal time series data set containing joint kinematics, foot pressure dynamics and lower limb muscle electro-physiological original time series data and corresponding clinical diagnosis information; performing gait cycle segmentation and time normalization to obtain standardized time series data; performing cross-modal data completion on samples with missing modalities to generate a complete multi-modal time series data set; extracting kinematics, dynamics and electromyography features from the complete data, and calculating cross-modal correlation features to generate a multi-modal fusion feature set; based on the feature set, a multi-branch deep learning model is trained by minimizing a multi-task loss function to generate a risk prediction model; and multi-modal time series data of a subject to be evaluated is input into the model to obtain a risk prediction result. The method can realize multi-dimensional and collaborative intelligent evaluation of femoral head necrosis risk, and effectively improve the comprehensiveness and accuracy of risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a method and system for predicting the risk of femoral head necrosis based on multimodal gait analysis. Background Technology

[0002] With the development of sports biomechanics and rehabilitation medicine engineering, gait analysis-based functional assessment techniques for orthopedic diseases are receiving increasing attention. These techniques aim to provide objective evidence for early disease detection, treatment evaluation, and rehabilitation guidance by quantitatively analyzing human movement patterns during walking. In the clinical diagnosis and rehabilitation management of femoral head necrosis, gait analysis technology, as an important supplementary means of functional assessment, is at a critical stage of transitioning from laboratory research to clinical practice.

[0003] Currently, the technical implementation in this field mainly relies on gait analysis systems to collect patients' kinematic or dynamic data. Specifically, existing solutions typically use optical motion capture systems or inertial sensors to acquire kinematic parameters such as joint angles and velocities, or utilize plantar pressure measurement systems to record pressure distribution and ground reaction forces. After preprocessing, the collected data extracts preset gait features (such as stride length, stride speed, and symmetry) and inputs them into classification or regression models for further processing, ultimately outputting a judgment on whether the patient has gait abnormalities or functional status.

[0004] However, current assessment methods have significant limitations. First, existing technologies largely rely on single gait data modalities (such as kinematics or kinetics only), lacking simultaneous acquisition and in-depth fusion analysis of multi-dimensional information such as kinematics, kinetics, and muscle electrophysiology. This "modal isolation" phenomenon results in a single assessment dimension, failing to comprehensively and synergistically reflect the multiple functional impairments that may coexist in patients with avascular necrosis of the femoral head, such as limited joint mobility, abnormal weight-bearing, and muscle compensation, making it difficult to construct an accurate individualized risk profile. Summary of the Invention

[0005] Therefore, it is necessary to provide a method and system for predicting the risk of femoral head necrosis based on multimodal gait analysis that can overcome the limitations of single-modal analysis and achieve effective fusion and collaborative interpretation of multi-dimensional gait biomechanical information, in order to address the above-mentioned technical problems.

[0006] Firstly, this application provides a method for predicting the risk of femoral head necrosis based on multimodal gait analysis, including:

[0007] S1. Obtain the initial multimodal time series dataset of the subjects; the initial multimodal time series dataset is associated with clinical diagnostic information and contains raw time series data of joint kinematics, raw time series data of plantar pressure dynamics and raw time series data of lower limb muscle electrophysiology;

[0008] S2. Perform gait cycle segmentation on the initial multimodal time series dataset to obtain segmented time series data, and perform time normalization on the segmented time series data to obtain standardized time series data.

[0009] S3. Perform cross-modal data completion on samples with missing modalities in the standardized time series data to generate a complete multimodal time series dataset;

[0010] S4. Extract kinematic features, dynamic features, and electromyographic features from the complete multimodal time-series dataset, and calculate cross-modal correlation features based on the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set;

[0011] S5. Based on the multimodal fusion feature set, the model parameters are optimized by minimizing the multi-task loss function through a multi-branch deep learning model to train the model and generate a trained risk prediction model.

[0012] S6. Input the multimodal time series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.

[0013] In one embodiment, S3 includes:

[0014] S31. Identify data samples that are missing at least one modality from standardized time series data, and use the existing modality data in the data samples as conditional data;

[0015] S32. Input the conditional data into the pre-trained generator network, and generate the completed missing modality time series data through the generator network;

[0016] S33. Combine the completed missing modal time series data and conditional data into generated samples;

[0017] S34. Input the generated samples and real complete data samples into the discriminator network to distinguish between real and fake samples, obtain the discrimination probability, and update the generator network parameters based on the discrimination probability through adversarial loss.

[0018] S35. Repeat S32 to S34 until the convergence condition is met, and output the complete multimodal time series dataset consisting of all the completed samples.

[0019] In one embodiment, S4 includes:

[0020] S41. Calculate the symmetry ratio of the bilateral gait parameters in the kinematic features and the bilateral pressure parameters in the dynamic features to generate gait symmetry feature data;

[0021] S42. Perform time delay correlation calculation on the activation time series data of the target muscle in the electromyographic features and the angle time series data of the corresponding joint in the kinematic features to generate muscle and joint coupling feature data.

[0022] S43. Perform instantaneous correlation calculation on the pressure center trajectory data in the dynamic characteristics and the corresponding muscle activation intensity data in the electromyographic characteristics to obtain muscle-pressure synergistic feature data; the muscle-pressure synergistic feature data consists of multiple instantaneous correlation coefficients; the instantaneous correlation coefficients are calculated using the following formula:

[0023]

[0024] in, The instantaneous correlation coefficient, This is the offset of the pressure center. For muscle activation intensity, The preset window half-width, and These are the mean values ​​of the data within the window;

[0025] S44. The gait symmetry feature data, muscle and joint coupling feature data, muscle and pressure coordination feature data, kinematic features, dynamic features and electromyographic features are vector-concatenated to obtain the concatenated feature data.

[0026] S45. Standardize the spliced ​​feature data to generate a multimodal fusion feature set.

[0027] In one embodiment, S5 includes:

[0028] S51. Based on a multi-branch deep learning model, kinematic features, dynamic features and electromyographic features are fused to obtain the final feature representation;

[0029] S52. Input the final feature representation into the classification output head to obtain the stage classification probability;

[0030] S53. Input the final feature representation into the regression output head to obtain the risk score prediction value;

[0031] S54. Calculate the classification loss value based on the phase classification probability and the actual phase label;

[0032] S55. Calculate the regression loss value based on the predicted risk score and the actual risk score;

[0033] S56. Construct a multi-task loss function based on the weighted sum of classification loss and regression loss values;

[0034] S57. Minimize the multi-task loss function using the backpropagation algorithm, update the parameters of the multi-branch deep learning model, and generate a risk prediction model.

[0035] In one embodiment, S51 includes:

[0036] S511. Input the kinematic features into the kinematic sub-network to obtain the high-level kinematic feature vector;

[0037] S512. Input the dynamic features into the dynamic sub-network to obtain the high-level dynamic feature vector;

[0038] S513. Input the electromyographic features into the electromyographic network to obtain the high-level feature vector of electromyography;

[0039] S514. By splicing the high-level kinematic feature vector, the high-level dynamic feature vector, and the high-level electromyographic feature vector, a fused feature vector is obtained.

[0040] S515. Input the fused feature vector into the fully connected fusion layer to obtain the final feature representation.

[0041] In one embodiment, S513 includes:

[0042] S5131. Transform the electromyographic features to obtain a two-dimensional matrix of temporal and channel sequences.

[0043] S5132. Input the temporal and channel two-dimensional matrix into the encoder layer containing a multi-head self-attention mechanism to obtain attention-weighted features;

[0044] S5133. Perform feedforward neural network calculation and layer normalization on the attention-weighted features to obtain enhanced features;

[0045] S5134. Perform global average pooling on the enhanced features to obtain the high-level electromyography feature vector.

[0046] In one embodiment, after S5, the following is also included:

[0047] S61. Obtain new subject data and real clinical outcome labels from new clinical centers to form a validation dataset;

[0048] S62. Input the validation dataset into the risk prediction model for prediction, and output the validation prediction results.

[0049] S63. Calculate and validate the consistency index between the predicted results and the actual clinical outcome labels;

[0050] S64. When the consistency index is lower than the preset threshold, the verification dataset is added to the training data to incrementally train the risk prediction model and generate an optimized risk prediction model.

[0051] In one embodiment, the method further includes:

[0052] S7. Extract the attention weight data and neuron activation data generated by the risk prediction model during the prediction process;

[0053] S8. Calculate the Shapley value of each feature based on the multimodal fusion feature set and risk prediction results to obtain feature contribution data;

[0054] S9. Identify a set of key features whose contribution exceeds a preset threshold based on the feature contribution data;

[0055] S10. Match the key feature set with the clinical knowledge rule base to obtain the matching results; and generate a pathological explanation in natural language based on the matching results; the clinical knowledge rule base is used to store the mapping relationship between features and explanation statements;

[0056] S11. Integrate pathological interpretation with risk prediction results to generate an assessment report.

[0057] Secondly, this application also provides a femoral head necrosis risk prediction system based on multimodal gait analysis, including:

[0058] The multimodal data acquisition module is used to acquire the initial multimodal time series dataset of the subjects; the initial multimodal time series dataset is associated with clinical diagnostic information and includes raw time series data of joint kinematics, raw time series data of plantar pressure dynamics, and raw time series data of lower limb muscle electrophysiology;

[0059] The gait cycle processing module is used to segment the initial multimodal time series dataset into gait cycles, obtain segmented time series data, and perform time normalization on the segmented time series data to obtain standardized time series data.

[0060] The cross-modal data completion module is used to complete cross-modal data for samples with missing modalities in standardized time series data, generating a complete multimodal time series dataset;

[0061] The multimodal feature extraction module is used to extract kinematic features, dynamic features, and electromyographic features from the complete multimodal time-series dataset, and calculate cross-modal correlation features based on the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set;

[0062] The risk prediction model training module is used to train the model based on a multimodal fusion feature set, by minimizing the multi-task loss function through a multi-branch deep learning model to optimize the model parameters and generate a trained risk prediction model.

[0063] The risk prediction module is used to input the multimodal time-series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.

[0064] The aforementioned method and system for predicting the risk of avascular necrosis of the femoral head based on multimodal gait analysis constructs an initial multimodal time-series dataset by simultaneously acquiring raw time-series data covering joint kinematics, plantar pressure dynamics, and lower limb muscle electrophysiology, along with corresponding clinical diagnostic information. This dataset is then segmented into gait cycles and normalized temporally to unify the temporal baseline, and cross-modal data completion is performed on samples with missing modalities to ensure data integrity. Based on this, multi-dimensional features are extracted from the completed data, and cross-modal correlation features are calculated to generate a deeply fused feature set. Subsequently, a multi-branch deep learning model designed specifically for multi-source data is used to train this feature set with the goal of minimizing the multi-task loss function, generating a risk prediction model with multimodal information fusion and understanding capabilities. Finally, this model is used to calculate the risk prediction results directly from the subject's multimodal time-series data. This method achieves an integrated and automated assessment of the risk of avascular necrosis of the femoral head, from multimodal data collaborative perception to deep feature fusion and intelligent model discrimination, effectively improving the comprehensiveness, robustness, and clinical applicability of risk prediction. Attached Figure Description

[0065] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 A flowchart illustrating a method for predicting the risk of femoral head necrosis based on multimodal gait analysis provided by this invention;

[0067] Figure 2 Gait trajectory data graph provided in the data preprocessing of this invention;

[0068] Figure 3 Plantar pressure cloud map provided for data preprocessing in this invention;

[0069] Figure 4 Electromyography signal map in data preprocessing provided by this invention;

[0070] Figure 5 This is a schematic diagram of the structure of a femoral head necrosis risk prediction system based on multimodal gait analysis provided by the present invention. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0072] In one embodiment, such as Figure 1 As shown, a method for predicting the risk of femoral head necrosis based on multimodal gait analysis is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S1 to S6:

[0073] S1. Obtain the initial multimodal time series dataset of the subjects; the initial multimodal time series dataset is associated with clinical diagnostic information and includes raw time series data of joint kinematics, raw time series data of plantar pressure dynamics and raw time series data of lower limb muscle electrophysiology.

[0074] Optionally, raw temporal data of joint kinematics acquired through a motion capture system (Mocap) or an inertial measurement unit (IMU) can be used. This data contains continuous temporal information on the angles, angular velocities, and displacements of key lower limb joints such as the hip and knee joints, reflecting the dynamic laws of joint movement. Figure 2 The gait trajectory data diagram shown intuitively presents the time-series data related to the motion trajectory of key lower limb joints (such as the hip and knee) during walking. The horizontal axis corresponds to the time or motion sequence dimension, and the vertical axis reflects the coordinates, angles, and other motion parameters of the joints, clearly demonstrating the dynamic changes in joint movement; for example... Figure 3 The plantar pressure cloud map shown visualizes the pressure distribution and dynamic changes in different areas of the foot during walking, intuitively demonstrating the differences in plantar pressure intensity and spatial distribution characteristics. High-resolution pressure sensing channels collect raw time-series data of plantar pressure dynamics. This device can record the pressure distribution and changing trends in different areas of the foot during walking, forming continuous time-series data. Surface electromyography (sEMG) is used to collect raw time-series data of lower limb muscle electrophysiology. Specifically, electrodes are precisely attached to target lower limb muscles related to femoral head necrosis, such as the gluteus medius and gluteus maximus, to capture the weak electrical signals generated during muscle contraction in real time and convert them into continuous time-series data. Figure 4As shown, the signal changes of the raw electrophysiological time-series data of the target muscles in the lower limbs (gluteus medius and gluteus maximus) after preprocessing are illustrated. The horizontal axis corresponds to the time dimension, and the vertical axis reflects the fluctuation of electrical signal intensity. Clinical diagnostic information includes key clinical information such as the subject's health status assessment, femoral head necrosis staging results, past orthopedic medical history, and treatment records. By assigning a unique identifier code to each subject, a correlation mapping is established between this unique identifier code and the three types of raw time-series data mentioned above, ensuring a one-to-one correspondence between clinical diagnostic information and multimodal raw data. Figure 4 As shown, the acquisition results of the raw time-series joint kinematics data can be referenced in the gait trajectory data diagram, which fully presents the original characteristics of the data.

[0075] S2. Perform gait cycle segmentation on the initial multimodal time series dataset to obtain segmented time series data, and perform time normalization on the segmented time series data to obtain standardized time series data.

[0076] Optionally, gait cycle segmentation is achieved based on feature inflection points in the raw time-series data of plantar pressure dynamics. Specifically, the point where plantar pressure values ​​change from zero to one is identified as the landing moment of the gait cycle, and the point where pressure values ​​change from one to zero is identified as the take-off moment of the gait cycle. Using these two moments as boundaries, the continuous raw time-series data of joint kinematics, raw time-series data of plantar pressure dynamics, and raw time-series data of lower limb muscle electrophysiology are segmented into time-series data corresponding to a single complete gait cycle. If the feature inflection points of the raw time-series data of plantar pressure dynamics are not obvious, the extreme points of joint angles (such as the maximum angle of hip flexion) in the raw time-series data of joint kinematics can be used as a reference for segmentation calibration to ensure the accuracy of segmentation. Time normalization employs a linear interpolation algorithm. Its core principle is to establish a mapping relationship between the original time axis and the target time axis, and to uniformly map the time series data of each segmented individual gait cycle onto a fixed-length time series. Regardless of the differences in the original gait cycle duration among different subjects, data points are supplemented or normalized through interpolation operations to ensure that the time series data length of all samples remains consistent, thereby eliminating the impact of gait speed differences on data comparability and obtaining standardized time series data.

[0077] S3. Perform cross-modal data completion on samples with missing modalities in the standardized time series data to generate a complete multimodal time series dataset.

[0078] Optionally, all samples in the standardized time-series data are first traversed, and each sample in the raw time-series data of joint kinematics, plantar pressure dynamics, and lower limb muscle electrophysiology is checked for modal missing conditions such as missing data segments or invalid values ​​exceeding a preset threshold. The samples with missing modalities and the specific types of missing modalities are clearly marked. Cross-modal data completion adopts a method based on a combination of memory units and kernel principal component analysis (KPCA). The memory units are used to record the feature prototypes of multimodal data. These feature prototypes are typical expressions of various modal features formed by clustering after extracting nonlinear features from complete multimodal samples using KPCA. The core principle of KPCA is to map the raw data to a high-dimensional feature space through kernel functions, thereby better capturing the nonlinear structural information in the data. During the completion process, complete modal data from the missing modal samples are input into the memory unit. By calculating the similarity between the complete modal features and the feature prototypes in the memory unit, the most valuable complete sample features are matched. Then, based on the potential correlation between cross-modal data, completion data matching the missing modality is generated. The completion data is then filled into the missing position, and finally, a complete multimodal time series dataset is generated.

[0079] S4. Extract kinematic features, dynamic features, and electromyographic features from the complete multimodal time series dataset, and calculate cross-modal correlation features based on the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set.

[0080] Optionally, kinematic features are extracted from raw time-series data of joint kinematics, specifically including mean, variance, peak value, and peak-to-peak value in the time domain, peak power spectral density, dominant frequency, and spectral bandwidth in the frequency domain, and approximate entropy, sample entropy, and fractal dimension in the nonlinear features. These features can comprehensively characterize the stability, regularity, and dynamic changes of joint movement. Dynamic features are extracted from raw time-series data of plantar pressure dynamics, including peak pressure, rate of pressure rise, rate of pressure fall, standard deviation of the pressure center trajectory, and distance traveled. At the same time, the pressure distribution non-uniformity coefficient is calculated to reflect the balance of plantar weight-bearing. Electromyographic features are extracted from raw time-series data of lower limb muscle electrophysiology, covering integrated electromyographic values, root mean square amplitude, spectral centroid, and average power frequency. These features can effectively characterize the activity intensity, fatigue level, and contraction pattern of muscles. Cross-modal correlation features are obtained by calculating the correlation between different modal features. Specifically, the extracted kinematic features, dynamic features, and electromyographic features are constructed as independent feature vectors. The linear correlation between any two modal feature vectors is calculated using the Pearson correlation coefficient, and the nonlinear correlation is calculated using mutual information. After flattening the obtained correlation coefficient matrix and mutual information values, they are concatenated and integrated with the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set. This process fully explores the synergistic change mechanism between gait, muscle, and pressure.

[0081] S5. Based on the multimodal fusion feature set, the model parameters are optimized by minimizing the multi-task loss function through a multi-branch deep learning model to train the model and generate a trained risk prediction model.

[0082] Optionally, the architecture of the multi-branch deep learning model is designed based on the feature characteristics of different modalities. The electromyography (EMG) and kinematics branches employ a structure combining a one-dimensional convolutional neural network (1D-CNN) with a gated recurrent unit (GRU), while the dynamics branch uses a one-dimensional convolutional neural network (1D-CNN) structure. The EMG branch receives temporal data corresponding to EMG features, extracts local spatiotemporal features through convolutional operations of the 1D-CNN, and then captures long-term dependencies in the temporal data through the update and reset gate mechanisms of the GRU, avoiding the gradient vanishing problem. The kinematics branch processes temporal data corresponding to kinematic features using the same network structure, accurately extracting the temporal patterns of joint movements. The dynamics branch extracts local features and spatial distribution features related to plantar pressure through multi-layer convolutional kernels of the 1D-CNN. The output feature vectors of the three branches are directly concatenated at the feature fusion layer to form a fused feature vector that comprehensively reflects multimodal information. This fused feature vector is then input into a fully connected layer for feature depth integration and dimensionality transformation. The multi-task loss function employs a weighted sum of classification and regression loss functions. The classification loss function uses cross-entropy loss to optimize classification tasks for healthy individuals, unilateral femoral head necrosis, and bilateral femoral head necrosis. The regression loss function uses mean squared error loss to optimize regression tasks for abnormality scoring. The weighting coefficients are determined through validation set performance tuning. During model training, all network parameters are iteratively updated using the backpropagation algorithm. After each iteration, the model's accuracy, F1 score, and regression mean squared error on the validation set are calculated. Training stops when the validation set performance remains stable for several consecutive iterations and no longer improves, generating a trained risk prediction model. The training process fully utilizes the constructed multimodal database to ensure the model learns specific feature patterns related to femoral head necrosis.

[0083] S6. Input the multimodal time series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.

[0084] Optionally, for the multimodal time-series dataset of the subjects to be assessed, gait cycle segmentation is first performed based on inflection points of plantar pressure features or extreme points of joint angles. Time normalization is achieved through linear interpolation to obtain standardized time-series data. If modalities are missing, a method combining memory units and KPCA is used to complete the cross-modal data. Then, kinematic features, kinetic features, and electromyographic features are extracted, and cross-modal correlation features are calculated to generate a multimodal fusion feature set to be assessed with the same format and feature dimensions as the training data. This multimodal fusion feature set to be assessed is input into the trained risk prediction model. The model extracts and processes features in a targeted manner through electromyographic, kinematic, and kinetic branches, respectively. The feature fusion layer achieves comprehensive integration of multimodal features, followed by deep computation through a fully connected layer. Finally, the multi-task output layer outputs the results. The classification result clarifies the health status or femoral head necrosis laterality (unilateral / bilateral) of the subject to be assessed, while the regression result provides an abnormality score. Integrating the two types of results yields a complete risk prediction result, directly addressing the clinical need for judging the laterality and severity of abnormalities.

[0085] The aforementioned method for predicting the risk of femoral head necrosis based on multimodal gait analysis comprehensively addresses the problems of modal isolation, data gaps, insufficient quantification, and poor model specificity by acquiring multimodal raw time-series data and performing segmentation and normalization, missing data completion, feature fusion, and dedicated model training. This achieves accurate classification and quantitative assessment of the risk of femoral head necrosis, providing objective and practical data for early clinical prevention, diagnosis, and rehabilitation, and effectively improving assessment efficiency and the specificity of treatment guidance.

[0086] In one embodiment, S3 includes:

[0087] S31. Identify data samples that are missing at least one modality from standardized time series data, and use the existing modal data in the data samples as conditional data.

[0088] Optionally, by traversing all samples in the standardized time-series data, each sample is checked to see if it simultaneously contains three complete modal data types: raw joint kinematics time-series data, raw plantar pressure dynamics time-series data, and raw lower limb muscle electrophysiology time-series data. The verification logic is to detect whether each modal data type has a continuous invalid value ratio exceeding a preset proportion, or a data length that does not reach the fixed length after time normalization. If any modal data type satisfies any of the above conditions, the sample is determined to be a data sample missing at least one modality. The verification of raw joint kinematics time-series data can be referenced as follows: Figure 2 The gait trajectory data plot shown in the figure has a standard format, and the original time-series data of plantar pressure dynamics must conform to the following: Figure 3 The effective data characteristics corresponding to the plantar pressure cloud map shown, the raw time-series data of lower limb muscle electrophysiology must be consistent with, as shown in the figure. Figure 4 The valid signal morphology of the electromyography signal diagrams shown is consistent. Conditional data refers to one or more types of modal data that have been verified to be complete and valid within the missing modal data sample. For example, if a sample is missing raw time-series data of lower limb muscle electrophysiology, but has complete and valid raw time-series data of joint kinematics and plantar pressure dynamics, then these two types of complete data are the conditional data of that sample. Conditional data must maintain the same time axis and data format as the standardized time-series data.

[0089] S32. Input the conditional data into the pre-trained generator network, and generate the completed missing modality time series data through the generator network.

[0090] Optionally, the pre-trained generator network employs an architecture combining a one-dimensional convolutional neural network and a Long Short-Term Memory (LSTM) network. Its pre-training process is based on complete multimodal temporal data. The training objective is to enable the generator network to learn the potential correlations between different modalities, particularly the collaborative change mechanisms between gait, muscle, and stress, thereby overcoming the modal isolation limitations of existing technologies. During pre-training, any one or two classes of modalities from the complete multimodal temporal data are used as input, with the corresponding other modalities as labels. The network parameters are optimized by minimizing the mean squared error loss function between the predicted value and the label, enabling the generator network to infer missing modal data based on known modal data. After inputting conditional data into the pre-trained generator network, the network first extracts local spatiotemporal features of the conditional data using a 1D-CNN, then captures long-term dependencies of the temporal data using an LSTM, and finally outputs completed missing modal temporal data with dimensions and temporal length completely consistent with the missing modal data. The time axis of the completed data is strictly aligned with the conditional data to ensure temporal consistency.

[0091] S33. Combine the completed missing modal time series data and conditional data into a generated sample.

[0092] Optionally, according to the correspondence between modal data types, the completed missing modal time-series data is combined with the conditional data to form a structurally complete generated sample. During the combination process, it is necessary to ensure that all types of modal data are completely synchronized in the time dimension; that is, each time point of the completed missing modal time-series data must match the corresponding time point in the conditional data, maintaining the same structural format as the actual complete data sample. For example, if the conditional data includes raw time-series data of joint kinematics and plantar pressure dynamics, and the completed missing modal time-series data is raw time-series data of lower limb muscle electrophysiology, then the three types of data are concatenated in parallel along the time axis, so that the generated sample simultaneously contains all three types of modal data.

[0093] S34. Input the generated samples and real complete data samples into the discriminator network to distinguish between real and fake samples, obtain the discrimination probability, and update the generator network parameters based on the discrimination probability through adversarial loss.

[0094] Optionally, the discriminator network adopts a 1D-CNN architecture. Its core function is to learn the distribution characteristics of real, complete multimodal time-series data, thereby distinguishing whether the input sample is a real, complete data sample or a generated sample. Generated samples and real, complete data samples are input into the discriminator network respectively. The network extracts the multimodal fusion features of the samples through multi-layer convolutional operations, and then maps the features to a discrimination probability between 0 and 1 through fully connected layers. A discrimination probability close to 1 indicates that the sample is judged as a real, complete data sample, while a probability close to 0 indicates that it is judged as a generated sample. The adversarial loss function uses the binary cross-entropy loss function, which calculates the predicted loss of the generated samples based on the discrimination probability. The loss value is passed to the generator network through the backpropagation algorithm, thereby updating the weight parameters of the generator network. The update allows the generator network to adjust its feature extraction and data generation strategies based on the loss feedback, making the generated missing modality time-series data closer to the distribution characteristics of real data, gradually improving the authenticity of the generated samples.

[0095] S35. Repeat S32 to S34 until the convergence condition is met, and output the complete multimodal time series dataset consisting of all the completed samples.

[0096] Optionally, the process from S32 to S34 is repeatedly executed, using a preset convergence condition as the termination criterion. The convergence condition is set as follows: the value of the adversarial loss function remains stable for multiple consecutive rounds, i.e., the change in the loss value between two adjacent rounds is less than a preset threshold, or the average discrimination probability of the generated samples is close to 0.5. This indicates that the discriminator network can no longer effectively distinguish between the generated samples and the real complete data samples, and the completed data generated by the generator network fully conforms to the distribution pattern of the real data. During each iteration, the parameters of the generator network are optimized according to the adversarial loss, and the authenticity and rationality of the completed data are gradually improved. When the convergence condition is met, the iteration process stops. At this time, all data samples missing at least one modality have completed the missing modality data through the optimized generator network. All the completed samples are integrated with the original complete samples to form a complete multimodal time series dataset. This complete multimodal time series dataset covers three complete modalities: joint kinematics, plantar pressure dynamics, and lower limb muscle electrophysiology, with a unified data format and consistent time series.

[0097] In the above embodiments, by accurately identifying missing modality samples and using adversarial training and iterative optimization between the pre-trained generator and discriminator, complete data that conforms to the real modality association rules is generated, which effectively solves the pain point of missing multimodal data in clinical practice and improves the completeness and robustness of the dataset.

[0098] In one embodiment, S4 includes:

[0099] S41. Calculate the symmetry ratio of the bilateral gait parameters in the kinematic features and the bilateral pressure parameters in the dynamic features to generate gait symmetry feature data.

[0100] Optionally, it is first clarified that the bilateral gait parameters in the kinematic features include parameters reflecting the movement state of both lower limbs, such as stride length, stride width, peak hip flexion angle, and peak knee extension angle. These parameters are all extracted from the original time-series data of joint kinematics in the complete multimodal time-series dataset. The bilateral pressure parameters in the dynamic features include the peak pressure, pressure rise rate, and pressure distribution non-uniformity coefficient of different regions of the bilateral soles, which are derived from the original time-series data of plantar pressure dynamics. The symmetry ratio calculation adopts the relative ratio method of corresponding parameters on both sides. Specifically, for each set of bilateral symmetry parameters, the ratio of the left parameter to the right parameter is calculated. If the right parameter is zero, the ratio of the right parameter to the left parameter is taken. At the same time, in order to avoid the influence of extreme values, the calculated ratio is truncated, and the ratio that exceeds the reasonable range is corrected to the range. Finally, the symmetry ratios of all bilateral gait parameters and bilateral pressure parameters together constitute gait symmetry feature data. This process can quantify the balance of bilateral lower limb movement and weight-bearing, accurately capture the unilateral compensatory gait characteristics that may occur in patients with femoral head necrosis, and make up for the deficiency that a single modality cannot reflect bilateral coordination differences.

[0101] S42. Perform time delay correlation calculation on the activation time series data of the target muscle in the electromyographic features and the angle time series data of the corresponding joint in the kinematic features to generate muscle and joint coupling feature data.

[0102] Optionally, for target muscles such as the gluteus medius and gluteus maximus, the activation timing data are time-series sequences of integral electromyography (EMG) values ​​and root mean square amplitudes from the EMG characteristics; for corresponding joints such as the hip joint, the angle timing data are the sequence of hip joint flexion, extension, adduction, and abduction angle changes from the kinematic characteristics. The core of time delay correlation calculation is to explore the temporal correlation between muscle activation and joint movement. First, a reasonable time delay range is set, which covers common time intervals where muscle activation may precede or lag joint movement. Then, for each time delay value, the Pearson correlation coefficient is used to calculate the linear correlation between the EMG activation timing data and the joint angle timing data at that delay. After traversing all time delay values, the maximum correlation coefficient is selected as the core indicator, and the time delay value corresponding to the maximum correlation coefficient is recorded. Both are used together as muscle-joint coupling feature data. This coupling feature data can effectively reflect the abnormal muscle-joint coordinated movement caused by joint dysfunction in patients with avascular necrosis of the femoral head, which meets the core requirements of multimodal collaborative analysis.

[0103] S43. Perform instantaneous correlation calculations on the pressure center trajectory data in the dynamic characteristics and the corresponding muscle activation intensity data in the electromyographic characteristics to obtain muscle-pressure synergistic feature data; the muscle-pressure synergistic feature data consists of multiple instantaneous correlation coefficients; the instantaneous correlation coefficients are calculated using the following formula:

[0104]

[0105] in, The instantaneous correlation coefficient, This is the offset of the pressure center. For muscle activation intensity, The preset window half-width, and These are the mean values ​​of the data within the window.

[0106] Optionally, the pressure center trajectory data in the dynamic features is a two-dimensional coordinate time sequence of the plantar pressure center during walking, clearly showing the movement path of the pressure center; the activation intensity data of the corresponding muscles in the electromyographic features are the root mean square amplitude time sequence of the gluteus medius and gluteus maximus muscles, derived from the raw time sequence data of lower limb muscle electrophysiology. Instantaneous correlation calculation uses the sliding window method, with a preset window half-width. The total length of the window is half the length of the sliding window. The window slides along the time series data in time steps, and each time it slides to a target time... Then extract the data before and after that point in time. Pressure center offset at each time step and muscle activation intensity First, calculate the value within the window. mean and mean Then calculate each one in the window separately. Moment and The product of The square of The square of the result is used to sum the results of these three types to obtain two terms in the numerator and denominator. Finally, the instantaneous correlation coefficient corresponding to time t is calculated by dividing the numerator by the square root of the denominator. As the window slides along the time sequence, each time point corresponds to an instantaneous correlation coefficient. All instantaneous correlation coefficients are arranged in chronological order to form muscle and stress synergistic feature data, accurately capturing the real-time correlation between stress changes and muscle activation, and fully exploring the synergistic change mechanism of gait-muscle-stress.

[0107] S44. The gait symmetry feature data, muscle and joint coupling feature data, muscle and pressure coordination feature data, kinematic features, dynamic features, and electromyographic features are vector-concatenated to obtain the concatenated feature data.

[0108] Optionally, the vector format of each feature data is first defined. Kinematic features, dynamic features, and electromyography (EMG) features are all one-dimensional feature vectors. Gait symmetry feature data, muscle-joint coupling feature data, and muscle-pressure coordination feature data are also processed into one-dimensional feature vectors, and the number of samples for all vectors remains the same. Vector concatenation adopts a sample-by-sample concatenation method. For each sample, the kinematic feature vector, dynamic feature vector, and EMG feature vector are concatenated sequentially. Then, the gait symmetry feature data vector, muscle-joint coupling feature data vector, and muscle-pressure coordination feature data vector are concatenated afterward to form a one-dimensional feature vector with higher dimensionality and more comprehensive information. During the concatenation process, the sample order of each feature vector is strictly guaranteed to correspond one-to-one to avoid misalignment of samples and features. Finally, the concatenated feature vectors of all samples together constitute the concatenated feature data. This concatenation method achieves deep fusion of the original single-modal features and cross-modal correlated features.

[0109] S45. Standardize the spliced ​​feature data to generate a multimodal fusion feature set.

[0110] Optionally, the standardization process employs the Z-score standardization method, the core of which is to convert features with different dimensions and numerical ranges into standard normal distribution data with a mean of 0 and a standard deviation of 1. First, all feature dimensions of the concatenated feature data are iterated through, calculating the mean and standard deviation of each feature dimension across all samples. Then, for each feature dimension of each sample, the original value of that feature dimension is subtracted from its mean, and then divided by its standard deviation to obtain the standardized feature value. If the standard deviation of a feature dimension is 0, meaning all samples have the same value in that dimension, all feature values ​​for that dimension are uniformly set to 0 to avoid calculation errors. After standardization, the numerical distributions of all feature dimensions are on the same order of magnitude, forming a multimodal fusion feature set. This multimodal fusion feature set retains the original kinematic, dynamic, and electromyographic feature information while integrating cross-modal correlation features.

[0111] In the above embodiments, by calculating three types of cross-modal correlation features—gait symmetry, muscle-joint coupling, and muscle-pressure synergy—and concatenating and standardizing them with the original single-modal features, the synergistic relationship between multimodal data is fully explored, solving the modal isolation problem, enriching the feature dimensions and information density, and generating a multimodal fusion feature set that provides accurate and comprehensive feature support for the femoral head necrosis risk prediction model, effectively improving the comprehensiveness and accuracy of model evaluation.

[0112] In one embodiment, S5 includes:

[0113] S51. Based on a multi-branch deep learning model, kinematic features, dynamic features, and electromyographic features are fused to obtain the final feature representation.

[0114] Optionally, the electromyography (EMG) branch employs a one-dimensional convolutional neural network (CNN) combined with a gated recurrent unit (GRU) structure to receive EMG features and extract them. The 1D-CNN traverses the temporal dimension of EMG features through the sliding convolutional kernels, capturing local spatiotemporal features. The GRU, through the synergistic effect of update and reset gates, memorizes the long-term temporal dependencies of EMG signals and filters redundant information. The kinematics branch also uses a 1D-CNN+GRU architecture to process kinematic features, accurately mining the dynamic changes and key temporal features in joint motion parameters. The dynamics branch uses a 1D-CNN structure, focusing on local features and spatial distribution information in dynamics features, and gradually abstracting features through multi-layer convolutional operations. After the three branches output their respective high-level feature vectors, they are directly concatenated in the feature fusion layer, integrating the high-level representations of EMG, kinematics, and dynamics features into a higher-dimensional and more comprehensive vector, which is the final feature representation.

[0115] S52. Input the final feature representation into the classification output head to obtain the stage classification probability.

[0116] Optionally, the classification output head consists of multiple fully connected layers and a Softmax function. The fully connected layers map the final feature representation to a feature dimension corresponding to the stage of femoral head necrosis. Femoral head necrosis is clinically classified into early, intermediate, and late stages; therefore, the output dimension of the fully connected layers is set to 3. The multiple fully connected layers gradually extract key stage-related information from the final feature representation through linear transformations of the weight matrix and nonlinear transformations of the activation function, mapping high-dimensional features to a low-dimensional classification space. Subsequently, the Softmax function normalizes the output of the fully connected layers, converting it into probability values ​​between 0 and 1. The sum of all probability values ​​is 1, and each probability value corresponds to the likelihood of the final feature representation belonging to a specific stage of femoral head necrosis. These probability values ​​collectively constitute the stage classification probability.

[0117] S53. Input the final feature representation into the regression output head to obtain the risk score prediction value.

[0118] Optionally, the regression output head employs a single-layer fully connected layer structure. The purpose is to map the final feature representation into a single-dimensional continuous numerical value, i.e., the risk score prediction value. The fully connected layer learns weight parameters to comprehensively calculate multimodal collaborative information and cross-modal correlation features in the final feature representation, transforming high-dimensional features into continuous values ​​that quantify the degree of risk of femoral head necrosis. The range of the risk score prediction value is consistent with clinically established risk scoring standards, directly reflecting the severity and progression risk of femoral head necrosis in the subject; for example, a higher score indicates a higher risk and a more severe condition.

[0119] S54. Calculate the classification loss value based on the phase classification probability and the actual phase label.

[0120] Optionally, the true stage labels are category labels constructed based on clinical diagnostic results, represented using one-hot encoding. For example, early stage corresponds to a vector with only the first dimension set to 1 and the rest to 0, mid-stage corresponds to a vector with only the second dimension set to 1 and the rest to 0, and late stage corresponds to a vector with only the third dimension set to 1 and the rest to 0. The construction of the labels strictly follows the clinical staging standards for avascular necrosis of the femoral head, ensuring a one-to-one correspondence with the dimensions of the staging classification probabilities. The classification loss value is calculated using the cross-entropy loss function. The core principle of this function is to measure the difference between the predicted distribution represented by the staging classification probabilities and the true distribution represented by the true stage labels. Specifically, the calculation process involves multiplying the true label of each category by the logarithm of the corresponding staging classification probability, summing the results, and then taking the negative value to obtain the classification loss value for a single sample. The average classification loss of all samples is the overall classification loss value. This loss function can effectively reflect the prediction error of the model in the staging classification task of avascular necrosis of the femoral head.

[0121] S55. Calculate the regression loss value based on the predicted risk score and the actual risk score.

[0122] Optionally, the true risk score is a quantitative indicator based on the subject's clinical examination results, medical history, and multimodal data characteristics. It is assessed by clinicians according to standardized criteria to ensure objectivity and consistency, and its value accurately reflects the actual risk level and functional impairment of the subject's femoral head necrosis. The regression loss is calculated using the mean squared error loss function. First, the difference between the predicted risk score and the true risk score for each sample is calculated. This difference is then squared to amplify the impact of larger errors. Finally, the average of the squared differences across all samples is calculated to obtain the overall regression loss. The mean squared error loss function effectively measures the degree of bias in continuous value prediction, ensuring the model's accuracy in risk score prediction tasks. Together with the classification loss value, it supports the optimization goal of multi-task learning.

[0123] S56. Construct a multi-task loss function based on the weighted sum of classification loss and regression loss values.

[0124] Optionally, the multi-task loss function is constructed using a weighted sum of classification and regression loss values. The classification and regression loss weights are preset hyperparameters, determined through multiple experiments on the validation set. The tuning principle is to achieve a balanced and optimal performance for the model on both the period classification and risk scoring prediction tasks, avoiding the loss from a single task dominating the model training process. The classification loss value is multiplied by its classification weight, and the regression loss value is multiplied by its regression weight; the two results are then added together to obtain the final value of the multi-task loss function.

[0125] S57. Minimize the multi-task loss function using the backpropagation algorithm, update the parameters of the multi-branch deep learning model, and generate a risk prediction model.

[0126] Optionally, the core principle of the backpropagation algorithm is based on the chain rule. Starting from the multi-task loss function, it calculates the gradient of the loss with respect to all trainable parameters in the multi-branch deep learning model layer by layer. Trainable parameters include the convolutional kernel weights of each branch, the hidden layer weights of GRU, the weights and biases of fully connected layers, etc. The gradient value reflects the degree of influence of parameter changes on the loss function. After obtaining the gradients of all parameters, an optimizer (such as the Adaptive Moment Estimation optimizer Adam) is used to update the model parameters. The update logic is to subtract the product of the learning rate and the gradient of each parameter. The learning rate is a hyperparameter that controls the step size of parameter updates and needs to be dynamically adjusted according to the loss changes during training. During model training, the process of feature extraction, loss calculation, backpropagation, and parameter update is repeated in each iteration. The model's staging classification accuracy, F1-score, and regression mean squared error on the validation set are continuously monitored. When the performance on the validation set remains stable for several consecutive iterations and no longer improves, training is stopped. The model obtained at this time is the trained risk prediction model.

[0127] In the above embodiments, multimodal features are extracted and fused through a multi-branch model, and classification and regression tasks are balanced by using a multi-task loss function. Parameters are optimized through backpropagation, achieving accurate classification and risk quantification scoring of femoral head necrosis staging. This solves the problems of modal isolation and single assessment, providing comprehensive and reliable risk prediction results for clinical practice and helping to improve the pertinence of diagnosis, treatment and rehabilitation guidance.

[0128] In one embodiment, S51 includes:

[0129] S511. Input the kinematic features into the kinematic sub-network to obtain the high-level kinematic feature vector.

[0130] Optionally, the kinematics subnetwork adopts an architecture of one-dimensional convolutional neural networks and gated recurrent units (GRUs) in series, specifically adapted to the temporal characteristics of kinematic features and joint motion patterns. Kinematic features include parameters extracted from the original temporal data of joint kinematics, such as stride length and joint angles. The network first processes the kinematic features through multiple 1D-CNN convolutional layers. The convolutional kernels slide along the temporal dimension, capturing the local spatiotemporal features of joint motion segment by segment. After each convolutional layer, a ReLU activation function is applied to introduce a nonlinear transformation, enhancing the feature representation capability. Max pooling layers are interspersed between convolutional layers to reduce data dimensionality and computational complexity while preserving key features. The feature maps output from the convolutional layers are flattened and then input into the GRU layer. The GRU layer controls the retention and forgetting of historical information through an update gate and adjusts the influence weight of the current input through a reset gate, effectively capturing the long-term temporal dependencies of joint motion in the kinematic features, filtering redundant noise information, and finally outputting a high-level kinematic feature vector with fixed dimensions and condensed information. This vector fully extracts the key joint motion patterns related to femoral head necrosis.

[0131] S512. Input the dynamic features into the dynamic sub-network to obtain the high-level dynamic feature vector.

[0132] Optionally, the dynamics subnetwork adopts a pure 1D-CNN architecture to specifically process the spatial distribution and local variation information of dynamic features. Dynamic features include parameters extracted from the original time-series data of plantar pressure dynamics, such as peak plantar pressure and pressure center trajectory. The network consists of multiple 1D-CNN convolutional layers, batch normalization layers, and pooling layers. The kernel size and number of the convolutional layers are adaptively set according to the dimension of the dynamic features. Sliding convolution operations are used to extract local features of plantar pressure in different time segments, as well as spatial correlation information of pressure distribution. The batch normalization layer standardizes the output of each convolutional layer, accelerating network training convergence and improving model stability. The pooling layer uses an average pooling strategy to further compress the feature dimension while preserving the global pressure change trend. Through iterative processing of multiple convolutions and pooling, the dynamic features are gradually abstracted into high-level semantic information, ultimately outputting a high-level dynamic feature vector that can characterize abnormal plantar pressure patterns, accurately capturing core features such as weight-bearing imbalance that may occur in patients with avascular necrosis of the femoral head.

[0133] S513. Input the electromyographic features into the electromyographic network to obtain the high-level feature vector of electromyography.

[0134] Optionally, the electromyography (EMG) network adopts the same 1D-CNN+GRU architecture as the kinesiology subnetwork, specifically adapted to the characteristics of raw temporal data of lower limb muscle electrophysiology. EMG features include parameters extracted from the surface EMG signals of the gluteus medius and gluteus maximus muscles, such as integral EMG values ​​and root mean square amplitude. The network first processes the EMG features through 1D-CNN convolutional layers. The convolutional kernels focus on local temporal fluctuations in the EMG signal, capturing the electrical signal changes during muscle activation and relaxation. After convolution, a ReLU activation function is used to introduce nonlinearity, followed by pooling layers to filter key features and reduce data dimensionality. Subsequently, the flattened feature sequence is input into the GRU layer. The GRU layer accurately captures the long-term temporal dependencies of the EMG signal through a gating mechanism, such as the duration and interval patterns of muscle activation, which are muscle compensatory features related to femoral head necrosis. This avoids feature distortion caused by noise interference in the EMG signal, ultimately outputting a high-level EMG feature vector that characterizes abnormal muscle activity patterns, fully reflecting the functional state of target muscles such as the gluteus medius and gluteus maximus related to femoral head necrosis.

[0135] S514. By splicing the high-level kinematic feature vector, the high-level dynamic feature vector, and the high-level electromyographic feature vector, a fused feature vector is obtained.

[0136] Optionally, the dimensionality consistency of the kinematic, dynamic, and electromyographic high-level feature vectors is first checked to ensure that the number of samples in the three vectors is completely matched and that the dimension of the feature vector corresponding to each sample is fixed. The concatenation operation adopts a method of concatenation in the order of feature dimensions. Specifically, for each sample, all elements of the kinematic high-level feature vector are arranged in their original order, followed by all elements of the dynamic high-level feature vector, and finally all elements of the electromyographic high-level feature vector are arranged in the final order to form a new vector with a dimension equal to the sum of the dimensions of the three vectors. This vector is the fused feature vector.

[0137] S515. Input the fused feature vector into the fully connected fusion layer to obtain the final feature representation.

[0138] Optionally, the fully connected fusion layer consists of multiple fully connected layers and dropout layers. The number of neurons in the fully connected layers gradually decreases from input to output, forming a progressive structure of feature compression and deep integration. After the fused feature vector is input into the fully connected fusion layer, it first undergoes a linear transformation in the first fully connected layer, mapping the high-dimensional fused features to an intermediate-dimensional feature space. Subsequently, a ReLU activation function is applied to introduce non-linearity, enhancing the non-linear expressive power of the features. The intermediate fully connected layers repeat the above linear transformation and non-linear activation process, further refining the key correlation information in the fused features. The dropout layer randomly deactivates some neurons, effectively preventing model overfitting and improving generalization ability. The final fully connected layer maps the intermediate-dimensional features to a fixed-dimensional feature space, outputting a final feature representation with unified dimensions and highly condensed information.

[0139] In the above embodiments, high-level information of multimodal features is extracted through dedicated sub-networks, and then deeply fused with vector concatenation and fully connected layers to fully break down modal isolation, extract multimodal collaborative correlation features, and generate a final feature representation that is comprehensive and highly targeted. This effectively solves the limitations of single-modal analysis, provides high-quality feature support for the femoral head necrosis risk prediction model, and helps improve the accuracy and comprehensiveness of the model's prediction.

[0140] In one embodiment, S513 includes:

[0141] S5131. Perform dimensional transformation on the electromyographic features to obtain a two-dimensional matrix of time sequence and channels.

[0142] Optionally, the electromyographic features are derived from the surface electromyography (sEMG) signals of the gluteus medius and gluteus maximus muscles, including parameters extracted from the raw temporal data of lower limb muscle electrophysiology such as integrated electromyographic values, root mean square amplitude, and spectral centroid. The core of the dimensionality transformation is to reconstruct the initial one-dimensional electromyographic feature vector into a temporal and channel two-dimensional matrix adapted for subsequent encoder layer processing. The initial one-dimensional electromyographic feature vector usually stacks all feature values ​​in temporal order, without clearly distinguishing the relationship between feature type and time dimension. First, the electromyographic features are divided into different channels according to type, for example, the integrated electromyographic value is one channel and the root mean square amplitude is another channel, with each channel corresponding to a complete temporal sequence. Then, with time step as the vertical dimension and feature channel as the horizontal dimension, the temporal sequences of all channels are arranged in parallel to form a temporal and channel two-dimensional matrix of (temporal length, number of channels), where the temporal length corresponds to the number of gait cycle temporal points after time normalization, and the number of channels is consistent with the number of electromyographic feature types. This transformation clearly separates the temporal structure of the electromyographic features from the feature type information, while also making them mutually related.

[0143] S5132. Input the temporal and channel two-dimensional matrix into the encoder layer containing the multi-head self-attention mechanism to obtain attention-weighted features.

[0144] Optionally, the core principle of the multi-head self-attention mechanism is to simultaneously capture the temporal dependencies of electromyographic features at different scales and time spans using multiple parallel attention heads. Compared to single-head self-attention, it can more comprehensively uncover the dynamic correlation patterns of muscle activation signals. The encoder layer takes a two-dimensional matrix of temporal and channel data as input. First, it performs a linear projection operation on this matrix to generate a query vector, a key vector, and a value vector, all with consistent dimensions and adapted to the number of channels. Then, the query vector, key vector, and value vector are each divided into a predetermined number of sub-vectors, each corresponding to an attention head. Each attention head independently calculates its attention weight: a similarity matrix is ​​obtained by performing a dot product operation between the query vector and the key vector, and after normalization using the Softmax function, the attention weight is obtained. This weight reflects the importance of each temporal point to other temporal points. The attention weight is then weighted and summed with the corresponding value vector to obtain the output feature of each attention head. Finally, the output features of all attention heads are concatenated along the channel dimension, and a single linear transformation is used to integrate multi-scale correlation information, generating attention-weighted features.

[0145] S5133. Perform feedforward neural network calculation and layer normalization on the attention-weighted features to obtain enhanced features.

[0146] Optionally, the feed-forward neural network (FFN) consists of two fully connected layers and a ReLU activation function. Layer normalization is used to standardize the feature distribution, improving the network's training stability and convergence speed. First, the attention-weighted features are subjected to layer normalization, calculating the mean and variance of each sample in the channel dimension. This transforms all feature values ​​into standardized data with a mean of 0 and a variance of 1, effectively mitigating the vanishing gradient problem and ensuring that the features maintain a reasonable numerical range during computation. Then, the normalized features are input into the feed-forward neural network. The first fully connected layer maps the features to a higher-dimensional feature space through a linear transformation, while the ReLU activation function introduces a non-linear transformation, filtering and enhancing key correlated features and suppressing redundant noise. The second fully connected layer maps the high-dimensional features back to the same dimension as the attention-weighted features, completing the non-linear enhancement of the features.

[0147] S5134. Perform global average pooling on the enhanced features to obtain the high-level electromyography feature vector.

[0148] Optionally, the core function of Global Average Pooling is to globally aggregate information from the temporal dimension of the augmented features, condensing them into a high-level electromyographic feature vector with fixed dimensions and dense information, while avoiding the risk of overfitting. For augmented features in the form of a two-dimensional matrix of temporal and channel data, the average value of the feature values ​​at all temporal points is calculated one by one along the channel dimension. That is, each channel corresponds to a global average feature value, which comprehensively reflects the overall activation level and trend of the channel feature throughout the entire gait cycle. For example, if a channel corresponds to the root mean square amplitude feature of the gluteus medius muscle, the globally averaged value represents the average activation intensity of the muscle during the complete gait cycle, effectively characterizing the muscle's functional state. The global average feature values ​​of all channels are arranged in channel order to form a one-dimensional high-level electromyographic feature vector. The dimension of this high-level electromyographic feature vector is consistent with the number of channels, which not only retains the core information of each type of electromyographic feature but also eliminates redundant details in the temporal dimension, achieving a high degree of feature condensation.

[0149] In the above embodiments, the electromyographic feature structure is reconstructed by dimensional transformation, temporal dependence is captured by multi-head self-attention, feature expression is enhanced by feedforward network and layer normalization, and core information is condensed by global average pooling. The dynamic laws and key features of electromyographic signals are mined out. The generated high-level feature vector information of electromyography is accurate and dimensionally adapted, effectively supporting the multimodal fusion process, improving the ability of the femoral head necrosis risk prediction model to identify muscle function abnormalities, and helping the model to predict accuracy and comprehensiveness.

[0150] In one embodiment, after S5, the following is also included:

[0151] S61. Obtain new subject data and real clinical outcome labels from new clinical centers to form a validation dataset.

[0152] Optionally, the new subject data includes raw time-series data on joint kinematics, plantar pressure dynamics, and lower limb muscle electrophysiology, as well as baseline information, from subjects at the new clinical center. The data acquisition methods are consistent with the initial multimodal time-series dataset: joint kinematic data is acquired via motion capture systems or inertial measurement units; plantar pressure dynamics data is acquired via high-resolution pressure-sensing walkways; and electromyography (EMG) data is acquired from the gluteus medius and gluteus maximus muscles. The subject baseline information includes basic information related to femoral head necrosis, such as age, gender, and past medical history, ensuring the completeness of the data dimensions. The real clinical outcome label consists of two parts: first, a femoral head necrosis staging label, assessed by clinicians according to unified clinical diagnostic criteria (such as the ARCO staging system), classifying it into early, intermediate, and late stages; second, a real risk score, quantitatively assessed by a professional medical team based on the subject's clinical examination results, degree of functional impairment, and disease progression trend. After the raw time-series data is collected, the same preprocessing procedure as the initial training data is performed on the raw time-series data of the new subjects, including gait cycle segmentation, time normalization, cross-modal data completion (if there are missing modalities), and multimodal fusion feature set generation. The processed feature data is then associated with the real clinical outcome labels to form a validation dataset. This validation dataset is specifically used to test the generalization ability of the risk prediction model on new clinical center data.

[0153] S62. Input the validation dataset into the risk prediction model for prediction, and output the validation prediction results.

[0154] Optionally, the feature data in the validation dataset is first validated to ensure that the dimensions and feature types of the multimodal fusion feature set are completely consistent with the input format during model training, avoiding the impact of data format mismatch on the prediction results. Then, the validated multimodal fusion feature set is input into the risk prediction model. The model extracts the corresponding features at a high level through kinematics subnetworks, dynamics subnetworks, and myoelectronic networks, respectively, and obtains the final feature representation through feature concatenation and a fully connected fusion layer. The final feature representation is simultaneously input into the classification output head and the regression output head. The classification output head outputs the stage classification probability (early, middle, or late stage) of femoral head necrosis to the subject through a Softmax function, while the regression output head outputs a quantified risk score prediction value through a fully connected layer. These two results together constitute the validation prediction result, retaining the complete probability distribution and specific score values ​​during output.

[0155] S63. Calculate and validate the consistency index between the predicted results and the actual clinical outcome labels.

[0156] Optionally, consistency metrics are designed separately for the two types of outputs in the validation prediction results. For classification tasks, consistency metrics include accuracy, F1-score, and ROC AUC (Receiver Operating Characteristic Area Under Curve). For regression tasks, consistency metrics include mean squared error and Pearson correlation coefficient. Accuracy is calculated by dividing the number of correctly classified samples by the total number of samples in the validation dataset, directly reflecting the overall accuracy of staging classification. The F1-score is calculated using the harmonic mean of precision and recall, balancing the model's recognition performance across different staging categories and avoiding evaluation bias caused by data imbalance. ROC AUC is obtained by plotting an ROC curve and calculating the area under the curve, measuring the model's ability to distinguish between different staging categories. Mean squared error is obtained by calculating the mean square of the difference between the predicted and actual risk scores, reflecting the degree of deviation between the predicted and actual scores. The Pearson correlation coefficient is obtained by calculating the ratio of the product of their covariance and standard deviation, measuring the strength of the linear correlation between the predicted and actual scores. All metrics are calculated based on all samples in the validation dataset, comprehensively evaluating the model's classification accuracy and regression precision on the new clinical center data.

[0157] S64. When the consistency index is lower than the preset threshold, the verification dataset is added to the training data to incrementally train the risk prediction model and generate an optimized risk prediction model.

[0158] Optionally, the preset thresholds are determined jointly based on the validation set performance and actual clinical needs during the initial model training. Independent thresholds are assigned to classification and regression tasks, such as the F1-score threshold for classification and the mean squared error threshold for regression. The threshold settings must ensure sufficient reliability of the model in clinical applications. If the consistency index for any task is lower than the corresponding preset threshold, the model is deemed to have insufficient generalization performance on new clinical center data. During incremental training, the multimodal fusion feature set of the validation dataset is first merged with the feature set of the original training data, while retaining the corresponding real clinical outcome labels, forming a new training dataset. During the merging process, it is necessary to ensure a balanced data distribution to avoid model bias caused by an excessively high or low proportion of new data. Subsequently, the original multi-task loss function, i.e., the weighted sum of classification and regression loss values, is used. The backpropagation algorithm and the Adaptive Moment Estimation (Adam) optimizer are employed to update the model parameters. The learning rate during training is set to a small proportion of the initial training learning rate to avoid over-updating and causing the model to forget existing knowledge. During iterative training, the consistency index of the new validation set (proportionally divided from the new training dataset) is continuously monitored. Training is stopped when the index is stable and higher than a preset threshold, and an optimized risk prediction model is generated.

[0159] In the above embodiments, the generalization of the model is tested by constructing a validation set using data from new clinical centers, and the performance is quantitatively evaluated using multi-dimensional consistency indicators. When the data is below a preset threshold, incremental training is used to optimize the model. This effectively solves the problems of data heterogeneity and clinical translation barriers, improves the reliability and applicability of the risk prediction model in different clinical scenarios, and provides more stable technical support for the clinical assessment of avascular necrosis of the femoral head.

[0160] In one embodiment, the method further includes:

[0161] S7. Extract the attention weight data and neuron activation data generated by the risk prediction model during the prediction process.

[0162] Optionally, the attention weight data originates from the multi-head self-attention mechanism encoder layer in the electromyography network. During the calculation of attention-weighted features in this layer, the weight matrix of each attention head is simultaneously saved. These weight matrices intuitively reflect the correlation importance of electromyographic features between different time points and channels. Subsequently, the weight matrices of all attention heads are integrated according to the channel dimension to form global attention weight data. Neuron activation data is received from the neuron outputs of each core layer of the risk prediction model, specifically including all neurons in the kinematics subnetwork, dynamics subnetwork, convolutional layers, GRU layers, fully connected fusion layers, classification output head, and regression output head of the electromyography network. When the risk prediction model performs prediction and inference on the multimodal fusion feature set, the output value of each neuron is recorded in real time. This data can reflect the response intensity of each layer of the model to different features and fully capture the internal operation state of the model from feature input to result output.

[0163] S8. Calculate the Shapley value of each feature based on the multimodal fusion feature set and risk prediction results to obtain feature contribution data.

[0164] Optionally, the Shapley value is calculated based on the marginal contribution principle in game theory. Its core is to quantify the average contribution of each feature to the prediction result across all possible feature subset combinations. The SHAP (Shapley Additive Explanations) algorithm used is a model-agnostic method, requiring no dependence on the specific structure of the risk prediction model; feature contributions can be derived solely from input features and output results. Using a multimodal fusion feature set as the set of all features participating in the "game," and the risk prediction result (including stage classification probabilities and risk score predictions) as the target outcome of the "game," the SHAP algorithm traverses all possible feature subsets for each feature, calculating the change in the prediction result when the feature is removed from or added to a subset—that is, the marginal contribution. The average of the marginal contributions across all subsets is then used to obtain the Shapley value for that feature. Feature contribution data is centered on the Shapley value, where positive values ​​indicate a promoting effect on the prediction result, negative values ​​indicate an inhibiting effect, and the absolute value reflects the strength of the contribution.

[0165] S9. Identify the set of key features whose contribution exceeds a preset threshold based on the feature contribution data.

[0166] Optionally, the determination of the preset threshold needs to combine the statistical distribution of Shapley values ​​of all features in the multimodal fusion feature set with actual clinical needs. By analyzing the overall distribution pattern of feature contribution data, noisy features with weak contributions are eliminated, while ensuring that no important features with substantial impact on the prediction results are missed. The threshold setting needs to be jointly verified by clinical experts and technicians to ensure its rationality and practicality. First, the feature contribution data (Shapley values) of all features are sorted by absolute value. Then, the contribution data of each feature is traversed, and the Shapley value is compared with the preset threshold. If the absolute value of the Shapley value of a feature is higher than the preset threshold, the feature is determined to be a key feature affecting the prediction of femoral head necrosis risk. After filtering out all such features, they are arranged in descending order of absolute Shapley value to form a key feature set. This key feature set usually includes core elements from gait symmetry feature data, muscle and joint coupling feature data, and muscle and pressure synergy feature data, as well as key indicators from kinematic features, dynamic features, and electromyographic features.

[0167] S10. Match the key feature set with the clinical knowledge rule base to obtain the matching result; and generate a pathological explanation in natural language based on the matching result; the clinical knowledge rule base is used to store the mapping relationship between features and explanation statements.

[0168] Optionally, the clinical knowledge rule base is a structured database built upon clinical diagnostic guidelines for femoral head necrosis, pathobiomechanical research findings, and the clinical experience of senior orthopedic surgeons. Its core function is to store fixed mapping relationships between key features and clinically understandable pathological explanations. Each mapping relationship has undergone multiple rounds of clinical validation to ensure a precise correspondence between technical features and pathological mechanisms. For example, "decreased root mean square amplitude of the gluteus medius" corresponds to "weakness of the gluteus medius may lead to decreased hip joint stability and compensatory gait." During the matching process, each feature in the key feature set is sequentially and precisely compared with the feature name in the clinical knowledge rule base. If a completely matching feature is found, the corresponding pathological explanation is extracted. If a key feature has multiple associated explanations, they are integrated according to clinical logical priority. If some features have no direct matches, the most similar explanation is associated based on the feature's pathological attributes. Subsequently, all extracted or associated explanations are sorted according to the contribution of the key features, and natural language processing technology is used to refine the statements, forming logically coherent and standardized natural language pathological explanations.

[0169] S11. Integrate pathological interpretation with risk prediction results to generate an assessment report.

[0170] Optionally, the fusion process prioritizes clinical practicality and readability. First, it clarifies the structural framework of the assessment report, comprising two core modules: risk prediction results and pathological interpretation. The risk prediction results must fully present the probabilities of femoral head necrosis staging (specific probability values ​​for early, intermediate, and late stages) and the predicted risk score, clearly defining the subject's disease stage and risk level. The pathological interpretation section describes the contribution of key features from highest to lowest, with each key feature corresponding to its pathological mechanism description, explaining how the feature affects the prediction results. The fusion process adopts a "result + cause" logical order, first providing a clear risk prediction conclusion, and then breaking down the core pathological basis behind the prediction in conjunction with the pathological interpretation. The final assessment report is written in concise and standardized language, containing both objective, quantitative prediction data and clear, easy-to-understand explanations of pathological mechanisms, conforming to the reading habits and decision-making needs of clinicians, and can be directly used as a reference for diagnosis and treatment.

[0171] In the above embodiments, by extracting internal model data, quantifying feature contributions, screening key features, and matching clinical rules to generate pathological interpretations, and finally integrating the prediction results to form an evaluation report, the "black box" problem of the model is effectively solved, making the interpretation of the prediction results consistent with clinical cognition, improving doctors' trust in the AI ​​system, and providing comprehensive support for the clinical diagnosis and treatment of femoral head necrosis that combines accuracy and interpretability.

[0172] The aforementioned method and system for predicting the risk of femoral head necrosis based on multimodal gait analysis acquires multimodal data on joint kinematics, plantar pressure dynamics, and lower limb muscle electrophysiology, along with associated clinical diagnostic information. After gait cycle segmentation and time normalization preprocessing, a generative adversarial network is used to complete the missing modal data. Single-modal features are then extracted and fused with cross-modal correlation features such as gait symmetry, muscle-joint coupling, and muscle-pressure synergy, and standardized. A multi-branch deep learning model with a multi-head self-attention mechanism is used, combined with a multi-task loss function to train the risk prediction model. The model's generalization ability is improved through validation and incremental optimization using data from a new clinical center. Finally, by extracting internal model data, calculating Shapley values ​​to select key features, and matching them with a clinical knowledge rule base, an interpretable assessment report is generated. This technical solution provides comprehensive, accurate, and interpretable quantitative data for the early prevention, clinical diagnosis, and rehabilitation of femoral head necrosis, effectively improving treatment outcomes and reducing related medical burdens.

[0173] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0174] Based on the same inventive concept, this application also provides a system for implementing the above-mentioned method for predicting the risk of femoral head necrosis based on multimodal gait analysis. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the system for predicting the risk of femoral head necrosis based on multimodal gait analysis provided below can be found in the limitations of the method for predicting the risk of femoral head necrosis based on multimodal gait analysis described above, and will not be repeated here.

[0175] In one exemplary embodiment, such as Figure 5 As shown, a femoral head necrosis risk prediction system 10 based on multimodal gait analysis is provided, including:

[0176] The multimodal data acquisition module 11 is used to acquire the initial multimodal time series dataset of the subject; the initial multimodal time series dataset is associated with clinical diagnostic information and includes raw time series data of joint kinematics, raw time series data of plantar pressure dynamics and raw time series data of lower limb muscle electrophysiology;

[0177] The gait cycle processing module 12 is used to perform gait cycle segmentation on the initial multimodal time series dataset to obtain segmented time series data, and to perform time normalization on the segmented time series data to obtain standardized time series data.

[0178] The cross-modal data completion module 13 is used to complete cross-modal data for samples with missing modalities in standardized time series data, generating a complete multimodal time series dataset;

[0179] The multimodal feature extraction module 14 is used to extract kinematic features, dynamic features and electromyographic features from the complete multimodal time series dataset, and calculate cross-modal correlation features based on kinematic features, dynamic features and electromyographic features to generate a multimodal fusion feature set;

[0180] The risk prediction model training module 15 is used to train the model based on a multimodal fusion feature set, by minimizing the multi-task loss function through a multi-branch deep learning model to optimize the model parameters and generate a trained risk prediction model.

[0181] The risk prediction module 16 is used to input the multimodal time series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.

[0182] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above-described method for predicting the risk of femoral head necrosis based on multimodal gait analysis.

[0183] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of a method for predicting the risk of femoral head necrosis based on multimodal gait analysis as described above.

[0184] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0185] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for predicting the risk of femoral head necrosis based on multimodal gait analysis, characterized in that, The method includes: S1. Obtain the initial multimodal time series dataset of the subject; the initial multimodal time series dataset is associated with clinical diagnostic information and includes raw time series data of joint kinematics, raw time series data of plantar pressure dynamics and raw time series data of lower limb muscle electrophysiology; S2. Perform gait cycle segmentation on the initial multimodal time series dataset to obtain segmented time series data, and perform time normalization on the segmented time series data to obtain standardized time series data. S3. Perform cross-modal data completion on samples with missing modalities in the standardized time series data to generate a complete multimodal time series dataset; S4. Extract kinematic features, dynamic features, and electromyographic features from the complete multimodal time-series dataset, and calculate cross-modal correlation features based on the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set; S5. Based on the multimodal fusion feature set, optimize the model parameters by minimizing the multi-task loss function through a multi-branch deep learning model to train the model and generate a trained risk prediction model. S6. Input the multimodal time series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.

2. The method according to claim 1, characterized in that, S3 includes: S31. Identify data samples lacking at least one modality from the standardized time series data, and use the existing modal data in the data samples as conditional data; S32. Input the conditional data into a pre-trained generator network, and generate completed missing modality time series data through the generator network; S33. Combine the completed missing modal time series data and the conditional data to generate a sample; S34. Input the generated sample and the real complete data sample into the discriminator network to distinguish between real and fake samples, obtain the discrimination probability, and update the generator network parameters based on the discrimination probability through adversarial loss. S35. Repeat S32 to S34 until the convergence condition is met, and output the complete multimodal time series dataset consisting of all the completed samples.

3. The method according to claim 1, characterized in that, S4 includes: S41. Calculate the symmetry ratio of the bilateral gait parameters in the kinematic features and the bilateral pressure parameters in the dynamic features to generate gait symmetry feature data; S42. Perform time delay correlation calculation on the activation time sequence data of the target muscle in the electromyographic features and the angle time sequence data of the corresponding joint in the kinematic features to generate muscle and joint coupling feature data. S43. Perform instantaneous correlation calculation on the pressure center trajectory data in the dynamic features and the corresponding muscle activation intensity data in the electromyographic features to obtain muscle-pressure synergistic feature data; the muscle-pressure synergistic feature data consists of multiple instantaneous correlation coefficients; the instantaneous correlation coefficients are calculated using the following formula: in, The instantaneous correlation coefficient, This is the offset of the pressure center. For muscle activation intensity, The preset window half-width, and These are the mean values ​​of the data within the window; S44. The gait symmetry feature data, the muscle and joint coupling feature data, the muscle and pressure coordination feature data, the kinematic features, the dynamic features, and the electromyographic features are vector-concatenated to obtain concatenated feature data. S45. Standardize the spliced ​​feature data to generate the multimodal fusion feature set.

4. The method according to claim 1, characterized in that, S5 includes: S51. Based on the multi-branch deep learning model, the kinematic features, dynamic features and electromyographic features are fused to obtain the final feature representation; S52. Input the final feature representation into the classification output head to obtain the stage classification probability; S53. Input the final feature representation into the regression output head to obtain the risk score prediction value; S54. Calculate the classification loss value based on the stage classification probability and the actual stage label; S55. Calculate the regression loss value based on the predicted risk score and the actual risk score; S56. Construct a multi-task loss function based on the weighted sum of the classification loss value and the regression loss value; S57. Minimize the multi-task loss function using the backpropagation algorithm, update the parameters of the multi-branch deep learning model, and generate the risk prediction model.

5. The method according to claim 4, characterized in that, S51 includes: S511. Input the kinematic features into the kinematic sub-network to obtain the high-level kinematic feature vector; S512. Input the dynamic features into the dynamic sub-network to obtain the high-level dynamic feature vector; S513. Input the electromyographic features into the electromyographic network to obtain the high-level feature vector of electromyography. S514. Concatenate the kinematic high-level feature vector, the dynamic high-level feature vector, and the electromyographic high-level feature vector to obtain a fused feature vector; S515. Input the fused feature vector into the fully connected fusion layer to obtain the final feature representation.

6. The method according to claim 5, characterized in that, S513 includes: S5131. Perform dimensional transformation on the electromyographic features to obtain a two-dimensional matrix of time sequence and channel; S5132. Input the time-series and channel two-dimensional matrix into the encoder layer containing a multi-head self-attention mechanism to obtain attention-weighted features; S5133. Perform feedforward neural network calculation and layer normalization on the attention-weighted features to obtain enhanced features; S5134. Perform global average pooling on the enhanced features to obtain the electromyography high-level feature vector.

7. The method according to claim 1, characterized in that, Following S5, the following is also included: S61. Obtain new subject data and real clinical outcome labels from new clinical centers to form a validation dataset; S62. Input the verification dataset into the risk prediction model for prediction, and output the verification prediction result; S63. Calculate the consistency index between the validation prediction results and the actual clinical outcome labels; S64. When the consistency index is lower than a preset threshold, the verification dataset is added to the training data to perform incremental training on the risk prediction model and generate an optimized risk prediction model.

8. The method according to claim 1, characterized in that, The method further includes: S7. Extract the attention weight data and neuron activation data generated by the risk prediction model during the prediction process; S8. Calculate the Shapley value of each feature based on the multimodal fusion feature set and the risk prediction result to obtain feature contribution data; S9. Identify a set of key features whose contribution is higher than a preset threshold based on the feature contribution data; S10. Match the set of key features with the clinical knowledge rule base to obtain the matching result; and generate a pathological explanation in natural language based on the matching result; the clinical knowledge rule base is used to store the mapping relationship between features and explanation statements; S11. Integrate the pathological interpretation with the risk prediction results to generate the assessment report.

9. A risk prediction system for femoral head necrosis based on multimodal gait analysis, characterized in that, The system includes: The multimodal data acquisition module is used to acquire the initial multimodal time-series dataset of the subjects; the initial multimodal time-series dataset is associated with clinical diagnostic information and includes raw time-series data of joint kinematics, raw time-series data of plantar pressure dynamics, and raw time-series data of lower limb muscle electrophysiology; The gait cycle processing module is used to perform gait cycle segmentation on the initial multimodal time series dataset to obtain segmented time series data, and to perform time normalization on the segmented time series data to obtain standardized time series data. The cross-modal data completion module is used to complete the cross-modal data of samples with missing modalities in the standardized time series data, and generate a complete multimodal time series dataset; The multimodal feature extraction module is used to extract kinematic features, dynamic features, and electromyographic features from the complete multimodal time-series dataset, and calculate cross-modal correlation features based on the kinematic features, dynamic features, and electromyographic features to generate a multimodal fusion feature set; The risk prediction model training module is used to train the model based on the multimodal fusion feature set, by minimizing the multi-task loss function through a multi-branch deep learning model to optimize the model parameters and generate a trained risk prediction model. The risk prediction module is used to input the multimodal time series dataset of the subjects to be evaluated into the trained risk prediction model to obtain the risk prediction results.