Fine-grained gait sub-phase recognition method and device based on spatiotemporal feature fusion

By deeply extracting and fusing the spatiotemporal characteristics of multi-source gait data, combining heterogeneous parallel convolutional architecture and bidirectional long and short-term memory network models, the problem of insufficient gait recognition accuracy of lower limb rehabilitation exoskeleton robots is solved, and high-precision and fine-grained gait subphase recognition is achieved.

CN119723676BActive Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510208954.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-09
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art has problems of complex multi-dimensional feature capture and insufficient overall recognition accuracy in fine-grained recognition of gait cycle of exoskeleton robots with lower limb rehabilitation.

Method used

By deeply extracting and fusion of the spatiotemporal features of gait multi-source data, a spatiotemporal feature fusion model based on heterogeneous parallel convolution architecture and bidirectional long and short-term memory networks is adopted, and a high-precision gait sub-phase recognition system is constructed based on a weighted classification cross-entropy loss function and a hierarchical K-fold cross-validation strategy.

Benefits of technology

It significantly improves the accuracy and fine-grainedness of gait phase recognition, and also has good real-time performance, which can accurately identify 8 types of gait phases, improving the gait recognition ability of lower limb rehabilitation exoskeleton robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723676B_ABST
    Figure CN119723676B_ABST
Patent Text Reader

Abstract

The present invention discloses a fine-grained gait sub-phase recognition method and device based on spatiotemporal feature fusion, which aims to accurately identify 8 types of gait sub-phases when a lower limb rehabilitation exoskeleton robot walks. The method includes: acquiring gait multi-source data by integrating multi-source heterogeneous sensors, designing a gait sub-phase label assignment algorithm, and constructing a training data set in combination with a sliding overlapping window technology; designing a feature extraction module based on a heterogeneous parallel convolution architecture, and combining it with BiLSTM to construct a spatiotemporal feature fusion model; introducing a weighted cross entropy loss function to improve the model's recognition ability for gait sub-phase categories with fewer samples, and training and evaluation are performed through a hierarchical K-fold cross-validation strategy; finally, an end-to-end gait sub-phase online recognition system is constructed based on Simulink to achieve real-time recognition of gait sub-phases. The present invention is superior to the prior art in terms of recognition accuracy, fine-grained analysis, and real-time performance, and significantly improves the perception ability of the lower limb rehabilitation exoskeleton robot to the walking mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of lower limb rehabilitation exoskeleton robot perception technology, and specifically to a fine-grained gait sub-phase recognition method and device based on spatiotemporal feature fusion. The focus is on solving the problems of difficulty in capturing complex multi-dimensional features and reduced overall recognition accuracy caused by the refinement of gait sub-phase recognition granularity. It covers key technologies such as multi-source data fusion, gait sub-phase label allocation, spatiotemporal feature extraction and fusion. Background Art

[0002] With the rapid development of artificial intelligence technology, lower limb rehabilitation exoskeleton robots have become an important research direction in the field of modern medical assistance. Through precise motion assistance and personalized training programs, rehabilitation exoskeleton robots can help patients restore their walking ability and significantly improve their quality of life. In the process of rehabilitation training, accurate recognition of walking patterns is the core link for robots to achieve efficient human-machine collaboration, and walking patterns are described by gait cycles. The gait cycle is usually divided into 2 to 8 sub-phases of granularity. From two types of coarse-grained phases (stance phase and swing phase), it is further refined to 8 sub-phases. The support phase is subdivided into: Initial Contact (IC), Loading Response (LR), Mid Stance (MSt), Terminal Stance (TSt) and Pre-swing (PSw), and the swing phase is subdivided into: Initial Swing (ISw), Mid Swing (MSw) and Terminal Swing (TSw). This fine-grained division method can not only comprehensively characterize gait characteristics, but also meet the needs of rehabilitation training for accuracy and personalization. Through high-precision recognition of fine-grained gait sub-phases, the rehabilitation exoskeleton robot can adjust the motion trajectory and assistance force in real time, thereby improving the patient's rehabilitation effect and comfort.

[0003] Although existing machine learning and deep learning methods have made some progress in the field of gait sub-phase recognition, the research mainly focuses on the recognition of the stance phase and the swing phase, and a few works have been extended to the recognition of five to six gait sub-phases. However, in the recognition task of seven or eight gait sub-phases, the bottleneck of model performance is difficult to break through. The fundamental reason is that human gait contains multi-dimensional, multi-scale fine-grained features, and further refinement of the recognition granularity requires capturing more complex and comprehensive feature information. At present, some methods try to use simple deep learning models for feature extraction, but these methods can often only extract features for a single dimension, and cannot effectively capture the intrinsic connection and complexity of multi-dimensional and multi-scale features in gait, which ultimately leads to low overall recognition accuracy. Therefore, it is urgent to design an efficient feature extraction and fusion mechanism that can fully exploit the fine-grained multi-dimensional features in the gait cycle, improve the high-precision recognition capability of complex gait sub-phases, and take into account a certain degree of real-time performance. The breakthrough in this problem not only has important academic research value, but also can provide strong technical support for the practical application of lower limb rehabilitation exoskeleton robots, thereby significantly improving patients' rehabilitation training effects and experience. Summary of the invention

[0004] In order to solve the problems of difficulty in capturing complex multi-dimensional features and insufficient overall recognition accuracy in fine-grained recognition of gait cycles of lower limb rehabilitation exoskeleton robots in the prior art, the present invention proposes a fine-grained gait sub-phase recognition method and device based on spatiotemporal feature fusion.

[0005] The present invention significantly improves the accuracy and granularity of gait sub-phase recognition by deeply extracting and fusing the spatiotemporal features of multi-source gait data, while also achieving good real-time performance.

[0006] In order to achieve the above effects, the first aspect of the present invention relates to a fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion, comprising the following steps:

[0007] S1: Integrate multi-source heterogeneous sensors on the lower limb rehabilitation exoskeleton robot, collect multi-source gait data of the subjects when they walk wearing the lower limb rehabilitation exoskeleton robot, and record the experimental video;

[0008] S2: The experimental video is extracted into static images, manually annotated to generate gold standard labels, and the time proportion of each sub-phase is calculated; the collected gait multi-source data is filtered and denoised, and a four-dimensional pressure distribution vector is constructed based on the ground contact force (GCF) data. At the same time, the surface electromyography (sEMG) signal and the inertial measurement unit (IMU) signal are normalized;

[0009] S3: Design a gait sub-phase label assignment algorithm to assign corresponding gait sub-phase labels to gait multi-source data; assign labels to the five sub-phases in the support phase based on the four-dimensional pressure distribution vector, and assign the remaining labels based on the time proportion of the three sub-phases in the swing phase;

[0010] S4: Using the sliding overlapping window technique, the processed gait multi-source data is divided into time windows of fixed length, and a unique label is assigned to each time window using the statistical majority method to construct a training dataset;

[0011] S5: Verify the reliability of the gait sub-phase label assignment algorithm based on gold standard labels;

[0012] S6: Design a feature extraction module based on a heterogeneous parallel convolutional architecture, and combine it with a bidirectional long short-term memory network to implement time series modeling. Then, cascade the output feature vectors to complete the classification decision and build a spatiotemporal feature fusion model.

[0013] S7: Introduce the weighted classification cross entropy loss function and dynamically adjust the category weights to improve the model's recognition ability for sub-phase categories with fewer samples;

[0014] S8: Use stratified K-fold cross-validation strategy for model training and evaluation, set hyperparameters reasonably and enable model checkpoint mechanism;

[0015] S9: Use Simulink to deploy the trained spatiotemporal feature fusion model, build an end-to-end gait sub-phase online recognition system, and realize real-time recognition of gait sub-phases.

[0016] The multi-source heterogeneous sensor described in step S1 includes a ground contact force insole, an inertial measurement unit, and a surface electromyography sensor;

[0017] The specific process of extracting the GCF data into a four-dimensional pressure distribution vector in step S2 is as follows:

[0018] The filtered GCF data is divided into four groups according to "forefoot", "arch", "foreheel" and "heel", and the average value of each group is calculated to obtain the average value of GCF of four groups; the average value of GCF of each group is subjected to threshold binarization processing, and the formula is as follows:

[0019] (1)

[0020] in is the GCF average value, is the threshold value, which is set to 5% of the subject's body weight; G represents the ground contact state, and its value is "1" for the ground state and "0" for the aboveground state; the G values ​​obtained after the threshold binarization processing are concatenated to construct a four-dimensional pressure distribution vector, so as to facilitate the subsequent gait sub-phase label allocation.

[0021] The specific design process of the gait sub-phase label allocation algorithm described in step S3 is as follows:

[0022] S31. Based on the four-dimensional pressure distribution vector, preliminarily determine whether the gait corresponding to the current gait multi-source data is in the support phase or the swing phase, that is, when any element in the four-dimensional pressure distribution vector is non-zero, the gait is determined to be in the support phase; when all elements in the four-dimensional pressure distribution vector are zero, the gait is determined to be in the swing phase;

[0023] S32. During the label assignment process, the four-dimensional pressure distribution vector is The corresponding gait multi-source data is initially assigned the label "9", indicating that it is in the swing phase, and the remaining data is initially assigned the label "0", indicating that it is in the support phase. Subsequently, the label assignment of the five sub-phases in the support phase is further clarified based on the distribution of "1" in the four-dimensional pressure distribution vector. If the four-dimensional pressure distribution vector is , assign the label "1", indicating that it is in "initial contact"; if the four-dimensional pressure distribution vector is or , assign label "2", indicating that it is in "load response"; if the four-dimensional pressure distribution vector is , assign the label "3", indicating "standing in the middle"; if the four-dimensional pressure distribution vector is or , assign the label "4", indicating "terminal standing"; if the four-dimensional pressure distribution vector is , assigned the label “5”, indicating that it is in “pre-swing”;

[0024] S33. Assign labels to the remaining gait multi-source data according to the time proportion of the three sub-phases in the swing phase. According to the time sequence and time proportion, the gait multi-source data initially assigned with label "9" are replaced with labels "6", "7" and "8" in sequence. Among them, label "6" indicates that it is in the "initial swing", label "7" indicates that it is in the "middle swing", and label "8" indicates that it is in the "final swing".

[0025] The specific process of constructing a high-quality training data set using the sliding overlapping window technology in step S4 is as follows:

[0026] S41. Set the window size to 50 and the sliding step size to 10 to segment the time series data; start from the first data point of the time series and intercept a time window of length 50 as the first window; then, start from the 11th data point and intercept data of the same length as the second window; slide the window at a fixed step size until the entire sequence is covered;

[0027] S42. The window labels are assigned using the statistical majority method, that is, the frequency of the labels at each time point in the window is counted, and the label with the highest frequency of occurrence is used as the unique label of the window; through the above method, the training data set generated includes: the input feature is a data matrix of size (50 × N) (where N is the number of signal channels), and the training label is the unique label corresponding to each window.

[0028] The specific process of designing a feature extraction module based on a heterogeneous parallel convolutional architecture in step S6 is as follows:

[0029] The input data of the model is 14-channel biosensor data consisting of 4-channel surface electromyography signals of both lower limbs and 3-channel inertial measurement unit signals. The input tensor dimension is: batch size 128×time step 50×sensor channel 14×signal dimension 1. The feature extraction module is composed of the following three heterogeneous convolution branches:

[0030] Spatiotemporal joint convolution branch: A 6×14 2D convolution kernel is used to slide along the spatiotemporal joint dimension. The number of convolution kernels is set to 16, and the ReLU activation function is used. The output feature tensor is regularized by the batch normalization layer, downsampled by the 2×2 maximum pooling layer, and finally compressed into a feature vector by the global average pooling layer.

[0031] Time-dominant convolution branch: A 14×1 time-dominant convolution kernel is used to slide along the time dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used. The output feature tensor is regularized by the batch normalization layer, downsampled by the 2×1 maximum pooling layer, and finally compressed into a feature vector by the global average pooling layer;

[0032] Spatial dominant convolution branch: A 1×14 spatial convolution kernel is used to slide along the sensor channel dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used. The output feature tensor is regularized by the batch normalization layer, and then the dimension is permuted by the Permute layer, reconstructed into a time series data structure, and finally compressed into a feature vector by the global average pooling layer.

[0033] The specific process of combining the bidirectional long short-term memory network to realize time series modeling in step S6 is as follows:

[0034] The output feature vectors of the three heterogeneous convolution branches are expanded into a three-dimensional temporal structure through the RepeatVector layer and then input into independently configured bidirectional long short-term memory networks. The number of hidden units of the spatiotemporal joint convolution branch corresponding to the bidirectional long short-term memory network is set to 128, and the forward and reverse outputs are concatenated into 256-dimensional temporal features, which are output through a 64-unit fully connected layer; the number of hidden units of the temporal dominant convolution branch corresponding to the bidirectional long short-term memory network is set to 64, and the output is a 64-dimensional temporal feature, which is output through a 32-unit fully connected layer; the number of hidden units of the spatial dominant convolution branch corresponding to the bidirectional long short-term memory network is set to 64, and the output is a 64-dimensional temporal feature, which is output through a 32-unit fully connected layer.

[0035] Step S6 describes cascading the output features and completing the classification decision. The specific process is as follows:

[0036] The output feature vectors of the bidirectional long short-term memory network corresponding to each branch are concatenated along the sensor channel dimension through the Concatenate layer to form a fused feature vector. It is mapped to a 256-dimensional semantic space through a 32-unit fully connected layer, and the weight matrix is ​​initialized using the Xavier normal distribution. Finally, the output layer uses the Softmax function to generate an 8×1 gait sub-phase probability distribution vector.

[0037] The weighted classification cross entropy loss function described in step S7 is introduced to dynamically adjust the category weights. The specific process is as follows:

[0038] Introducing weighted categorical cross entropy As a loss function, different categories are assigned weights that are proportional to the inverse of their sample numbers, which is mathematically expressed as:

[0039] (2)

[0040] in, is the number of categories, Representation sample One-hot encoding value of the label. If the label is Class, then ,otherwise ; The model is The predicted probability of the class output is calculated by the Softmax function. It is Class weight; Class weight Calculated by the following formula:

[0041] (3)

[0042] in, Represents the training set The total number of samples in the class.

[0043] The stratified K-fold cross-validation strategy described in step S8 has the following specific process:

[0044] During formal model training, the dataset is divided into 10 non-overlapping folds, and each fold is divided into training set, validation set, and test set in a ratio of 8:1:1, ensuring that the sample category distribution in each fold is consistent with the original dataset.

[0045] The hyperparameters described in step S8 are set and the model checkpoint mechanism is started. The specific process is as follows: the model is optimized using the adaptive moment estimation Adam algorithm, the initial learning rate is set to 0.001, and it is dynamically adjusted according to the average loss; to prevent the model from overfitting, an early stopping mechanism is adopted, that is, when the loss of the validation set fails to improve within 10 consecutive cycles, the training is terminated in advance; the model checkpoint is enabled during the training process, and the model with the smallest validation set loss is automatically saved at the end of each round of training, so as to ensure that the best model parameters are captured and avoid performance degradation due to overfitting; all models are trained on the NVIDIA 4060ti GPU, and the batch size is set to 128; the above hyperparameter settings are based on the pre-training effect of the previous model.

[0046] The step S9 described in which the gait sub-phase online recognition system is constructed based on the trained spatiotemporal feature fusion model is specifically performed as follows:

[0047] S91. Real-time data stream processing: Surface electromyography sensors and inertial measurement units are integrated on the lower limb rehabilitation exoskeleton robot, and 4-channel surface electromyography signals and 3-channel inertial measurement unit signals of both lower limbs are collected synchronously in real time at a sampling rate of 1000Hz to form 14 channels of raw gait data. A circular buffer management strategy is adopted, and the data cache queue length is set to 100 milliseconds. The raw gait data is divided into time windows by sliding overlapping window technology, with the window size set to 50 and the sliding step set to 10. The same filtering and denoising processing as step S2 is performed in units of windows. Subsequently, the dynamic range of the filtered and denoised gait data is normalized, that is, the normalization coefficient is updated in real time based on the maximum and minimum values ​​of the data in the last 10 seconds. Finally, the input tensor dimensions are configured as: batch size 128×time step 50×sensor channel 14×signal dimension 1, and all labels are subtracted by one to match the one-hot vector format.

[0048] S92. Online model reasoning: Convert the trained spatiotemporal feature fusion model into ONNX format, and use MATLAB's importONNXNetwork function to verify the compatibility of the model's input and output nodes, tensor dimensions, and data types. Build an end-to-end gait sub-phase real-time recognition system in Simulink, integrating the signal acquisition module, preprocessing module, data buffer module, model reasoning module, and result display module. Load the ONNX format file to the model reasoning module through Deep Learning Toolbox, perform online model reasoning, and output an 8×1 gait sub-phase probability distribution vector. At the same time, enable CUDA acceleration and single-sample batch processing mode. Finally, perform the maximum value decision on the gait sub-phase probability distribution vector output by the model, and its mathematical expression is:

[0049] (4)

[0050] in, represents the output label after executing the maximum value decision, Indicates time Input data Belongs to category The posterior probability of the maximum decision is obtained; the output label of the maximum decision is mapped from the numerical value to the gait sub-phase, and finally the real-time recognition result of the current gait sub-phase is obtained.

[0051] The second aspect of the present invention relates to a fine-grained gait sub-phase recognition device based on spatiotemporal feature fusion, which includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion of the present invention.

[0052] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion of the present invention.

[0053] The innovation of the present invention is as follows: by collecting and preprocessing multi-source gait data, combining the gait sub-phase label assignment algorithm and the sliding overlapping window technology, a high-quality training data set is effectively constructed; a spatiotemporal feature fusion model is constructed based on a heterogeneous parallel convolutional architecture and a bidirectional long short-term memory network to improve the model's ability to capture complex multidimensional features of gait; a weighted classification cross entropy loss function is introduced to solve the problem of gait sub-phase imbalance, thereby improving the model's ability to recognize sub-phase categories with fewer samples; the model is optimized, trained and evaluated through a stratified K-fold cross-validation strategy to further improve the robustness and generalization ability of the model; the trained spatiotemporal feature fusion model is deployed on a Simulink platform, an end-to-end gait sub-phase online recognition system is constructed, and real-time recognition of gait sub-phases is realized.

[0054] The working principle of the present invention is as follows: multi-source gait data is collected synchronously through multimodal sensors, and pre-processing techniques such as filtering, denoising, and normalization are used to significantly improve data quality; a gait sub-phase label allocation algorithm is designed to solve the data labeling problem, and the gait sub-phase recognition task is converted into a supervised learning problem, thereby constructing a high-quality training data set; a feature extraction module is designed based on a heterogeneous parallel convolution structure, and a bidirectional long short-term memory network (BidirectionalLong Short-Term Memory Network, BiLSTM) is combined to effectively extract local spatial features, capture motion patterns at different time points, and capture time series information through bidirectional modeling, explore long-term dependencies, and overcome the nonlinear and high-dimensional challenges of gait features; a weighted classification cross entropy loss function is introduced to dynamically adjust category weights, solve the category imbalance problem, and improve the model's recognition ability for gait sub-phase categories with fewer samples; a stratified K-fold cross-validation strategy is used to optimize model training and evaluation, further enhancing the stability and accuracy of the model; an end-to-end real-time gait sub-phase recognition system is constructed by deploying the model in Simulink to achieve real-time recognition of gait sub-phases.

[0055] Compared with the existing methods, the beneficial effects of the present invention are mainly reflected in: through the spatiotemporal feature fusion model, the present invention can efficiently extract the complex multidimensional features of gait data and realize high-precision recognition of eight types of gait sub-phases; by designing a gait sub-phase label assignment algorithm, the gait sub-phase recognition task is converted into a supervised learning problem, which effectively solves the data labeling problem; the introduction of the weighted classification cross entropy loss function effectively solves the category imbalance problem and significantly improves the model's recognition ability for gait sub-phase categories with fewer samples; based on the heterogeneous parallel convolutional architecture, the feature extraction module is designed, and combined with the bidirectional long short-term memory network, the spatiotemporal feature fusion model is constructed to realize high-precision recognition of eight types of gait sub-phases; through the hierarchical K-fold cross-validation strategy, the training and evaluation of the model are optimized, and the robustness and generalization ability of the model are significantly enhanced; after the model is deployed on the Simulink platform, the present invention can realize real-time and accurate recognition of gait sub-phases.

[0056] The present invention can accurately identify eight types of gait sub-phases when a lower limb rehabilitation exoskeleton robot is walking. It is superior to the existing technology in terms of recognition accuracy, fine-grained analysis and real-time performance. It significantly improves the lower limb rehabilitation exoskeleton robot's perception of walking patterns and has strong practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the overall process of a specific embodiment of the present invention;

[0058] Figure 2 A schematic diagram of an experimental hardware platform in a specific implementation mode of the present invention;

[0059] Figure 3 It is a schematic diagram of the surface electromyography collection position of the lower limb muscles in a specific embodiment of the present invention;

[0060] Figure 4 It is a schematic diagram of the division of four groups of areas of the ground contact force insole in a specific embodiment of the present invention;

[0061] Figure 5 A box plot for evaluating the reliability of a gait sub-phase label assignment algorithm in a specific embodiment of the present invention;

[0062] Figure 6 A detailed network architecture of a spatiotemporal feature fusion model in a specific embodiment of the present invention;

[0063] Figure 7 This is a step-by-step comparison diagram of the recognition results of the model in real-time gait sub-phase recognition in a specific embodiment of the present invention and the true label. DETAILED DESCRIPTION

[0064] The following embodiments of the invention are described in detail in terms of experimental scheme, data processing, model construction and training, and performance verification in conjunction with the accompanying drawings to further clarify the application field, design ideas, and technical solutions of the invention.

[0065] Example 1

[0066] This embodiment relates to a fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion, such as Figure 1 As shown, the following steps are included:

[0067] Step S1: Integrate multi-source heterogeneous sensors on the lower limb rehabilitation exoskeleton robot, collect multi-source gait data of the subject when walking wearing the lower limb rehabilitation exoskeleton robot, and record the experimental video.

[0068] like Figure 2 As shown in the figure, by integrating multi-source heterogeneous sensors on the robot, including ground contact force insoles, inertial measurement units, and surface electromyography sensors, the synchronous acquisition of multi-source gait data is achieved, and a camera is configured next to the treadmill to record the walking process. Real-time data acquisition and monitoring are completed through Simulink software to ensure efficient management and analysis of experimental data. The subjects were required to walk on the treadmill at a fixed speed of 1.5 km / h for 60 seconds to complete an experimental cycle. Each subject was required to complete 5 experimental cycles, and they could get enough rest after each cycle to relieve muscle fatigue.

[0069] A six-lead electromyography sensor (SICHIRAY Rev2.0, Sizhirui Technology Co., Ltd., Wuxi) was used to record the surface electromyography (sEMG) signals of four muscles (8 channels in total) of both legs. The electrodes used were bipolar active Ag / AgCl electrodes and were precisely placed according to the recommendations of the Surface Electromyography for the Non-Invasive Assessment of Muscles (SENIAM). Figure 3As shown in the figure, the muscles for collecting electromyographic signals include the rectus femoris (RF), vastus medialis (VM), vastus lateralis (VL) and medial gastrocnemius (MG), and the sampling rate is set to 1000Hz; the inertial measurement unit (IMU) signal is collected using the DETA10 series micro inertial navigation system of FDISYSTEMS, and the IMU is fixed at the knee joint of the robot to monitor the dynamic changes of posture during walking. The V series equipment of this system is selected to record the yaw angle (Yaw), roll angle (Roll) and pitch angle (Pitch), and the sampling rate is also 1000Hz; in addition, the commercial ground contact force insole RX-ES42-18 (GT International Electronic Technology Co., Ltd., Hong Kong) is installed on the sole of the lower limb rehabilitation exoskeleton robot to record the ground contact force (GCF) data during walking at a sampling rate of 100Hz.

[0070] Step S2: The experimental video is extracted as a static image, manually annotated to generate a gold standard label, and the time proportion of each sub-phase is calculated; the collected gait multi-source data is filtered and denoised, and a four-dimensional pressure distribution vector is constructed based on the GCF data. At the same time, the surface electromyography signal and the inertial measurement unit signal are normalized.

[0071] 10% of the experimental videos were extracted as static images at a frequency of 20 frames per second. Based on the definition of 8 types of gait sub-phases, the static images were labeled with their respective gait sub-phases, and the time proportion of each sub-phase was calculated. The collected sEMG signals were preprocessed as follows: a 50Hz notch filter was applied to remove power line interference, and then a 5th-order Butterworth filter was used for 25Hz-400Hz bandpass filtering, and finally a maximum-minimum normalization process was performed; the IMU signal was low-pass filtered at 5Hz through a 4th-order Butterworth filter, and then baseline drift was removed and maximum-minimum normalization was performed; the GCF data was low-pass filtered at 40Hz and baseline drift was removed, and upsampled to 1000Hz by cubic spline interpolation to achieve alignment with the sampling rate of sEMG and IMU signals.

[0072] The filtered and denoised GCF data is Figure 4 The specified "Forefoot (FF)", "Arch of the Foot (AF)", "Rearfoot (RF)" and "Heel (HL)" are divided into four groups, and the average value of each group is calculated to obtain the average value of GCF of the four groups; the average value of GCF of each group is subjected to threshold binarization processing, and the formula is as follows:

[0073] (5)

[0074] in is the GCF average value, is the threshold value, which is set to 5% of the subject's body weight; G represents the ground contact state, and its value is "1" for the ground state and "0" for the aboveground state; the G values ​​obtained after the threshold binarization processing are concatenated to construct a four-dimensional pressure distribution vector, so as to facilitate the subsequent gait sub-phase label allocation.

[0075] Step S3: Design a gait sub-phase label assignment algorithm to assign corresponding gait sub-phase labels to gait multi-source data; assign labels to the five sub-phases in the support phase based on the four-dimensional pressure distribution vector, and assign the remaining labels based on the time proportion of the three sub-phases in the swing phase.

[0076] First, based on the four-dimensional pressure distribution vector, it is preliminarily determined whether the gait corresponding to the current gait multi-source data is in the support phase or the swing phase. That is, when any element in the four-dimensional pressure distribution vector is non-zero, the gait is determined to be in the support phase; when all elements in the four-dimensional pressure distribution vector are zero, the gait is determined to be in the swing phase.

[0077] Secondly, during the label assignment process, the four-dimensional pressure distribution vector is The corresponding gait multi-source data is initially assigned the label "9", indicating that it is in the swing phase, and the remaining data is initially assigned the label "0", indicating that it is in the support phase. Subsequently, according to the distribution of "1" in the four-dimensional pressure distribution vector, the label assignment of the five sub-phases in the support phase is further clarified. The mapping relationship between the four-dimensional pressure distribution vector and the assigned label is shown in Table I.

[0078] Finally, labels are assigned to the remaining gait multi-source data according to the time proportion of the three sub-phases in the swing phase. According to the time sequence and time proportion, the gait multi-source data initially assigned with label "9" are replaced with labels "6", "7" and "8" in sequence. Among them, label "6" indicates that it is in the "initial swing", label "7" indicates that it is in the "middle swing", and label "8" indicates that it is in the "final swing".

[0079] Table I Mapping relationship between four-dimensional pressure distribution vector and assigned label

[0080]

[0081] Step S4: Using the sliding overlapping window technique, the processed gait multi-source data is divided into time windows of fixed length, and a unique label is assigned to each time window using the statistical majority method to construct a training data set.

[0082] The window size is set to 50 and the sliding step is set to 10 to retain some overlapping information between adjacent windows, enhance continuity and context association, and ensure a certain degree of independence, which helps the model learn the dynamic characteristics of the time series. Specifically, starting from the first data point of the time series, a time window of length 50 is intercepted as the first window; then, starting from the 11th data point, data of the same length is intercepted as the second window. The window is slid at a fixed step size until the entire sequence is covered.

[0083] To ensure the consistency of supervised training, a unique label needs to be assigned to each window. The statistical mode method is used to assign window labels, that is, the frequency of the corresponding labels at each time point in the statistical window is counted, and the label with the highest frequency (mode) is used as the unique label of the window to ensure its representativeness. Through the above method, the generated training data set includes: the input feature is a data matrix of size (50 × N) (where N is the number of signal channels), and the training label is the unique label corresponding to each window.

[0084] Step S5: Verify the reliability of the proposed gait sub-phase label assignment algorithm based on the gold standard label.

[0085] 500 complete gait cycles are randomly selected from the dataset containing gold standard labels for label assignment and divided into 100 groups, each containing 5 gait cycles. The reliability of the proposed gait sub-phase label assignment algorithm is verified by comparing the consistency of each group of assigned labels with the gold standard labels.

[0086] The precision, recall and F1 score of the gait sub-phase label assignment algorithm on each gait sub-phase are statistically analyzed. As shown in Table II, the assignment algorithm performs well in all indicators, with a macro-average precision of 96.96±2.53%, a macro-average recall of 91.08±4.09%, and a macro-average F1 score of 93.89±2.89%. In addition, Figure 5 The stability of the gait sub-phase label assignment algorithm in the label assignment process is intuitively demonstrated. As can be seen from the figure, the F1 score distribution of the gait sub-phase label assignment algorithm on most gait sub-phases is compact and consistent, with few outliers, and it performs well in multiple performance indicators, fully demonstrating its reliability in the gait sub-phase label assignment task.

[0087] Table II Reliability verification of gait sub-phase label assignment algorithm

[0088]

[0089] Step S6: Design a feature extraction module based on a heterogeneous parallel convolutional architecture, and combine it with a bidirectional long short-term memory network to implement time series modeling. Then, cascade the output feature vectors to complete the classification decision and build a spatiotemporal feature fusion model.

[0090] The input data of the model is 14-channel biosensor data consisting of 4-channel surface electromyographic signals and 3-channel inertial measurement unit signals of both lower limbs, and the input tensor dimension is 128×50×14×1 (batch size×time step×sensor channel×signal dimension). The feature extraction module consists of three heterogeneous convolution branches: the spatiotemporal joint convolution branch uses 6×14 two-dimensional convolution kernels to slide along the spatiotemporal joint dimension, the number of convolution kernels is set to 16, and the ReLU activation function is used. The output feature tensor is regularized by the batch normalization layer, downsampled by the 2×2 maximum pooling layer, and finally compressed into a feature vector by the global average pooling layer; the time-dominant convolution branch uses a 14×1 time-dominant convolution kernel to slide along the time dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used. After the output feature tensor is regularized by the batch normalization layer, it is downsampled by the maximum pooling layer of 2×1, and finally compressed into a feature vector by the global average pooling layer; the spatial dominant convolution branch uses a 1×14 spatial convolution kernel to slide along the sensor channel dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used. After the output feature tensor is regularized by the batch normalization layer, it is dimensionally permuted by the Permute layer, reconstructed into a time series data structure, and finally compressed into a feature vector by the global average pooling layer.

[0091] Subsequently, the output feature vectors of the three heterogeneous convolution branches are expanded into a three-dimensional temporal structure through the RepeatVector layer and then input into independently configured bidirectional long short-term memory networks. Among them, the number of hidden units of the bidirectional long short-term memory network corresponding to the spatiotemporal joint convolution branch is set to 128, and the forward and reverse outputs are spliced ​​into 256-dimensional temporal features, which are output through a 64-unit fully connected layer; the number of hidden units of the bidirectional long short-term memory network corresponding to the time-dominant convolution branch is set to 64, and the 64-dimensional temporal features are output through a 32-unit fully connected layer; the number of hidden units of the bidirectional long short-term memory network corresponding to the space-dominant convolution branch is set to 64, and the 64-dimensional temporal features are output through a 32-unit fully connected layer.

[0092] Finally, the output feature vectors of the bidirectional long short-term memory network corresponding to each branch are concatenated along the sensor channel dimension through the Concatenate layer to form a fused feature vector. The 32-unit fully connected layer is used to map it to a 256-dimensional semantic space. The weight matrix is ​​initialized using the Xavier normal distribution, and the model output layer uses the Softmax function to generate an 8×1 gait sub-phase probability distribution vector. At this point, the spatiotemporal feature fusion model is built, and its detailed network architecture is as follows: Figure 6shown.

[0093] Step S7: Introduce the weighted classification cross entropy loss function and dynamically adjust the category weights to improve the model's recognition ability for sub-phase categories with fewer samples.

[0094] In order to alleviate the problem of category imbalance, weighted classification cross entropy is introduced As a loss function, by assigning weights proportional to the inverse of the number of samples to different categories, the dominant effect of the majority class on the loss is reduced and the model's sensitivity to the minority class is enhanced. Its mathematical expression is:

[0095] (6)

[0096] in, is the number of categories, Representation sample One-hot encoding value of the label (if the true label is Class, then ,otherwise ), The model is The predicted probability of the class output (calculated by the Softmax function), It is The weight of the class. It is calculated by the following formula:

[0097] (7)

[0098] in, Represents the training set Through this weighting mechanism, less common categories will receive greater weights, thereby improving the recognition ability of categories with fewer samples.

[0099] Table III shows the classification cross entropy of the model in pre-training without introducing the weighting mechanism. And the weighted classification cross entropy The recognition results of each sub-phase when using the two loss functions. It can be seen that after using the weighted mechanism, the recognition accuracy of the model is improved in almost all sub-phases, especially in the recognition of the IC sub-phase, where the accuracy improvement is particularly significant.

[0100] Table III The recognition accuracy of each sub-phase when the model uses two loss functions in pre-training (%)

[0101]

[0102] Step S8: Use the stratified K-fold cross-validation strategy to train and evaluate the model, set hyperparameters reasonably, and enable the model checkpoint mechanism.

[0103] In the formal model training process, a stratified K-fold cross-validation strategy is used for model training and evaluation. Specifically, the dataset is divided into 10 non-overlapping folds, and each fold is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. At the same time, the sample distribution of each fold subset is ensured to be consistent with the original dataset, effectively reducing the deviation caused by class imbalance and improving the generalization performance of the model.

[0104] The adaptive moment estimation (Adam) algorithm was used to optimize the model. The initial learning rate was set to 0.001 and dynamically adjusted according to the average loss. To prevent overfitting of the model, an early stopping mechanism was adopted, that is, when the loss of the validation set failed to improve within 10 consecutive cycles, the training was terminated early. Model checkpoints were enabled during the training process, and the model with the smallest validation set loss was automatically saved at the end of each round of training to ensure that the optimal model parameters were captured and performance degradation due to overfitting was avoided. All models were trained on an NVIDIA 4060ti GPU with a batch size of 128. The above hyperparameter settings were based on the pre-training results of the previous models.

[0105] Step S9: Use Simulink to deploy the trained spatiotemporal feature fusion model, build an end-to-end gait sub-phase online recognition system, and realize real-time recognition of gait sub-phases.

[0106] The surface electromyography sensor and inertial measurement unit were integrated on the lower limb rehabilitation exoskeleton robot. The 4-channel surface electromyography signals and 3-channel inertial measurement unit signals of the two lower limbs were collected synchronously in real time at a sampling rate of 1000Hz to form 14 channels of raw gait data. The circular buffer management strategy was adopted, and the data cache queue length was set to 100 milliseconds. The raw gait data was divided into time windows by sliding overlapping window technology. The window size was set to 50 and the sliding step was set to 10. The same filtering and denoising process as step S2 was performed in units of windows. Subsequently, the dynamic range of the filtered and denoised gait data was normalized, that is, the normalization coefficient was updated in real time based on the maximum and minimum values ​​of the data in the last 10 seconds. Finally, the input tensor dimension was configured as: 128×50×14×1 (batch size×time step×sensor channel×signal dimension), and all labels were subtracted by one to match the one-hot vector format.

[0107] Subsequently, the trained spatiotemporal feature fusion model was converted into ONNX format, and the compatibility of the model's input and output nodes, tensor dimensions, and data types was verified using MATLAB's importONNXNetwork function. An end-to-end gait sub-phase real-time recognition system was built in Simulink, integrating the signal acquisition module, preprocessing module, data buffer module, model inference module, and result display module. The ONNX format file was loaded into the model inference module through Deep Learning Toolbox, and online model inference was performed to output an 8×1 gait sub-phase probability distribution vector. At the same time, CUDA acceleration and single-sample batch processing mode were enabled. Finally, the maximum value decision was performed on the gait sub-phase probability distribution vector output by the model, and its mathematical expression is:

[0108] (8)

[0109] in, represents the output label after executing the maximum value decision, Represents the input data at time t Belongs to category The posterior probability of the maximum decision is mapped from the numerical value to the gait sub-phase, and finally the real-time recognition result of the current gait sub-phase is obtained and saved.

[0110] In addition, the gait multi-source data in the gait sub-phase real-time recognition task was exported, and the gait sub-phase label assignment algorithm proposed in the present invention was applied to obtain the real label. By comparing with the recognition result, the accuracy of gait sub-phase real-time recognition was calculated to be 93.78%±2.06%. At the same time, a step comparison diagram was drawn (such as Figure 7 As shown in the figure, it can be seen that the method proposed in the present invention can improve the recognition granularity to 8 categories in the real-time gait sub-phase recognition task, especially in the IC sub-phase where samples are relatively scarce, the present invention can still achieve high-precision recognition and improve the integrity of gait sub-phase recognition.

[0111] Example 2

[0112] Reference Figure 1 , 2 , 3, 6, this embodiment relates to a fine-grained gait sub-phase recognition device based on spatiotemporal feature fusion. The device includes a display interface, a memory, and one or more processors. The memory stores executable code, and when the processor executes the code, it can realize real-time recognition of gait sub-phases. The device transmits gait data to the processor in real time through Bluetooth technology for processing, and displays the recognition results in real time through the display interface. The device combines the display interface, memory, processor and executed executable code to work together to realize a fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion.

[0113] Example 3

[0114] Reference Figure 1-6 This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion of the present invention is implemented.

[0115] In summary, the present invention aims at the problem of difficulty in capturing complex multi-dimensional features and decreased overall accuracy caused by the granularity refinement of gait sub-phase recognition of lower limb rehabilitation exoskeleton robots, and proposes a fine-grained gait sub-phase recognition method and device based on spatiotemporal feature fusion. The present invention realizes the synchronous acquisition of multi-source gait data by integrating multi-source heterogeneous sensors (including ground contact force insoles, inertial measurement units, and surface electromyography sensors) on lower limb rehabilitation exoskeleton robots, and designs a gait sub-phase label allocation algorithm to transform the gait sub-phase recognition task into a supervised learning problem. In order to accurately capture the spatiotemporal dynamic pattern and local details of gait, the present invention innovatively designs a feature extraction module based on a heterogeneous parallel convolutional architecture, and constructs a spatiotemporal feature fusion model in combination with a bidirectional long short-term memory network. By introducing a weighted classification cross entropy loss function, the model's recognition ability for gait sub-phase categories with fewer samples is effectively improved. The spatiotemporal feature fusion model that has been trained is deployed using Simulink, and an end-to-end gait sub-phase online recognition system is constructed to realize real-time recognition of gait sub-phases. The present invention has achieved remarkable breakthroughs in overall accuracy, recognition granularity and real-time performance, providing key technical support for gait recognition of lower limb rehabilitation exoskeleton robots, and has broad application prospects and potential development value.

[0116] The above embodiment describes an implementation of the present invention. For those skilled in the art, various adjustments, changes and improvements may be made to the present invention. All adjustments, changes and improvements made within the scope of the core principle of the present invention are considered to be within the scope of the present invention.

Claims

1. A fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion, characterized in that: The following steps are involved: S1: Integrate multi-source heterogeneous sensors on the lower limb rehabilitation exoskeleton robot to collect multi-source gait data of the subjects when they walk wearing the lower limb rehabilitation exoskeleton robot; S2: Extract the experimental video into static images, manually annotate them to generate gold standard labels, and calculate the time proportion of each sub-phase; filter and denoise the collected gait multi-source data, construct a four-dimensional pressure distribution vector based on the ground contact force GCF data, and normalize the surface electromyography signal and inertial measurement unit signal; S3: Design a gait sub-phase label assignment algorithm to assign corresponding gait sub-phase labels to gait multi-source data; assign labels to the five sub-phases in the support phase based on the four-dimensional pressure distribution vector, and assign the remaining labels based on the time proportion of the three sub-phases in the swing phase; S4: Using the sliding overlapping window technique, the processed gait multi-source data is divided into time windows of fixed length, and a unique label is assigned to each time window using the statistical majority method to construct a training dataset; S5: Verify the reliability of the gait sub-phase label assignment algorithm based on gold standard labels; S6: Design a feature extraction module based on a heterogeneous parallel convolutional architecture, and combine it with a bidirectional long short-term memory network to implement time series modeling. Then, cascade the output feature vectors to complete the classification decision and build a spatiotemporal feature fusion model. S7: Introduce the weighted classification cross entropy loss function and dynamically adjust the category weights to improve the model's recognition ability for sub-phase categories with fewer samples; S8: Use stratified K-fold cross-validation strategy for model training and evaluation, set hyperparameters reasonably and enable model checkpoint mechanism; S9: Use Simulink to deploy the trained spatiotemporal feature fusion model, build an end-to-end gait sub-phase online recognition system, and realize real-time recognition of gait sub-phases.

2. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The multi-source heterogeneous sensor described in step S1 includes a ground contact force insole, an inertial measurement unit, and a surface electromyography sensor.

3. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The specific process of constructing the four-dimensional pressure distribution vector based on the ground contact force GCF data in step S2 is as follows: The ground contact force GCF data after filtering and denoising are divided into four groups according to "forefoot", "arch", "front heel" and "heel", and the average value of each group is calculated to obtain the average value of the four groups of GCF; the threshold binarization process is performed on the average value of each group of GCF, and the formula is as follows: (1) in is the GCF average value, is the threshold value, which is set to 5% of the subject's body weight; G represents the ground contact state, and its value is "1" for the ground state and "0" for the aboveground state; the G values ​​obtained after the threshold binarization processing are concatenated to construct a four-dimensional pressure distribution vector, so as to facilitate the subsequent gait sub-phase label allocation.

4. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The specific design process of the gait sub-phase label assignment algorithm described in step S3 is as follows: S31. Based on the four-dimensional pressure distribution vector, preliminarily determine whether the gait corresponding to the current gait multi-source data is in the support phase or the swing phase, that is, when any element in the four-dimensional pressure distribution vector is non-zero, the gait is determined to be in the support phase; When all elements of the four-dimensional pressure distribution vector are zero, the gait is judged to be in the swing phase; S32. During the label assignment process, the four-dimensional pressure distribution vector is The corresponding gait multi-source data were initially assigned the label "9", indicating that they were in the swing phase, and the remaining data were initially assigned the label "0", indicating that they were in the stance phase; Then, according to the distribution of "1" in the four-dimensional pressure distribution vector, the label assignment of the five sub-phases in the support phase is further clarified; if the four-dimensional pressure distribution vector is , assign label "1", indicating "initial contact"; if the four-dimensional pressure distribution vector is or , assign label "2", indicating that it is in "load response"; if the four-dimensional pressure distribution vector is , assign label "3", indicating "standing in the middle"; if the four-dimensional pressure distribution vector is or , assign label "4", indicating "terminal standing"; if the four-dimensional pressure distribution vector is , assign the label "5", indicating that it is in "pre-swing"; S33. Assign labels to the remaining gait multi-source data according to the time proportion of the three sub-phases within the swing phase; replace the gait multi-source data initially assigned with label "9" with labels "6", "7" and "8" in sequence according to the time sequence and time proportion; wherein label "6" indicates that it is in the "initial swing", label "7" indicates that it is in the "middle swing", and label "8" indicates that it is in the "final swing".

5. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The specific process of constructing the training data set using the sliding overlapping window technology in step S4 is as follows: S41. Set the window size to 50 and the sliding step size to 10 to segment the time series data; start from the first data point of the time series and intercept a time window of length 50 as the first window; then, start from the 11th data point and intercept data of the same length as the second window; slide the window at a fixed step size until the entire sequence is covered; S42. The window label is assigned using the statistical majority method, that is, the frequency of the label at each time point in the window is counted, and the label with the highest frequency of occurrence is used as the only label of the window.

6. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The detailed architecture and hierarchical configuration of the spatiotemporal feature fusion model in step S6 specifically include: The feature extraction module designed based on the heterogeneous parallel convolutional architecture described in step S6 specifically includes: The input data of the model is 14-channel biosensor data consisting of 4-channel surface electromyographic signals of both lower limbs and 3-channel inertial measurement unit signals. The input tensor dimension is: batch size 128 × time step 50 × sensor channel 14 × signal dimension 1; the feature extraction module includes the following three heterogeneous convolution branches: Spatiotemporal joint convolution branch: A 6×14 two-dimensional convolution kernel is used to slide along the joint time-space dimension. The number of convolution kernels is set to 16, and the ReLU activation function is used. The output feature tensor is regularized by the batch normalization layer, downsampled by the 2×2 maximum pooling layer, and finally compressed into a feature vector by the global average pooling layer. Time-dominant convolution branch: A 14×1 time-dominant convolution kernel is used to slide along the time dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used; the output feature tensor is regularized by the batch normalization layer, downsampled by a 2×1 maximum pooling layer, and finally compressed into a feature vector by a global average pooling layer; Spatial dominant convolution branch: A 1×14 spatial convolution kernel is used to slide along the sensor channel dimension, the number of convolution kernels is set to 8, and the ReLU activation function is used; the output feature tensor is regularized by the batch normalization layer, and then the dimension is permuted by the Permute layer to reconstruct it into a time series data structure, and finally compressed into a feature vector by the global average pooling layer; The step S6 of combining a bidirectional long short-term memory network to realize time series modeling specifically includes: The output feature vectors of the three heterogeneous convolution branches are expanded into a three-dimensional time series structure through the RepeatVector layer and input into independently configured bidirectional long short-term memory networks respectively; among them, the number of hidden units of the bidirectional long short-term memory network corresponding to the spatiotemporal joint convolution branch is set to 128, and the forward and reverse outputs are spliced ​​into 256-dimensional time series features, which are output through a 64-unit fully connected layer; the number of hidden units of the bidirectional long short-term memory network corresponding to the time series dominant convolution branch is set to 64, and the 64-dimensional time series features are output through a 32-unit fully connected layer; the number of hidden units of the bidirectional long short-term memory network corresponding to the space dominant convolution branch is set to 64, and the 64-dimensional time series features are output through a 32-unit fully connected layer; Step S6 of cascading the output features and completing the classification decision specifically includes: The output feature vectors of the bidirectional long short-term memory network corresponding to each branch are concatenated along the sensor channel dimension through the Concatenate layer to form a fused feature vector; they are mapped to a 256-dimensional semantic space through a 32-unit fully connected layer, and the weight matrix is ​​initialized using the Xavier normal distribution; finally, the output layer uses the Softmax function to generate an 8×1 gait sub-phase probability distribution vector.

7. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The weighted classification cross entropy loss function is introduced in step S7 to dynamically adjust the category weights, specifically including: Introducing weighted categorical cross entropy As a loss function, different categories are assigned weights that are proportional to the inverse of their sample numbers, which can be mathematically expressed as: (2) in, is the number of categories, Representation sample One-hot encoding value of the label. If the label is Class, then ,otherwise ; The model is The predicted probability of the class output is calculated by the Softmax function. It is Class weight; Class weight Calculated by the following formula: (3) in, Represents the training set The total number of samples in the class.

8. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The stratified K-fold cross-validation strategy described in step S8 specifically includes: When conducting formal model training, the dataset is divided into 10 non-overlapping folds, and each fold is divided into training set, validation set and test set in a ratio of 8:1:1, and the sample category distribution in each fold is guaranteed to be consistent with the original dataset; The hyperparameter setting and starting the model checkpoint mechanism described in step S8 specifically include: using the adaptive moment estimation Adam algorithm to optimize the model, the initial learning rate is set to 0.001, and dynamically adjusted according to the average loss; to prevent the model from overfitting, an early stopping mechanism is adopted, that is, when the loss of the validation set fails to improve within 10 consecutive cycles, the training is terminated in advance; the model checkpoint is enabled during the training process, and the model with the smallest validation set loss is automatically saved at the end of each round of training, so as to ensure that the best model parameters are captured and avoid performance degradation due to overfitting.

9. The fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as claimed in claim 1, characterized in that: The step S9 uses Simulink to deploy the trained spatiotemporal feature fusion model to build an end-to-end gait sub-phase online recognition system, which specifically includes: S91. Real-time data stream processing: Surface electromyography sensors and inertial measurement units are integrated on the lower limb rehabilitation exoskeleton robot, and 4-channel surface electromyography signals and 3-channel inertial measurement unit signals of both lower limbs are collected synchronously in real time at a sampling rate of 1000 Hz to form 14 channels of raw gait data; a circular buffer management strategy is adopted to set the data cache queue length to 100 milliseconds, and the raw gait data is divided into time windows by sliding overlapping window technology, with the window size set to 50 and the sliding step set to 10, and the same filtering and denoising processing as step S2 is performed on a window basis; then, the gait data after filtering and denoising is subjected to dynamic range normalization, that is, the normalization coefficient is updated in real time based on the maximum and minimum values ​​of the data in the last 10 seconds; finally, the input tensor dimension is configured as follows: batch size 128×time step 50×sensor channel 14×signal dimension 1, and all labels are subtracted by one to match the one-hot vector format; S92. Online model reasoning: Convert the trained spatiotemporal feature fusion model into ONNX format, and use MATLAB's importONNXNetwork function to verify the compatibility of the model's input and output nodes, tensor dimensions, and data types; build an end-to-end gait sub-phase real-time recognition system in Simulink, integrating the signal acquisition module, preprocessing module, data buffer module, model reasoning module, and result display module; load the ONNX format file to the model reasoning module through Deep Learning Toolbox, perform online model reasoning, and output an 8×1 gait sub-phase probability distribution vector; at the same time, enable CUDA acceleration and single sample batch processing mode; finally, perform maximum value decision on the gait sub-phase probability distribution vector output by the model, and its mathematical expression is: (4) in, represents the output label after executing the maximum value decision, Indicates time Input data Belongs to category The posterior probability of the maximum decision is obtained; the output label of the maximum decision is mapped from the numerical value to the gait sub-phase, and finally the real-time recognition result of the current gait sub-phase is obtained.

10. A fine-grained gait sub-phase recognition device based on spatiotemporal feature fusion, characterized in that: It includes a display interface, a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the fine-grained gait sub-phase recognition method based on spatiotemporal feature fusion as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Exoskeleton gait feature identification method based on multi-modal information fusion representation

    CN117235660A

  • Fetal three-vessel ultrasonic image segmentation method based on multi-scale feature collaborative learning

    CN117333500A