Gait analysis and anomaly detection method and system and medium
By constructing a joint modeling framework for spatiotemporal and time-series of multi-scale deep residual scaling network and Transformer timing encoder, the existing gait analysis technology has solved the problem of insufficient spatial and temporal modeling, noise processing and medical explanatory performance in space-time, and efficient gait analysis and disease diagnosis are achieved.
Patent Information
- Application Number
- CN202510566713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
The existing gait analysis technology has shortcomings in spatiotemporal modeling, noise processing, computing efficiency and medical interpretability, and it is difficult to meet the needs of accurate gait analysis and disease diagnosis.
A joint modeling framework for space-time based on ground reaction force signals is constructed, combining multi-scale depth residual scaling networks and Transformer timing encoders, and multi-scale feature extraction and global timing dependency modeling of GRF signals through time slice segmentation and multi-head self-attention mechanisms.
It improves the accuracy and robustness of gait analysis, enhances the medical interpretation ability of gait characteristics, is suitable for resource-constrained equipment, and supports early detection and evaluation of diseases such as Parkinson's disease.
Smart Images

Figure CN120337009A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - field of medical signal processing and artificial intelligence. Specifically, it relates to a gait analysis and anomaly detection method, system and medium based on ground reaction force signals and spatio - temporal feature learning. Background Art
[0002] Gait analysis evaluates health status by quantifying human movement behavior and has become an important means for the research of neurodegenerative diseases and movement disorders. Ground reaction force (GRF) signals, which can capture the interaction between plantar pressure and the ground, are a key feature source for clinical gait analysis and play an important role in the diagnosis and rehabilitation of diseases such as Parkinson's disease and musculoskeletal lesions, and can assist in differentiating patient groups and evaluating disease severity.
[0003] Traditional gait analysis uses classical machine learning techniques such as support vector machines (SVM), k - nearest neighbor classifiers (KNN), and artificial neural networks (ANN), which rely on manually extracting multi - dimensional features in the time - frequency domain for classification. However, these methods have obvious limitations: feature engineering requires a large amount of manual annotation and selection; they are insufficient in modeling the variation law of the time dimension of gait signals; manual feature extraction is prone to the curse of dimensionality and difficult to process large - scale and diverse data.
[0004] In recent years, deep learning techniques have brought new breakthroughs to gait analysis. Convolutional neural networks (CNN) can capture local spatial features of data, and Transformer models can effectively model global time - dependent information. However, the existing methods still have the following technical bottlenecks:
[0005] 1. Limitations in modeling spatio - temporal characteristics: GRF signals have both high - dimensional time series and spatio - temporal correlations. For example, it not only shows rapid dynamic changes in a short time locally but also exhibits global time - dependence over the gait cycle. Existing models mostly process the time and space dimensions independently and lack the ability of joint modeling. Limited by the fixed receptive field of the convolution kernel, traditional methods are difficult to capture the feature consistency of different signal lengths or gait phases, resulting in insufficient representation of gait periodic laws.
[0006] 2. Insufficient signal noise suppression and redundant feature elimination: GRF signals are vulnerable to device differences, individual characteristics, and experimental environment interference. However, existing methods lack the ability of dynamic noise suppression and multi - scale feature mining, and common padding or truncation operations will exacerbate noise propagation and reduce classification accuracy.
[0007] 3. Poor real - time computing and model adaptability: The application demand for gait analysis in mobile or embedded devices (such as in telemedicine scenarios) is increasing day by day. However, existing deep network models consume a large amount of computing resources and are difficult to meet the deployment requirements of low power consumption and high real - time performance.
[0008] 4. Insufficient medical interpretability: Most gait classification models based on deep learning are in the "black box" mode, unable to intuitively show the specific reasons for abnormal gaits, difficult to locate the disease-affected areas or kinetic characteristics, severely restricting their clinical application value.
[0009] Through the retrieval of patent literature, it is found that the invention patent with the publication number CN 116473514 discloses a gait Parkinson's disease detection based on an adaptive directed spatio-temporal graph neural network. The detection steps are as follows: S1. Signal preprocessing: Divide the obtained gait signal into a predetermined number of time steps; S2. Model construction: Model the topological structure of the plantar sensors and process the signals obtained by the sensors into a two-stream modality; S3. Feature extraction network: Through multiple adaptive directed spatio-temporal graph neural network units, use the message passing mechanism in space to obtain local and global information of the sole, and use 1D convolution in time to obtain temporal information, so as to analyze gait changes in the spatio-temporal domain; S4. Classifier: Adopt the cross-entropy loss function as the classifier; S5. Model fusion: Perform linear fusion on the two-stream modality; S6. Diagnostic result: Average all the segmented results of the subjects to obtain the final diagnostic result. This patent only adopts the two-stream modality in model construction, which is relatively single. In terms of means to improve model performance, it does not involve non-linear signal representation and optimization of computational efficiency.
[0010] In summary, aiming at the problems of the above-mentioned existing technologies, researching a gait analysis and anomaly detection method, system and medium based on ground reaction force signals and spatio-temporal feature learning has become a key task that needs to be solved urgently at present. Summary of the Invention
[0011] Aiming at the defects in the existing technologies, the purpose of the present invention is to provide a gait analysis and anomaly detection method, system and medium based on ground reaction force signals and spatio-temporal feature learning.
[0012] According to a gait analysis and anomaly detection method provided by the present invention, it includes the following steps:
[0013] Step S1, preprocess the original GRF signal, and divide the preprocessed GRF signal into a time slice sequence of a fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle;
[0014] Step S2, input the time slice sequence into a multi-scale deep residual scaling network to extract multi-scale spatio-temporal features;
[0015] Step S3, splice the multi-scale spatio-temporal features into a temporal feature sequence in chronological order, and input the temporal feature sequence into a Transformer encoder, use the multi-head self-attention mechanism to model the global temporal dependence relationship, and at the same time adopt a masking mechanism to mask invalid time steps, and output global time series features;
[0016] Step S4: Classify the global time series features and output the gait classification probability distribution for Parkinson's disease recognition, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
[0017] Preferably, step S1 includes the following sub-steps:
[0018] Step S1.1: Obtain the original GRF signal X = {X1, X2, …, X N}, where N represents the number of samples, and X i represents the GRF signal of the i-th sample. Each GRF signal contains C channels, denoted as representing the GRF signal of the c-th channel. The signal of each channel where L i represents the length of the original GRF signal of the i-th sample;
[0019] Step S1.2: Through a padding operation on all the original GRF signals, normalize the length to a fixed length L to obtain the preprocessed GRF signal;
[0020] Step S1.3: According to the sampling rate and the number of data points, divide the preprocessed GRF signal into T time slices, denoted as where the number of slices T = L / L s and L s is the predetermined time slice length, representing the t-th GRF time slice signal of the c-th channel of the i-th sample.
[0021] Preferably, step S2 includes the following sub-steps:
[0022] Step S2.1: Input the time slice sequence into the multi-scale deep residual scaling network. Extract local detail features through the 5×5 convolutional kernel DRSN-CW branch of the multi-scale deep residual scaling network, and extract overall dynamic characteristics through the 17×17 convolutional kernel DRSN-CW branch of the multi-scale deep residual scaling network;
[0023] Step S2.2: Concatenate the features output by the 5×5 convolutional kernel DRSN-CW branch and the 17×17 convolutional kernel DRSN-CW branch respectively to obtain the multi-scale spatio-temporal features.
[0024] Preferably, each DRSN-CW branch in step S2.1 includes stacked RSBU-CW modules, and the RSBU-CW module includes the following operations:
[0025] Step S2.1.1: Perform global average pooling operation on the input time slice s, compress it into a one-dimensional vector, and input the one-dimensional vector into a two-layer fully connected network to calculate the scaling parameter α c :
[0026]
[0027] where z c is the output eigenvalue of the FC network corresponding to the c-th channel;
[0028] Step S2.1.2: Calculate the channel threshold τ according to formula (2) c :
[0029] τ c = α c · average i,j |x i,j,c | (2)
[0030] where average i,j represents the absolute average value of the feature map on channel c, and i and j are the indices of width and height respectively, and x i,j,c is the eigenvalue of the feature map at the c-th channel, width position i, and height position j;
[0031] Step S2.1.3: Use the channel threshold to perform soft threshold shrinkage operation on each position of the input s, defined as follows:
[0032]
[0033] where x represents the feature map (the original GRF signal at the first layer), and y represents the output feature after shrinkage;
[0034] Step S2.1.4: Inside each DRSN-CW, the RSBU-CW modules are connected in series through a residual connection strategy. For a certain residual unit, the feature transfer form is:
[0035] H(s) = F(s,{W i}) + s (4)
[0036] where H(s) represents the output of the residual unit, F(s,{W i}) represents the feature transformation through the RSBU-CW module, s represents the input feature map of this residual unit, and W i represents all the learnable parameters in the RSBU-CW module.
[0037] Preferably, in step S2.2, the features F small (x) and Flarge (x) The formula for splicing multi-scale spatio-temporal features is as follows:
[0038] F multi (s) = concat(F small (s), F large (s)) (5).
[0039] Preferably, step S3 includes the following sub-steps:
[0040] Step S3.1, splicing the multi-scale spatio-temporal features into a time series feature sequence in chronological order;
[0041] Step S3.2, inputting the time series feature sequence into a Transformer encoder,
[0042] Step S3.3, calculating the attention weights at each time step through the multi-head self-attention layer of the Transformer encoder;
[0043] Step S3.4, performing a masking operation based on the padding position to output the global time series features.
[0044] Preferably, in step S3.3, the input of the self-attention layer includes a query Q, a key K, and a value V, and the self-attention calculation method is:
[0045]
[0046] where d k is a scaling factor depending on the layer size, and Q, K, and V are respectively obtained by linear transformation of the input sequence features.
[0047] Preferably, in step S4, the global time series features are mapped through a multi-layer fully connected layer to output a gait classification probability distribution, and the gait probability distribution is used for Parkinson's disease identification, Parkinson's severity assessment, and musculoskeletal injury gait classification.
[0048] The present invention also provides a gait analysis and anomaly detection system, including:
[0049] A GRF time segmentation module that preprocesses the original GRF signal and divides the preprocessed GRF signal into a time slice sequence of a fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle;
[0050] A multi-scale deep residual shrinkage network (DRSN) module that inputs the time slice sequence into a multi-scale deep residual scaling network to extract multi-scale spatio-temporal features;
[0051] The Transformer time series encoder module concatenates multi-scale spatio-temporal features in chronological order into a time series feature sequence, and inputs the time series feature sequence into the Transformer encoder. The multi-head self-attention mechanism is used to model the global time series dependencies, and at the same time, a masking mechanism is adopted to mask invalid time steps, and the global time series features are output;
[0052] The downstream task classification module classifies the global time series features and outputs the gait classification probability distribution, which is used for Parkinson's disease identification, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
[0053] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned gait analysis and anomaly detection method are realized.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. The present invention proposes a joint modeling framework that combines time dynamic representation and spatial feature extraction, introduces a segmentation mechanism based on time slices, reasonably organizes variable-length signals, and realizes flexible mapping and representation of different gait cycles.
[0056] 2. The present invention designs a multi-scale robust feature extraction network, strengthens the noise suppression function, and enhances the combination ability of global and local features under different receptive fields.
[0057] 3. The present invention optimizes the model structure, constructs a lightweight computing framework, combines complex feature processing with real-time requirements, and provides efficient adaptation support for resource-constrained devices.
[0058] 4. The present invention uses the attention mechanism to mine gait cycle information, reveals the contributions of different gait phases to the classification task, improves the medical interpretability of the model, and provides a transparent basis for medical diagnosis.
[0059] In summary, the present invention comprehensively models spatio-temporal characteristics, captures periodic change laws by combining gait anatomical characteristics; constructs a robust feature extraction method to suppress noise and enhance the spatial and dynamic characteristics of classification; improves the adaptability and generalization ability in large-scale and multi-scenario data; enhances the medical interpretability of gait feature modeling, thereby improving the efficiency and diagnostic accuracy of gait analysis, providing solid technical support for the early detection, evaluation and treatment of movement-related diseases such as Parkinson's disease, and comprehensively improving the performance and practicality of gait analysis technology. Description of the Drawings
[0060] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more apparent:
[0061] Figure 1 This is the overall architecture diagram of T-GaitNet in the embodiments of the present invention;
[0062] Figure 2 This is the schematic diagram of GRF time slices in the embodiments of the present invention;
[0063] Figure 3 This is the architecture diagram of DRSN-CW in the embodiments of the present invention;
[0064] Figure 4 This is the schematic diagram of the attention mechanism in the embodiments of the present invention. Detailed implementation manners
[0065] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.
[0066] In the field of gait analysis, existing ground reaction force (GRF) signal gait analysis technologies have deficiencies in spatio-temporal modeling, noise processing, computational efficiency, and medical interpretability. The present invention aims to propose a gait analysis and anomaly detection method, system, and medium based on ground reaction force signals and spatio-temporal feature learning. Through innovative technical means, accurate identification of abnormal gaits and disease detection are realized, providing new technical support for clinical diagnosis, rehabilitation evaluation, and sports science research.
[0067] GRF data has the characteristics of multi-dimensional and serialized, and there are complex problems such as inconsistent signal lengths and large inter-individual differences. This makes it difficult for traditional methods to effectively extract spatial and temporal features therein, and cannot meet the needs of accurate gait analysis and disease diagnosis.
[0068] The core points of the present invention are as follows:
[0069] 1. Construct a spatio-temporal joint modeling framework based on GRF signals, design a new deep learning framework T-GaitNet, and combine a multi-scale deep residual scaling network (MS-DRSN) and a Transformer temporal encoder. Through this modular integration, the spatial domain and global time dependencies are uniformly modeled, comprehensively integrating the local and global characteristics of GRF signals, thereby improving the accuracy and robustness of abnormal gait detection and disease diagnosis.
[0070] 2. Design a time slice segmentation and signal segmentation mechanism. In the T-GaitNet framework, the GRF time segmentation module is one of the core innovations. After padding the GRF signals of non-fixed length to a unified length, this module divides them into multiple time slices, with each slice corresponding to a specific phase of the gait cycle, forming a GRF time slice sequence of fixed length. This mechanism can not only analyze gait signals of different lengths but also maintain the integrity of gait cycle features, enhance the feature expression ability of different gait phases, and is applicable to the analysis of signals of various lengths.
[0071] 3. MS-DRSN plays a key role in T-GaitNet. It captures the dynamic characteristics of GRF signals within different time windows through a multi-scale convolution module, combines a soft threshold shrinkage mechanism to suppress signal noise, enhances the robustness and expression ability of key features, and realizes fine-grained modeling of spatial domain features.
[0072] 4. The Transformer time series encoder applies the multi-head self-attention mechanism to model the temporal dependencies of GRF signals globally, and uses a masking mechanism to mask the invalid time steps generated by padding, ensuring that the integrity features and dynamic patterns of the gait cycle are accurately mined. This method innovatively uses the multi-head self-attention mechanism and the masking mechanism to solve the problem of processing variable-length gait signals.
[0073] 5. Establish a medical interpretability enhancement mechanism. Through attention weight visualization and dynamic analysis of the gait cycle, it can locate the key time points and dynamic features in abnormal gait signals. This visualization technology enhances clinical applicability, provides a transparent basis for disease diagnosis, and helps to locate the pathological features of gait abnormalities.
[0074] The technical solution of the present invention, T-GaitNet, is as follows:
[0075] The overall architecture of T-GaitNet consists of four core modules, namely the GRF time segmentation module, the multi-scale deep residual shrinkage network (DRSN) module, the Transformer time series encoder module, and the downstream task classification module. Each module works together to achieve efficient analysis of GRF signals and disease diagnosis.
[0076] The GRF time segmentation module is used to enhance the T-GaitNet's ability to model the time dimension of GRF signals and capture the gait cycle pattern. It divides the GRF signals padded to a unified length into multiple time slices, with each slice corresponding to a specific phase of the gait cycle, thus structuring the global signal into a sequence of GRF time slices of a fixed length. In this way, the periodic pattern of the gait signal can be retained by concatenating multiple time slices, and at the same time, the local features of the gait at different phases can be focused on for extraction. Meanwhile, the attention mechanism in T-GaitNet can intuitively reflect the contribution of different gait periods to the final prediction, providing interpretability for gait recognition and analysis. Moreover, for signals of different lengths, this module can flexibly adjust the division granularity of time segments. When dealing with long signals, it avoids the problem of feature dimension explosion caused by the fixed-size convolutional kernel in traditional CNNs; when dealing with short signals, it can continuously extract rich local spatial features.
[0077] The divided time segments are fed into the multi-scale deep residual shrinkage network (MS-DRSN) module. The MS-DRSN module realizes the robust extraction of GRF features through multi-scale convolution, and at the same time adopts an adaptive shrinkage mechanism (i.e., the soft threshold shrinkage mechanism) to dynamically suppress noise, further enhancing the effective information of the signal. This module takes into account the extraction of spatial features of local details and global dependencies, effectively strengthening the extraction of effective information of GRF signals by T-GaitNet.
[0078] The feature combination generated by the MS-DRSN module is input into the Transformer temporal encoder module. The Transformer temporal encoder module uses the multi-head self-attention mechanism to extract the temporal dependence relationship and gait transition pattern in the GRF signal. In addition, since directly padding or truncating GRF signals of different lengths will cause information redundancy or loss, this module uses a masking mechanism to automatically ignore the invalid time steps generated by padding when extracting temporal features, thereby improving the modeling accuracy of T-GaitNet.
[0079] Finally, the downstream task classification module maps the global temporal representation output by the Transformer temporal encoder module to specific gait diagnosis and classification problems, realizing functions such as the diagnosis of Parkinson's disease and the recognition of abnormal gaits.
[0080] Example 1:
[0081] Figure 1 This is the overall architecture diagram of T-GaitNet in the embodiments of the present invention.
[0082] As Figure 1 shown, this embodiment provides a gait analysis and anomaly detection method, including the following steps:
[0083] Step S1. To model the gait cycle information of the ground reaction force (GRF) signal in the time dimension, the original GRF signal is preprocessed and divided into a sequence of time slices of a fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle.
[0084] Figure 2 This is a schematic diagram of the GRF time slice in the embodiment of the present invention.
[0085] As Figure 2 shows how to divide the GRF time slice from the original GRF signal. Specifically, step S1 includes the following sub-steps:
[0086] Step S1.1. Obtain the original GRF signal X = {X1, X2, …, X N}, where N represents the number of samples, and X i represents the GRF signal of the i-th sample. Each GRF signal contains C channels (corresponding to different sensors or different GRF directions), denoted as represents the GRF signal of the c-th channel. The signal of each channel where L i represents the length of the original GRF signal of the i-th sample; in actual gait data acquisition, due to the differences in the acquisition duration and the gait characteristics of patients, the lengths of the original GRF signals of different samples are often different.
[0087] Step S1.2. To facilitate the unified processing of the GRF signal by the T-GaitNet model, all the original GRF signals are padded to normalize the length to a fixed length L to obtain the preprocessed GRF signal.
[0088] Step S1.3. According to the sampling rate and the number of data points, the preprocessed GRF signal is divided into T time slices, denoted as where the number of slices T = L / L s and L s is the predetermined length of the time slice, represents the t-th GRF time slice signal of the c-th channel of the i-th sample.
[0089] The rationality of the time slice division is mainly reflected in two aspects: gait medical characteristics and calculation methods:
[0090] First, from a biomedical perspective, the division of time slices corresponds to specific phases in the gait cycle, enabling each time slice to focus on extracting local features of that phase. By sequentially concatenating multiple time slices, not only can the overall periodic pattern of the GRF signal be retained, but also the dynamic changes between phases can be deeply captured, thereby more accurately revealing the gait pattern.
[0091] Second, in terms of computational methods, dividing time slices can effectively avoid the curse of dimensionality problem when processing long GRF signals. Additionally, by reasonably setting the division granularity, a balance can be achieved between retaining signal details and reducing computational overhead. Shorter time slices can more sensitively capture the rapidly changing features in the GRF signal, while longer time slices can better highlight the periodic pattern of the GRF signal.
[0092] However, in the prior art, the continuous raw GRF signal sequence of all sensors is directly used as input without slicing the raw GRF signal into shards of specific periods. This approach simplifies the preprocessing process compared to this embodiment. However, since the raw GRF signal is not preprocessed, directly inputting high-dimensional time series may lead to an increase in computational complexity, especially in applications on large-scale datasets. In addition, different from the gait cycle features of the GRF signal corresponding to the time slices in this embodiment, the raw GRF signal lacks clear time segment separation information, which may not be conducive to clinicians' understanding and interpretation of pathological features.
[0093] Step S2: Input the time slice sequence into a multi-scale deep residual scaling network (MS-DRSN) to extract multi-scale spatio-temporal features.
[0094] In this embodiment, the GRF signal has complex spatio-temporal characteristics and multi-scale dynamic patterns. However, during the acquisition process, noise interference will significantly reduce the quality of feature representation of the GRF signal. To more effectively extract the spatial features of GRF time slices, the present invention designs a multi-scale deep residual scaling network (MS-DRSN), which further optimizes the T-GaitNet model's ability to understand complex sequence signals by combining feature patterns under different receptive fields and a channel adaptive soft threshold shrinkage mechanism.
[0095] Specifically, step S2 includes the following sub-steps:
[0096] Step S2.1: Input the time slice sequence into a multi-scale deep residual scaling network (MS-DRSN). Extract local detail features through the DRSN-CW branch with a 5×5 convolutional kernel of the multi-scale deep residual scaling network, and extract overall dynamic characteristics through the DRSN-CW branch with a 17×17 convolutional kernel of the multi-scale deep residual scaling network.
[0097] In this embodiment, MS-DRSN consists of two parallel Deep Residual Scaling Networks (DRSN-CW). The first DRSN-CW branch uses a 5×5 convolutional kernel to emphasize the detailed features of local time segments, while the second DRSN-CW branch uses a 17×17 convolutional kernel to capture the overall dynamic characteristics of time segments.
[0098] Specifically, each DRSN-CW branch is based on the ResNet-18 framework, and the residual blocks therein are replaced by stacked RSBU-CW modules (Residual Shrinkage Building Unit with Channel-wisethresholds) to enhance the anti-noise performance of MS-DRSN for feature samples.
[0099] The RSBU-CW module is the basic building block of DRSN-CW, and its main function is to combine the soft-thresholding shrinkage mechanism to filter out redundant features and enhance significant features through adaptively calculated channel thresholds.
[0100] Figure 3 It is the architecture diagram of DRSN-CW in the embodiment of the present invention.
[0101] As Figure 3 shown, the RSBU-CW module includes the following operations:
[0102] Step S2.1.1: Perform global average pooling (GAP) on the input time slice s to compress it into a one-dimensional vector, and input the one-dimensional vector into a two-layer fully connected (FC) network to calculate the scaling parameter α c :
[0103]
[0104] where z c is the output feature value corresponding to the c-th channel of the FC network.
[0105] Step S2.1.2: Calculate the channel threshold τ according to formula (2) c :
[0106] τ c = α c ·average i,j |x i,j,c | (2)
[0107] where average i,j represents the absolute average value of the feature map on channel c, and i and j are the indices of width and height respectively, and x i,j,c is the feature value of the feature map at the c-th channel, width position i, and height position j.
[0108] Step S2.1.3: Perform a soft-thresholding shrinkage operation on each position of the input s using the channel threshold, defined as follows:
[0109]
[0110] where x represents the feature map (the original GRF signal in the first layer), and y represents the output feature after shrinkage.
[0111] Such a design can achieve flexible denoising between channels, thereby enhancing the ability of MS-DRSN to capture important features.
[0112] Step S2.1.4: Inside each DRSN-CW, the RSBU-CW modules are connected in series through a residual connection strategy, which ensures that signal features can be effectively retained during transmission and significantly accelerates the network optimization process.
[0113] For a certain residual unit, the form of feature transmission is:
[0114] H(s) = F(s, {W i ) + s (4)
[0115] where H(s) represents the output of the residual unit, F(s, {W i ) represents the feature transformation through the RSBU-CW module, s represents the input feature map of this residual unit, and W i represents all the learnable parameters in the RSBU-CW module.
[0116] Step S2.2: Concatenate the features output by the 5×5 convolutional kernel DRSN-CW branch and the 17×17 convolutional kernel DRSN-CW branch respectively to obtain multi-scale spatio-temporal features.
[0117] Specifically, in Step S2.2, the formula for concatenating the features F small (x) and F large (x) output by the 5×5 convolutional kernel DRSN-CW branch and the 17×17 convolutional kernel DRSN-CW branch respectively to obtain multi-scale spatio-temporal features is as follows:
[0118] F multi (s) = concat(F small (s), F large (s)) (5).
[0119] However, in the existing technology, the alternative methods of the multi-scale deep residual scaling network (MS-DRSN) mainly include some classic convolutional neural networks, such as GoogLeNet, ResNet, standard convolutional network (CNN), etc. These classic networks have shown robust performance in extracting spatial features and dealing with complex tasks, but there are some deficiencies in dealing with the multi-scale characteristics and noise interference of GRF data. The following is an analysis of these alternative solutions:
[0120] First of all, GoogLeNet, with its structure design based on the Inception module, can integrate multi-scale convolutional kernels in a model to extract features under different receptive fields, which is similar to the multi-scale characteristics in MS-DRSN. However, the structure of GoogLeNet is relatively fixed. When dealing with the tasks of time and space mixed associations in GRF signals, it is difficult to flexibly adjust the size of the convolutional kernel, and the noise suppression performance for processing non-linear plantar pressure data is insufficient. In addition, the computational complexity of GoogLeNet is relatively high, which is not friendly to resource-constrained devices.
[0121] Secondly, ResNet relies on the design of residual skip connections, which can effectively alleviate the gradient vanishing problem when dealing with deep networks and demonstrate powerful feature extraction capabilities. However, the performance of ResNet depends on the residual units with a fixed structure. Its convolutional kernels are usually used to extract local features, and the modeling of the global multi-scale characteristics of plantar pressure data and noise suppression are weak. Especially when dealing with the irregularity and length change of GRF signals, it shows certain adaptability deficiencies. In addition, the classic ResNet does not consider the dynamic feature shrinkage mechanism, which is a disadvantage for significantly enhancing the noise processing ability and improving the feature robustness.
[0122] The standard convolutional neural network (CNN) is a type of basic network structure, which is usually good at extracting local spatial features from fixed-length signals. Although its simple structure design is easy to implement and has a certain degree of flexibility, most CNNs are based on convolutional kernels of fixed size, with a limited range of action, and it is difficult to capture both the global characteristics of gait with long-term changes and the local features with short-term changes at the same time. In addition, CNN lacks a dynamic screening mechanism for invalid information and redundant features in the data, and additional preprocessing steps are required to alleviate the noise problem.
[0123] Step S3: Concatenate the multi-scale spatio-temporal features in chronological order to form a time series feature sequence, and input the time series feature sequence into the Transformer encoder. Use the multi-head self-attention mechanism to model the global time series dependencies, and at the same time adopt a masking mechanism to mask the invalid time steps, and output the global time series features.
[0124] In this embodiment, step S3 includes the following sub-steps:
[0125] Step S3.1: Concatenate the multi-scale spatio-temporal features in chronological order to form a time-series feature sequence.
[0126] Specifically, after feature extraction by MS-DRSN, the feature vectors of all time slices are concatenated in chronological order to generate a unified time-series representation. In the time series, each feature vector corresponds to a time slice of the GRF signal, and the length of the time series is equal to the number of time slices.
[0127] Step S3.2: Input the time-series feature sequence into the Transformer encoder.
[0128] In this embodiment, the Transformer encoder is used to model the dynamic dependencies between time steps and the gait transition rules. The Transformer encoder includes multiple layers of multi-head self-attention layers and feed-forward neural networks, and its core part is the multi-head self-attention layer.
[0129] Step S3.3: Calculate the attention weights for each time step through the multi-head self-attention layer of the Transformer encoder.
[0130] Figure 4 This is the schematic diagram of the attention mechanism in the embodiment of the present invention.
[0131] As Figure 4 shown, the input of the self-attention layer includes query Q, key K, and value V, and the self-attention calculation method is:
[0132]
[0133] where d k is a scaling factor depending on the layer size, and Q, K, and V are obtained by linear transformation from the input sequence features respectively. This mechanism enables the Transformer encoder to learn the weighted relationships between different time steps, thereby effectively capturing the temporal dependencies of the time series.
[0134] Step S3.4: For the invalid time steps introduced by time alignment padding in the time series, perform a masking operation based on the padding positions to output the global time-series features. This masking operation ensures that the features at the masked positions do not affect subsequent attention allocation or feature learning, thereby further improving the ability of the Transformer encoder to extract key features of the GRF signal.
[0135] However, in the prior art, alternative solutions such as Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), or Bidirectional Long Short-Term Memory Network (BiLSTM) are good at handling time-series tasks, but still have obvious deficiencies compared with the Transformer encoder. When dealing with long-term dependence problems, RNN is prone to gradient vanishing or gradient explosion, and cannot perform parallel computing, resulting in low efficiency; LSTM effectively improves long-time series dependence modeling through gating mechanisms, but has a complex structure and high computational cost, especially lacking flexibility when dealing with GRF signals of non-fixed length; as a simplified version of LSTM, GRU has improved computational efficiency, but its feature modeling ability is limited, and its ability to represent non-linear complex dynamics is weak; BiLSTM enhances the extraction ability through bidirectional processing, but the computational overhead is doubled, and it still suffers from the limitation of step-by-step calculation over time. In addition, these methods lack a masking mechanism compared with the Transformer encoder, and cannot dynamically mask invalid time steps, thus lacking a complete periodic modeling of GRF gait signals and accurate positioning of abnormal gaits.
[0136] Step S4: Classify the global time series features and output the gait classification probability distribution for Parkinson's disease recognition, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
[0137] Specifically, map the global time series features through a multi-layer fully connected layer to output the gait classification probability distribution, which is used for Parkinson's disease recognition, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
[0138] However, in the prior art, traditional machine learning classifiers such as Support Vector Machine (SVM) and k-Nearest Neighbor algorithm (KNN) are used as alternative solutions. SVM and KNN directly classify the extracted plantar pressure features, and have the advantages of simple implementation and high computational efficiency, especially performing well in scenarios with small-scale datasets and limited feature quantities. However, compared with the downstream classification module based on deep learning methods in this embodiment, SVM and KNN are limited in dealing with high-dimensional complex features and large-scale data, and are insufficient in capturing the non-linear relationships and spatio-temporal dynamic patterns between complex features, resulting in lower recognition accuracy for Parkinson's disease recognition, severity assessment, and gait classification of musculoskeletal injuries.
[0139] Embodiment 2:
[0140] The present invention also provides a gait analysis and anomaly detection system, which can be implemented by executing the process steps of the gait analysis and anomaly detection method. That is, those skilled in the art can understand the gait analysis and anomaly detection method as a preferred implementation manner of the gait analysis and anomaly detection system.
[0141] Specifically, the gait analysis and anomaly detection system includes:
[0142] The GRF time segmentation module preprocesses the original GRF signal and divides the preprocessed GRF signal into a sequence of time slices of fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle.
[0143] The multi-scale deep residual shrinkage network (DRSN) module inputs the time slice sequence into the multi-scale deep residual scaling network to extract multi-scale spatio-temporal features.
[0144] The Transformer temporal encoder module concatenates the multi-scale spatio-temporal features in chronological order into a temporal feature sequence, inputs the temporal feature sequence into the Transformer encoder, models the global temporal dependence using the multi-head self-attention mechanism, and at the same time uses the masking mechanism to mask invalid time steps, and outputs the global time series features.
[0145] The downstream task classification module classifies the global time series features and outputs the gait classification probability distribution for Parkinson's disease identification, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
[0146] Specifically, the GRF time segmentation module includes the following sub-modules:
[0147] Module M1.1 obtains the original GRF signal X = {X1, X2, …, X N}, where N represents the number of samples, and X i represents the GRF signal of the i-th sample. Each GRF signal contains C channels (corresponding to different sensors or different GRF directions), denoted as represents the GRF signal of the c-th channel, and the signal of each channel where L i represents the length of the original GRF signal of the i-th sample; in actual gait data acquisition, due to the differences in acquisition duration and patient gait characteristics, the lengths of the original GRF signals of different samples are often different.
[0148] Module M1.2, to facilitate the unified processing of GRF signals by the T-GaitNet model, all original GRF signals are padded to normalize the length to a fixed length L to obtain the preprocessed GRF signal.
[0149] Module M1.3 divides the preprocessed GRF signal into T time slices according to the sampling rate and the number of data points, denoted as where the number of slices T = L / L s , L sis the length of a predetermined time slice, represents the t-th GRF time slice signal of the c-th channel of the i-th sample.
[0150] Specifically, the multi-scale deep residual shrinkage network (DRSN) module includes the following sub-modules:
[0151] Module M2.1 inputs the time slice sequence into the multi-scale deep residual scaling network (MS-DRSN), extracts local detail features through the 5×5 convolutional kernel DRSN-CW branch of the multi-scale deep residual scaling network, and extracts overall dynamic characteristics through the 17×17 convolutional kernel DRSN-CW branch of the multi-scale deep residual scaling network.
[0152] RSBU-CW includes the following operations:
[0153] Module M2.1.1 performs a global average pooling (GAP) operation on the input time slice s, compresses it into a one-dimensional vector, and inputs the one-dimensional vector into a two-layer fully connected (FC) network to calculate the scaling parameter α c :
[0154]
[0155] where z c is the output eigenvalue corresponding to the c-th channel of the FC network.
[0156] Module M2.1.2 calculates the channel threshold τ according to formula (2) c :
[0157] τ c = α c · average i,j |x i,j,c | (2)
[0158] where average i,j represents the absolute average value of the feature map on channel c, and i and j are the indices of width and height respectively, and x i,j,c is the eigenvalue of the feature map at the c-th channel, with width position i and height position j.
[0159] Module M2.1.3 performs a soft threshold shrinkage operation on each position of the input s using the channel threshold, defined as follows:
[0160]
[0161] where x represents the feature map (the original GRF signal in the first layer), and y represents the output feature after shrinkage.
[0162] Such a design can achieve flexible noise cancellation between channels, thereby enhancing the ability of MS-DRSN to capture important features.
[0163] Module M2.1.4, inside each DRSN-CW, the RSBU-CW modules are connected in series through a residual connection strategy, which ensures that signal features can be effectively retained during transmission and significantly accelerates the network optimization process.
[0164] For a certain residual unit, the form of feature transmission is:
[0165] H(s) = F(s, {W i ) + s (4)
[0166] where H(s) represents the output of the residual unit, F(s, {W i ) represents the feature transformation through the RSBU-CW module, s represents the input feature map of this residual unit, and W i represents all the learnable parameters in the RSBU-CW module.
[0167] Module M2.2, concatenates the features output by the 5×5 convolutional kernel DRSN-CW branch and the 17×17 convolutional kernel DRSN-CW branch respectively to obtain multi-scale spatio-temporal features.
[0168] Specifically, in module M2.2, the features F small (x) and F large (x) output by the 5×5 convolutional kernel DRSN-CW branch and the 17×17 convolutional kernel DRSN-CW branch respectively are concatenated, and the formula for the multi-scale spatio-temporal features is as follows:
[0169] F multi (s) = concat(F small (s), F large (s)) (5).
[0170] In this embodiment, the Transformer temporal encoder module includes the following sub-modules:
[0171] Module M3.1, concatenates the multi-scale spatio-temporal features in chronological order into a temporal feature sequence.
[0172] Specifically, after the feature extraction by MS-DRSN, the feature vectors of all time slices are concatenated in chronological order to generate a unified time series representation. In the time series, each feature vector corresponds to a time slice of the GRF signal, and the length of the time series is equal to the number of time slices.
[0173] Module M3.2, inputs the temporal feature sequence into the Transformer encoder.
[0174] In this embodiment, the Transformer encoder is used to model the dynamic dependencies between time steps and the gait transition rules. The Transformer encoder includes multiple layers of multi-head self-attention layers and a feed-forward neural network, and its core part is the multi-head self-attention layer.
[0175] Module M3.3 calculates the attention weights for each time step through the multi-head self-attention layer of the Transformer encoder.
[0176] The inputs of the self-attention layer include query Q, key K, and value V, and the self-attention calculation method is:
[0177]
[0178] where d k is a scaling factor depending on the layer size, and Q, K, and V are respectively obtained by linear transformation of the input sequence features. This mechanism enables the Transformer encoder to learn the weighted relationships between different time steps, thereby effectively capturing the temporal dependencies of the time series.
[0179] Module M3.4 performs a masking operation based on the padding position for the invalid time steps introduced by time alignment padding in the time series, and outputs the global time series features. This masking operation ensures that the features at the masked positions do not affect subsequent attention allocation or feature learning, thereby further improving the ability of the Transformer encoder to extract key features of the GRF signal.
[0180] Embodiment 3:
[0181] This embodiment provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps of a gait analysis and anomaly detection method in the above Embodiment 1.
[0182] Comparative example:
[0183] In this study, we evaluated the performance of T-GaitNet on two public datasets: the PhysioBank gait dataset and the GaitRec gait dataset. The PhysioBank dataset mainly focuses on the gait patterns of Parkinson's disease patients, and we conducted two classification tasks: a binary classification task of Parkinson's disease patients and healthy individuals, and a four-class classification task of the severity of Parkinson's disease. The GaitRec dataset contains gait data of healthy individuals and individuals with musculoskeletal injuries. On the GaitRec dataset, we conducted a four-class classification task (hip, knee, ankle, and calcaneus) and a five-class classification task (normal, hip, knee, ankle, and calcaneus) respectively.
[0184] The PhysioBank database collected gait measurement data from 93 idiopathic PD patients and 73 healthy controls in three studies, for a total of 306 records. These three studies respectively investigated the effects of dual-task, rhythmic auditory stimulation (RAS), and treadmill walking on gait, covering gait characteristics under various experimental conditions. The database also includes demographic information and the Hoehn & Yahr staging of disease severity. The gait data in this database were collected by 8 sensors (Ultraflex Computer Dyno Graphy, Infotronic Inc.) under each foot, which recorded the vertical GRF (unit: Newton) over time as the subjects walked on flat ground at their self-selected speed for about 2 minutes. The left-foot sensors were sequentially numbered from L1 to L8, and the right-foot sensors were numbered from R1 to R8.
[0185] The GaitRec dataset includes bilateral GRF gait data of 2,084 patients with various musculoskeletal disorders and 211 healthy controls. Two force plates (Kistler, Type 9281B12, Winterthur, CH) recorded the center of pressure (COP) and GRF signals in three degrees of freedom, namely vertical, anteroposterior, and mediolateral directions. The patient data covered their walking tests during hospitalization in the rehabilitation center, involving joint replacement, fractures, ligament ruptures, and diseases related to the hip, knee, ankle, or calcaneus, with a total of 75,732 bilateral walking records. These data were manually annotated by physical therapists with more than a decade of clinical experience according to the existing medical diagnoses of each patient, and were divided into two major categories: healthy controls and gait disorders. Among them, gait disorders were further subdivided into four categories according to the anatomical location of the injury, including the hip, knee, ankle, and calcaneus.
[0186] The baseline methods for comparison include three traditional machine learning methods and two advanced deep learning models. The three machine learning methods are: SVM that directly uses the original signal as input, and KNN based on time-domain features and frequency-domain features respectively. The two deep learning models are CNN-LSTM and GaitRec-Net. To comprehensively evaluate the performance of the models, five evaluation metrics were used in the present invention: ACC (accuracy), Precision (precision), Recall (recall), F1-score, and AUC.
[0187] 1. Comparison results of machine learning
[0188] In the binary classification task between Parkinson's disease patients and healthy individuals (see Table 1), T-GaitNet demonstrated significant performance advantages, with all indicators far higher than those of traditional machine learning methods. Specifically, the SVM based on the original signal only achieved an ACC of 0.641, while the ACC of T-GaitNet was as high as 0.958. In addition, T-GaitNet was also significantly superior to the baseline model in the three indicators of Precision, Recall, and F1-score, reaching 0.964, 0.976, and 0.970 respectively. It is worth noting that the AUC of T-GaitNet was as high as 0.984, indicating its extremely high discrimination ability in classifying Parkinson's disease patients and healthy individuals. In contrast, even KNN-Time (ACC = 0.758), which performed relatively well, still could not be compared with T-GaitNet in terms of overall performance.
[0189] In the four-classification task of Parkinson's disease severity (see Table 2), T-GaitNet still demonstrated excellent performance, with all five indicators reaching leading levels. For example, the ACC of T-GaitNet was 0.933, nearly 30 percentage points higher than that of the best-performing KNN-Time (ACC = 0.627). At the same time, T-GaitNet achieved 0.933 and 0.974 in F1-score and AUC respectively, significantly superior to KNN based on frequency-domain features (F1-score = 0.489, AUC = 0.674). This indicates that T-GaitNet can stably capture the characteristics of gait patterns at different levels in the sub-classification of Parkinson's disease severity, reflecting its strong modeling ability for complex classification tasks.
[0190] For the classification task of musculoskeletal injury gaits (hip, knee, ankle, and calcaneus) (see Table 3), the performance of T-GaitNet was similar to that of the KNN method with time-domain and frequency-domain features, but it was slightly inferior in some indicators. Specifically, the ACC, Precision, Recall, and F1-score of T-GaitNet were all 0.923, while the corresponding indicators of KNN-Time were 0.930 and those of KNN-Freq were 0.928. It should be noted that although T-GaitNet was slightly behind KNN, its AUC value was close to that of KNN-Time, being 0.989 and 0.991 respectively, indicating that T-GaitNet still maintained a high classification accuracy and had strong characteristic discrimination ability.
[0191] When performing a five-classification task (normal, hip, knee, ankle, and calcaneus) on gait data of healthy individuals and individuals with musculoskeletal injuries (see Table 4), T-GaitNet once again demonstrated significant performance advantages. The ACC of T-GaitNet was 0.933, higher than that of KNN-Freq (ACC = 0.919) and KNN-Time (ACC = 0.915), and achieved an overall lead in key indicators such as Precision, Recall, F1-score, and AUC, reaching 0.938, 0.936, 0.933, and 0.974 respectively. In contrast, SVM based on the original signal performed poorly in the five-classification task, with an ACC of only 0.322 and low Precision and F1-score, further indicating the limitations of traditional methods in complex classification tasks. In comparison, T-GaitNet can better capture the subtle differences existing in gait signals and maintain high robustness under the increased complexity of the classification structure.
[0192] Table 1 Comparison results of T-GaitNet and traditional machine learning methods in the Parkinson's disease recognition task
[0193] Method ACC Precision Recall F1-score AUC SVM Based on Original Signal 0.641 0.731 0.771 0.749 0.621 KNN Based on Time-Domain Features 0.758 0.830 0.758 0.769 0.890 KNN Based on Frequency-Domain Features 0.670 0.770 0.752 0.759 0.614 T-GaitNet 0.958 0.964 0.976 0.970 0.984
[0194] Table 2 Comparison results of T-GaitNet and traditional machine learning methods in the Parkinson's disease severity grading task
[0195] Method ACC Precision Recall F1-score AUC SVM Based on Original Signal 0.476 0.772 0.627 0.635 0.633 KNN Based on Time-Domain Features 0.627 0.830 0.758 0.769 0.890 KNN Based on Frequency-Domain Features 0.552 0.501 0.507 0.489 0.674 T-GaitNet 0.933 0.938 0.936 0.933 0.974
[0196] Table 3 Comparison results of T-GaitNet and traditional machine learning methods in the musculoskeletal injury gait classification task (four-classification)
[0197] Method ACC Precision Recall F1-score AUC SVM Based on Original Signal 0.436 0.470 0.419 0.398 0.716 KNN Based on Time-Domain Features 0.930 0.930 0.930 0.930 0.991 KNN Based on Frequency-Domain Features 0.928 0.928 0.927 0.928 0.952 T-GaitNet 0.923 0.924 0.922 0.923 0.989
[0198] Table 4 Comparison results of T-GaitNet and traditional machine learning methods in the musculoskeletal injury gait classification task (five-classification)
[0199] Method ACC Precision Recall F1-score AUC SVM Based on Original Signal 0.322 0.411 0.392 0.325 0.713 KNN Based on Time-Domain Features 0.915 0.915 0.915 0.915 0.990 KNN Based on Frequency-Domain Features 0.919 0.918 0.916 0.917 0.948 T-GaitNet 0.933 0.938 0.936 0.933 0.974
[0200] 2. Deep learning comparison results
[0201] In the binary classification task between Parkinson's disease patients and healthy individuals (see Table 5), T-GaitNet achieved a significant performance improvement, with an ACC as high as 0.958, Precision, Recall, and F1-score being 0.964, 0.976, and 0.970 respectively, significantly outperforming deep learning methods such as CNN-LSTM and GaitRec-Net. In contrast, the overall performance of CNN-LSTM was poor, with an ACC of only 0.700 and an AUC of 0.654. Although GaitRec-Net was significantly better than CNN-LSTM in all metrics (e.g., ACC = 0.856, AUC = 0.925), there was still a large gap compared to T-GaitNet. T-GaitNet demonstrated higher accuracy and more stable classification performance in this task with its fine-grained modeling ability for temporal features.
[0202] In the four-class classification task of Parkinson's disease severity (see Table 6), T-GaitNet achieved the optimal classification performance, with an ACC reaching 0.933, Precision, Recall, and F1-score all exceeding 0.93, and the AUC reaching 0.974. In contrast, the ACC of GaitRec-Net was 0.758. Although its AUC (0.907) was relatively high, neither Precision nor F1-score exceeded 0.74. The overall performance of CNN-LSTM was not satisfactory, with an ACC of only 0.401, and Precision, Recall, and F1-score all at relatively low levels, indicating obvious limitations in the feature extraction ability and modeling effect of CNN-LSTM in fine-grained complex multi-classification tasks. The excellent performance of T-GaitNet in this task not only reflects its strong temporal feature extraction ability but also proves its high sensitivity to the changes in the severity of Parkinson's disease symptoms.
[0203] In the four-class classification task related to hip, knee, ankle, and calcaneus injuries (see Table 7), T-GaitNet also performed excellently, with an ACC of 0.923, Precision, Recall, and F1-score all being 0.923, and the AUC reaching 0.989. In contrast, the performance of CNN-LSTM was slightly lower, with an ACC of 0.879 and an F1-score of 0.881, while the ACC of GaitRec-Net was 0.857, the F1-score was 0.859, and the AUC was 0.973. This shows that T-GaitNet can better capture the complex features of gait data, especially performing more superiorly in the multi-class subdivision of gait disorders. At the same time, its AUC is close to 1.0, indicating that the model has extremely high robustness and reliability in distinguishing different gait injury categories.
[0204] In the five-classification task (normal gait, hip joint, knee joint, ankle joint, and calcaneus injury), the classification performance of T-GaitNet is still superior to other deep learning methods (see Table 8). The ACC of T-GaitNet is 0.933, both Precision and Recall are 0.938, and the F1-score is also 0.933. Its comprehensive classification performance is much higher than that of CNN-LSTM (ACC = 0.853) and GaitRec-Net (ACC = 0.868). Although GaitRec-Net performs slightly better in terms of AUC (AUC = 0.982, while T-GaitNet is 0.974), T-GaitNet significantly leads in other metrics, demonstrating its stable adaptability to multi-class tasks. It should be noted that in the case of increased task complexity, the performance of CNN-LSTM further deteriorates, indicating its limited ability to model the five-classification task, while T-GaitNet can effectively overcome this limitation by virtue of its optimized structural design.
[0205] Table 5 Comparison results of T-GaitNet and deep learning methods in the Parkinson's disease recognition task
[0206]
[0207]
[0208] Table 6 Comparison results of T-GaitNet and deep learning methods in the Parkinson's disease severity grading task
[0209] Method ACC Precision Recall F1-score AUC CNN-LSTM 0.401 0.248 0.273 0.246 0.526 GaitRec-Net 0.758 0.734 0.737 0.706 0.907 T-GaitNet 0.933 0.938 0.936 0.933 0.974
[0210] Table 7 Comparison results of T-GaitNet and deep learning methods in the musculoskeletal injury gait classification task (four-classification)
[0211] Method ACC Precision Recall F1-score AUC CNN-LSTM 0.879 0.884 0.879 0.881 0.979 GaitRec-Net 0.857334 0.856227 0.862752 0.858646 0.973279 T-GaitNet 0.923 0.924 0.922 0.923 0.989
[0212] Table 8 Comparison results of T-GaitNet and deep learning methods in the musculoskeletal injury gait classification task (five-classification)
[0213] Method ACC Precision Recall F1-score AUC CNN-LSTM 0.853 0.858 0.858 0.857 0.976 GaitRec-Net 0.868 0.865 0.876 0.869 0.982 T-GaitNet 0.933 0.938 0.936 0.933 0.974
[0214] Those skilled in the art know that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.
[0215] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A gait analysis and anomaly detection method, characterized in that, It includes the following steps: Step S1, preprocess the original GRF signal, and divide the preprocessed GRF signal into a sequence of time slices with a fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle; Step S2, input the time slice sequence into a multi-scale deep residual scaling network to extract multi-scale spatio-temporal features; Step S3, concatenate the multi-scale spatio-temporal features in chronological order to form a temporal feature sequence, and input the temporal feature sequence into a Transformer encoder. Use the multi-head self-attention mechanism to model the global temporal dependence relationship, and at the same time adopt a masking mechanism to mask invalid time steps, and output the global time series features; Step S4, perform classification processing on the global time series features, and output the gait classification probability distribution, which is used for Parkinson's disease recognition, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
2. The gait analysis and anomaly detection method according to claim 1, wherein The step S1 includes the following sub-steps: Step S1.1, obtain the original GRF signal X = {X1, X2, …, X N}, where N represents the number of samples, and X i represents the GRF signal of the i-th sample. Each GRF signal contains C channels, denoted as represents the GRF signal of the c-th channel. The signal of each channel where L i represents the length of the original GRF signal of the i-th sample; Step S1.2, perform padding operations on all original GRF signals to normalize the length to a fixed length L to obtain the preprocessed GRF signal; Step S1.3, according to the sampling rate and the number of data points, divide the preprocessed GRF signal into T time slices, expressed as where the number of slices T = L / L s , L s is the predetermined length of the time slice, representing the t-th GRF time slice signal of the c-th channel of the i-th sample.
3. A gait analysis and anomaly detection method according to claim 1, characterized in that, The step S2 includes the following sub-steps: Step S2.1, input the time slice sequence into a multi-scale deep residual scaling network, and extract local detail features through the DRSN-CW branch with a 5×5 convolutional kernel of the multi-scale deep residual scaling network, and extract the overall dynamic characteristics through the DRSN-CW branch with a 17×17 convolutional kernel of the multi-scale deep residual scaling network; Step S2.2, concatenate the features output by the DRSN-CW branch with a 5×5 convolutional kernel and the DRSN-CW branch with a 17×17 convolutional kernel respectively to obtain multi-scale spatio-temporal features.
4. A gait analysis and anomaly detection method according to claim 3, characterized in that Each DRSN-CW branch in the step S2.1 includes stacked RSBU-CW modules, and the RSBU-CW module includes the following operations: Step S2.1.1, perform global average pooling operation on the input time slice s, compress it into a one-dimensional vector, and input the one-dimensional vector into a two-layer fully connected network to calculate the scaling parameter α c : where z c is the output eigenvalue corresponding to the c-th channel of the fully connected network; Step S2.1.2, calculate the channel threshold τ according to formula (2) c :[[-END]] τ c = α c · average i,j |x i,j,c |(2) where average i,j represents the absolute average value of the feature map on channel c, and i and j are the indices of the width and height respectively, and x i,j,c is the feature value of the feature map at the width position i and height position j on the c-th channel; Step S2.1.3, use the channel threshold to perform a soft threshold shrinkage operation on each position of the input s, which is defined as follows: where x represents the feature map and y represents the shrunk output feature; Step S2.1.4, inside each DRSN-CW, the RSBU-CW modules are connected in series through a residual connection strategy. For a certain residual unit, the feature transfer form is: H(s) = F(s, {W i}) + s(4) where H(s) represents the output of the residual unit, F(s, {W i}) represents the feature transformation through the RSBU-CW module, s represents the input feature map of the residual unit, and W i represents all learnable parameters in the RSBU-CW module.
5. The gait analysis and anomaly detection method according to claim 3, characterized in that, In step S2.2, the formula for concatenating the multi-scale spatio-temporal features of the features F small (x) and F large (x) output by the 5×5 convolution kernel DRSN-CW branch and the 17×17 convolution kernel DRSN-CW branch respectively is as follows: F multi (s) = concat(F small (s), F large (s))(5).
6. The gait analysis and anomaly detection method according to claim 1, characterized in that, The step S3 includes the following sub-steps: Step S3.1, concatenate the multi-scale spatio-temporal features in chronological order to form a temporal feature sequence; Step S3.2, input the temporal feature sequence into a Transformer encoder, Step S3.3, calculate the attention weights of each time step through the multi-head self-attention layer of the Transformer encoder; Step S3.4, adopt a masking operation based on the padding position to output the global time series features.
7. The gait analysis and anomaly detection method according to claim 6, wherein, In the step S3.3, the input of the self-attention layer includes the query Q, the key K, and the value V, and the self-attention calculation method is: where d k is a scaling factor depending on the layer size, and Q, K, and V are obtained from the input sequence features through linear transformations respectively.
8. A gait analysis and anomaly detection method according to claim 1, characterized in that In the step S4, map the global time series features through multiple fully connected layers to output the gait classification probability distribution, and the gait probability distribution is used for Parkinson's disease recognition, Parkinson's severity assessment, and gait classification of musculoskeletal injuries.
9. A gait analysis and abnormality detection system, characterized in that It includes: The GRF time segmentation module preprocesses the original GRF signal and divides the preprocessed GRF signal into a sequence of time slices of a fixed length according to the sampling rate and the number of data points, where each time slice corresponds to a specific phase of the gait cycle; The multi-scale deep residual shrinkage network module inputs the time slice sequence into the multi-scale deep residual scaling network to extract multi-scale spatio-temporal features; The Transformer temporal encoder module concatenates the multi-scale spatio-temporal features into a temporal feature sequence in chronological order, inputs the temporal feature sequence into the Transformer encoder, uses the multi-head self-attention mechanism to model the global temporal dependence, and at the same time adopts the masking mechanism to mask the invalid time steps, and outputs the global time series features; The downstream task classification module classifies the global time series features and outputs the gait classification probability distribution for Parkinson's disease recognition, Parkinson's severity assessment, and musculoskeletal injury gait classification.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the gait analysis and anomaly detection method according to any one of claims 1 to 8.
Citation Information
Cited By
Personnel abnormal data early warning method and system
CN120612736A