Gait recognition method based on PFTgait model

By adopting the PFTgait model in gait recognition, combining the parallel multi-layer convolution feature extractor, frequency domain feature fusion module and timing-sequence correlation feature encoder, the problem of manual feature engineering dependence and global dependency modeling in the existing technology is solved, and more efficient gait feature extraction and classification accuracy is achieved.

CN120148113APending Publication Date: 2025-06-13CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510223342.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing gait recognition methods rely on manual feature engineering, making it difficult to fully mine the complex patterns of gait data, and frequency domain analysis is susceptible to noise interference. It is difficult for traditional RNN structures to effectively model global dependencies when processing long-term gait sequences.

Method used

The gait recognition method based on the PFTgait model is adopted, and the gait features are extracted and fused to capture the global dependency relationship through parallel multi-layer convolution feature extractor, frequency domain feature fusion module and timing correlation feature encoder.

Benefits of technology

It improves the adaptability and characterization ability of the model, reduces noise interference, enhances the capture of gait timing dependence, and improves the accuracy and robustness of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148113A_ABST
    Figure CN120148113A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of gait recognition, in particular to a gait recognition method based on a PFTgait model, and the method comprises the steps: collecting plantar pressure data; the multi-layer convolution feature extractor, the frequency domain feature fusion module and the time sequence correlation feature encoder are cascaded; gait feature extraction is carried out by using a multi-layer convolution feature extractor; performing frequency domain feature extraction and frequency-time feature conversion on the gait features by using a frequency domain feature fusion module; capturing a global dependency relationship of gaits by using a time sequence correlation feature encoder; and outputting a gait category by using the classifier. The method solves the problems that existing manual feature engineering depends on prior knowledge, a complex mode of gait data is difficult to comprehensively mine, and the representation capacity of a model is limited; an existing frequency domain analysis method is prone to noise interference, and the stability and generalization ability of features are affected. A traditional RNN structure is difficult to effectively model a global dependency relationship when processing a long-time gait sequence, resulting in the problems of attenuation or loss of key information and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gait recognition, and particularly to a gait recognition method based on the PFTgait model. Background Art

[0002] Parkinson's disease is a chronic progressive neurological disease that mainly affects the motor control function of patients.

[0003] Y. Liu et al. used a human pose estimator to extract the joint coordinates of the human body from a video dataset, then obtained relevant features through feature engineering and used relevant algorithms for prediction; A. Sabo et al. proposed a spatio-temporal graph convolutional network (ST-GCN); such methods can capture complete gait features; however, the acquisition of video data is affected by factors such as environmental illumination, occlusion, and camera perspective; in addition, privacy protection and data storage costs are also important factors restricting their practical applications.

[0004] Deep learning methods have been widely used in the field of gait analysis; Hu et al. used a two-stream three-dimensional convolutional neural network for gait pattern recognition; this method makes full use of the spatial information of gait data and fuses different features through a two-stream structure; however, the three-dimensional convolution has a high computational complexity, this method requires a large amount of computing resources, and the training process is time-consuming; in addition, directly mapping gait data into a class image form may lead to the loss of time series features, making the model have certain limitations in capturing gait temporal dependencies.

[0005] The method of ElMaachi et al. does not fully consider the temporal dependencies of gait data, while GaitNet further combines a recurrent neural network (RNN) to enhance the temporal modeling ability; however, due to the limitations of the RNN structure in capturing the long-range dependencies of long gait sequences, this method may lead to the loss of information of long-term gait patterns, affecting the integrity of feature expression. Summary of the Invention

[0006] 1. Aiming at the deficiencies of existing methods, the present invention solves the problems that existing manual feature engineering depends on prior knowledge, is difficult to comprehensively mine the complex patterns of gait data, and limits the representation ability of the model; existing frequency domain analysis methods are vulnerable to noise interference, affecting the stability and generalization ability of features; traditional RNN structures are difficult to effectively model global dependencies when processing long gait sequences, resulting in the attenuation or loss of key information.

[0007] The technical solution adopted by the present invention is that the gait recognition method based on the PFTgait model includes the following steps:

[0008] Step 1: Collect plantar pressure data and preprocess the pressure data;

[0009] As a preferred embodiment of the present invention, the preprocessing includes: missing value elimination, data windowing, and setting an overlap rate.

[0010] Step 2: Construct a PFTgait model, including: cascading a plurality of multi-layer convolutional feature extractors, a frequency domain feature fusion module, and a temporal correlation feature encoder; using the multi-layer convolutional feature extractors to extract gait features; using the frequency domain feature fusion module to perform frequency domain feature extraction and frequency-time feature conversion on the gait features; using the temporal correlation feature encoder to capture the global dependencies of the gait;

[0011] As a preferred embodiment of the present invention, the multi-layer convolutional feature extractor includes: a first convolutional layer, a first activation function, a second convolutional layer, a second activation function, a max pooling layer, a third convolutional layer, and a third activation function; wherein, the first to third convolutional layers are 1D convolutions, the convolutional kernel is 3, and the number of output features is 16;

[0012] As a preferred embodiment of the present invention, the number of multi-layer convolutional feature extractors is 18.

[0013] As a preferred embodiment of the present invention, the frequency domain feature fusion module includes:

[0014] First, divide the feature map into several subgroups [v 0 , v 1 , v 2 , ···, v n-1 along the channel dimension, where v i = R 1 ×L and n = N v ;

[0015] Second, output the frequency vectors [Freq 0 , Freq 1 , ···, Freq n-1 for each subgroup in ascending order of frequency, and stack all the frequency vectors through stack operations;

[0016] Third, input the stacked frequency vectors into a fully connected layer;

[0017] Finally, perform an inverse transformation on the frequency domain features output by the fully connected layer.

[0018] As a preferred embodiment of the present invention, the temporal correlation feature encoder includes:

[0019] The time step data X t is L x × d model , where L x represents the sequence length, and d modelis the dimension of the model;

[0020] X t After passing through the multi-head attention mechanism, it is processed through residual connection and layer normalization and then fed into the feed-forward neural network;

[0021] Then, data compression is performed through the combination of convolution and pooling.

[0022] As a preferred embodiment of the present invention, the number of temporal correlation feature encoders is 3.

[0023] Step 3: Use the classifier to output the gait category;

[0024] As a preferred embodiment of the present invention, the classifier is a fully connected layer of 1 neuron and 5 neurons.

[0025] As a preferred embodiment of the present invention, the gait recognition system based on the PFTgait model includes: a memory for storing instructions executable by a processor; a processor for executing the instructions to implement the gait recognition method based on the PFTgait model.

[0026] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code, the computer program code implementing the gait recognition method based on the PFTgait model when executed by a processor.

[0027] Advantages of the present invention:

[0028] 1. Through the design of the parallel multi-layer convolution module, the multi-scale convolution structure enhances the learning ability of the local patterns of gait data, reduces the dependence on artificial feature engineering, and improves the self-adaptability of the model;

[0029] 2. Combining the frequency domain information enhancement mechanism, using the discrete cosine transform to extract the frequency domain features of the gait signal, and through adaptive weight adjustment, dynamically focusing on the key frequency components, effectively suppressing the interference of noise on feature extraction;

[0030] 3. The introduction of the Transformer structure breaks through the bottleneck of the traditional RNN model in modeling long-range dependencies, enabling the model to capture the global dependencies in the gait signal, and improving the accuracy and robustness of classification. Description of the Drawings

[0031] Figure 1 is the flowchart of the gait recognition method based on the PFTgait model of the present invention;

[0032] Figure 2 is the system connection diagram of the measuring device of the present invention;

[0033] Figure 3It is a schematic structural diagram of the measuring device of the present invention;

[0034] Figure 4 It is a structural diagram of the channel in the measurement section of the present invention;

[0035] Figure 5 It is a structural diagram of the fixed steel needles in the channel of the measurement section of the present invention;

[0036] Figure 6 It is the VGRF spectrogram of Parkinson's patients with different severities;

[0037] Figure 7 It is the data distribution diagram of the subjects;

[0038] Figure 8 It is the result of the Mann-Whitney U test;

[0039] Figure 9 It is the Parkinson classification confusion matrix. Detailed implementation manners

[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.

[0041] As Figure 1 、 2 shown, a gait recognition method based on the PFTgait model includes the following steps:

[0042] Step 1: Collect plantar pressure data and preprocess the pressure data;

[0043] Adopt the pressure data of 16 sensors (8 points on each of the left and right) and the sum of the pressure values of the sensors on the left and right feet;

[0044] First, input the one-dimensional vertical ground reaction force (VGRF) signal provided by the foot sensors when the subject is walking, that is, the pressure data, denoted as s i (t), where t represents time and i is the index of the sensor; the VGRF signal represents the vertical ground reaction force measured by the foot sensors (unit: Newton) and is recorded as a function of time.

[0045] In the data preprocessing stage, first process the missing values in the data; for the records with missing values, choose to remove them to ensure the integrity of the data; then, perform windowing on the data; considering that the subtle changes in gait usually occur in a very short period of time, set the window size to 100 data points and set an overlap rate of 50% to ensure that the existing data can be utilized to the maximum extent; divide the signal data into n different segments, and the segments are represented as:

[0046]

[0047] Among them, i is the index of the sensor, (i ∈ [1, 16]); j is the segment index, and n is the number of elements.

[0048] Step 2: Construct the PFTgait (Parallel Frequency-Temporal Gait Model) model, which is mainly divided into three parts: a parallel multi-layer convolutional feature extractor, a frequency-domain feature fusion module, and a temporal correlation feature encoder;

[0049] The preprocessed data is subjected to feature extraction by a feature extractor with a multi-layer convolutional structure; subsequently, the feature data enters the frequency-domain feature fusion module to further enrich the feature representation; finally, the data is passed into the temporal correlation feature encoder, which not only performs temporal modeling but also completes the final classification and prediction tasks through a classifier;

[0050] Parallel multi-layer convolutional feature extractor (Parallel Multi-Layer Convolutional Feature Extractor); The starting part of the network is a parallel multi-layer convolutional feature extractor, located as Figure 2 shown in the upper left part, and its internal detailed structure is shown in Figure 3 ; Since the present invention uses multi-feature data, a feature extractor with a parallel structure is used to extract relevant important features, aiming to analyze the patterns that change over time in time-series data, especially the gait patterns captured by each foot sensor;

[0051] The parallel multi-layer convolutional feature extractor is composed of several multi-layer convolutional feature extractors. The multi-layer convolutional feature extractor includes: a first convolutional layer, a first activation function, a second convolutional layer, a second activation function, a max pooling layer, a third convolutional layer, and a third activation function; among them, the first to third convolutional layers are 1D convolutions, the convolutional kernel is 3, and the number of output features is 16; the first to third activation functions are selu functions.

[0052] The output features of several multi-layer convolutional feature extractors are input into the fusion features through a concat operation;

[0053] Adopt a method similar to depthwise convolution to perform separate convolutions on each channel of the input data, that is, the features, so as to obtain an output feature map with the same number of channels as the input feature map; however, this may lead to too few output feature maps, thus affecting the effectiveness of information. Therefore, the idea of pointwise convolution is added to capture the correlation between different channels while maintaining spatial features and enhance the feature expression ability; in this way, not only can the changes in gait patterns and the diversity of features be obtained, but also the complex information in gait data can be understood and utilized more effectively; the parallel multi-layer convolution feature extractor consists of 18 parallel convolution blocks, which are used to process the 8 independent sensor data of each foot, as well as the additional features composed of the sum of these 8 sensors; compared with the traditional convolution structure, the structure of the parallel multi-layer convolution feature extractor can reduce the number of input channels due to the parallel processing method, thus effectively reducing the parameters required for the convolution layer and having a faster running speed at the same time.

[0054] Define k as the size of the convolution kernel, C in as the number of channels of the input feature map, C out as the number of channels of the output feature map, H is the height of the input feature map, W is the width of the input feature map, and the parameter ratio and calculation amount ratio formulas are as follows;

[0055]

[0056]

[0057] Input the fused features output by the multi-layer convolution feature extractor into the frequency domain feature fusion module and perform residual connection with the fused features;

[0058] Frequency Domain Feature Fusion Module. Since applying the relevant feature extraction method only in the time dimension may lead to insufficient information extracted from the time series and even cause information loss; to alleviate this problem, more frequency domain information needs to be introduced;

[0059] Although the Fourier transform and its inverse transform can obtain the frequency domain representation and reconstruct the time domain signal, redundant information is usually easily introduced during the Fourier transform process; for this reason, the present invention uses the discrete cosine transform (DCT) to replace the Fourier transform to reduce this potential impact; the discrete cosine transform is a mathematical transformation method used to transform a real number sequence from the time domain to the frequency domain; compared with the Fourier transform in the complex domain, the DCT is easier to understand and process; it can not only effectively capture the frequency information of the signal, but also has lower complexity and higher execution efficiency in terms of calculation;

[0060] The DCT transforms a real-valued sequence x[n] of length N into another real-valued sequence X[k] of length N, expressed as:

[0061]

[0062] where X[k] is the k-th coefficient in the frequency domain, representing the amplitude of the original signal x[n] at frequency k; α(k) is the normalization coefficient; and x[n] is the n-th sample value in the time domain.

[0063] The frequency-domain feature fusion module effectively captures the subtle features hidden in the frequency domain of the time series.

[0064] The input feature map is split into n subgroups along the channel dimension, denoted as [v 0 , v 1 , v 2 , ···, v n-1 , where each v i = R 1×L , and n = N v , N v represents the number of partitions of the sequence of length N along the channel dimension; then, for each subgroup, its corresponding DCT frequency components are processed in the order from low frequency to high frequency to obtain the set of processed frequency vectors [Freq 0 , Freq 1 , ···, Freqn -1 ; finally, all the processed frequency components are stacked through a stack operation to form a complete frequency vector:

[0065] Freq = stack([Freq 0 , Freq 1 , …, Freqn -1 )

[0066] When processing the complete frequency vector, traditional filters can usually only process local frequency components, and these filters share a fixed set of parameters, which may limit the comprehensive capture of frequency-domain information; to overcome these limitations, learnable filters are introduced, which essentially adopt the design of a fully connected layer (FC layer); different from traditional fixed-parameter filters, this module can dynamically adjust its parameters according to the feedback during the training process, thereby optimizing the extraction of frequency-domain information; through this method, the module can extract the features most valuable for the model task, improving the flexibility and effectiveness of feature extraction.

[0067] The obtained frequency-domain features are inversely transformed and fully fused with the original time-domain data. This not only fully exploits the frequency-domain information, improves the diversity and representation accuracy of the data, but also enhances the learning ability of the model, thus showing better performance when dealing with complex time-series data.

[0068] Temporal Correlation Feature Encoder. To reduce the within-class variance and effectively capture the global dependencies, the temporal correlation feature encoder of the present invention is as Figure 4 shown; To effectively capture the global dependencies, fixed-position encoding based on appropriate segment lengths and fixed step sizes is introduced; Fixed-position encoding can not only effectively model the global dependencies in these data, but also enhance the performance of the model when dealing with time-series data;

[0069] First, the time-series data is input into the temporal correlation feature encoder of the model; The input data X at each time step t is organized in matrix form, with dimensions of L x ×d model , where, L x represents the sequence length, and d model is the dimension of the model;

[0070] After the input data passes through the multi-head attention mechanism, the relationship information between the elements in the input sequence is obtained; Then, it is processed through residual connection and layer normalization, which ensures the stability of the model and the efficiency of training; Subsequently, the data is fed into the feed-forward neural network for further in-depth processing of the extracted features;

[0071] To speed up the calculation, the data processed each time is compressed, mainly by combining convolution and pooling, to compress the data length to half of the original; By stacking several temporal correlation feature encoders, a pyramid-like structure is achieved; The number of temporal correlation feature encoders is 3;

[0072] This method can effectively retain and integrate the key information in the data, making the final output contain rich features, and at the same time enhancing the expressive ability of the model, enabling it to better capture the complex patterns in the data.

[0073] To diagnose Parkinson's disease and predict its severity, two different fully connected layers are used as the classifier of the model, containing 1 neuron and 5 neurons respectively; These layers receive the output of the upper temporal correlation feature encoder and generate the relevant probability distributions for finally determining the classification category of Parkinson's disease; This design enables the model to effectively process the input data and output the prediction results related to the severity of Parkinson's disease.

[0074] Experimental process:

[0075] Use the PhysioNet gait dataset; it contains data of 93 Parkinson's disease (PD) patients and 73 healthy subjects; each participant walked independently on flat ground for two minutes without assistive devices; during data collection, 8 pressure sensors were installed on each foot to measure the plantar pressure distribution in real time; the relative positions of these sensors when the subject stands with feet parallel are as Figure 5 shown, and the sensor positions on the left foot are symmetric with respect to the Y-axis; the data of all pressure sensors were collected at a sampling frequency of 100 Hz; in addition, the recorded data also includes the total signals from all 8 sensors on the left and right feet to reflect the overall pressure changes on the left and right plantar surfaces; in addition, the Unified Parkinson's Disease Rating Scale (UPDRS), a commonly used scale for evaluating the severity of Parkinson's disease, was used; the continuous UPDRS scores were divided into five grades, class 1: UPDRS < 5, class 2: 5 ≤ UPDRS < 15, class 3: 15 ≤ UPDRS < 25, class 4: 25 ≤ UPDRS < 35, class 5: 35 ≤ UPDRS; the number of subjects in each category is shown in Table 1.

[0076] Table 1 Number of people with different severities

[0077]

[0078] Comparison of VGRF readings differences, Figure 6 showing the differences in the vertical ground reaction force (VGRF) readings of the foot positions of the control group and Parkinson's disease (PD) subjects; for convenient visualization, the sensor data at the heel position of the right foot was selected and the raw data was transformed from the time domain to the frequency domain through the fast Fourier transform (FFT); since the frequency domain signal is symmetric, only the positive frequency part was plotted; from Figure 6 it can be seen that there are obvious differences between the spectrograms of PD patients with different severities; as the disease severity increases, the deviation shown in the spectrogram becomes more obvious; indicating that the physiological characteristics of Parkinson's disease patients in gait pattern and gait stability change with different disease severities; this further proves the necessity of gait analysis in the diagnosis and research of Parkinson's disease.

[0079] It is well known that different individuals will show various differences when performing the same action, and these differences are affected by many factors such as physiology, psychology, experience and environment. Taking push-ups as an example, people with stronger muscle strength may find it easier to maintain a standard posture, while those with weaker strength may need to adjust their posture to complete the action. Similar situations are also observed in gait datasets. Early Parkinson's patients mainly show a slightly slower pace and a reduced stride. In the middle and late stages, they show an unstable gait, a further reduction in stride, a dragging pace, and even easy loss of balance when walking. Even at the same severity, patients with taller stature or heavier weight face greater challenges in maintaining balance and coordination. Therefore, before using the dataset, it is crucial to conduct a proper analysis of its existing differences.

[0080] Analyzing these differences in data sets can provide a deep understanding of the diversity and complexity of the data, which is crucial for choosing an appropriate data model architecture. If there are significant and complex differences in the data set, a deep neural network model that can handle this complexity may be needed. These models can better capture changes and associations within the data, thereby improving the model's ability to represent data features. Therefore, when designing and selecting models, understanding the differences between data sets can guide finding a balance between model complexity and performance to achieve better prediction and analysis results.

[0081] First, samples were selected from different individuals with the same severity, and the data distribution characteristics were analyzed in detail. The data distribution of different subjects was visualized using the T-SNE tool. T-SNE is a nonlinear dimensionality reduction technique. In the T-SNE graph, each point represents a sample in the original dataset, and the coordinates of the point represent its position in the low-dimensional space. This method preserves the local structure of the data by mapping high-dimensional data to low-dimensional space while keeping the distance between classes as large as possible. Figure 7 In a, the blue and yellow points represent different subject data, and the distribution of these points in two-dimensional space shows the similarities and differences between them; similar points will cluster together, while points with large differences will be far apart; it can be clearly seen from the figure that for the same subject, their data are almost all clustered together, even if there are some overlaps, but they are very rare, which reflects that there are significant differences between the two; further analysis shows that in Figure 7 The same conclusion can be drawn from the figure by comparing more specific statistics in b; for example, the median, etc. These observations reveal the differences in gait data between individuals, which may be caused by individual physiological characteristics, exercise habits or other factors.

[0082] Then, the hypothesis testing method was used to conduct a specific difference analysis on the data set. Before conducting any hypothesis test, it is necessary to ensure that the data meets the basic assumptions of normal distribution and homogeneity of variance. The significance level is set to 0.05. If the P value of the hypothesis test is less than this significance level, the null hypothesis is rejected, indicating that the data does not meet the normal distribution or the variance is significantly different. By performing normality and homogeneity of variance tests on each sensor data, it is found that the data does not meet these assumptions.

[0083] Therefore, the present invention uses non-parametric test methods, such as Mann-Whitney U test, to compare the significant differences between the data of two different individuals. First, the data set to be tested for hypothesis is simply processed. For each row of data, the average value is taken as the basis. Due to the large number of data points, some sample points are selected by random sampling. In total, subjects 1 and 2 with the same severity and subject 3 with a different severity are selected. After setting the significance level to 0.05, the data are subjected to pairwise Mann-Whitney U test. The final results are as follows: Figure 8 As shown in the figure, it can be clearly seen that the P-Value between each two subjects is much smaller than the set significance level, so there are great differences between them. Based on this, it can not only help understand the differences in data between individuals, but also provide an important basis for selecting appropriate data modeling methods, ensuring that the model can accurately capture the complexity of the data set and the differences between individuals.

[0084] To evaluate the performance of the proposed method, several commonly used indicators were calculated and reported; for the binary classification problem (whether or not someone has Parkinson's disease), the results are expressed in terms of sensitivity, specificity, and accuracy; Sensitivity, also known as True Positive Rate, measures the proportion of positive examples correctly identified by the model among all actual positive examples, reflecting the ability of the model to detect actual positive examples; Specificity, also known as True Negative Rate, measures the proportion of negative examples correctly identified by the model among all actual negative examples, reflecting the accuracy of the model in excluding negative examples; and Accuracy indicates the proportion of samples correctly predicted by the model among all predicted samples, providing an overview of the overall performance of the model.

[0085] For multi-classification problems (distinguishing the severity of Parkinson's disease), they are represented by Precision, Recall, F1-Score, and Accuracy; Precision measures the proportion of true positives among all samples predicted as positive. Recall measures the proportion of positive examples correctly identified by the classifier among all actual positive examples. The F1-score is the harmonic mean of Precision and Recall, and its value ranges from 0 to 1, with 1 indicating the best classification effect. The following are the definitions of the performance metrics used, including TP (number of true positives), TN (number of true negatives), FP (number of false positives), and FN (number of false negatives).

[0086] Before conducting an in-depth analysis of the one-dimensional signals in the Physionet gait dataset, appropriate data preprocessing is first carried out; the sliding window technique plays a crucial role in this process; to adapt to the complexity of the severity classification task, the parameters selected for the sliding window include the window length and the overlap ratio, and the specific details are shown in Tables 2 and 3; as shown in Table 2, the window lengths of 50, 100, and 200 data points are compared, and it is found that when the window length is set to 100 data points, the model achieves the best performance; in addition, based on this optimal window length, different overlap ratios are evaluated, namely 30%, 50%, and 70%; the results show that when the overlap ratio is set to 50%, the model performs best.

[0087] The original gait signals are cut into windows with a length of 100 data points, generating a total of approximately 64,000 data windows, and ensuring a 50% overlap rate between adjacent windows; such an overlap setting can better capture the continuity and local features of the signals, reducing information loss due to window boundaries; after windowing, label matching is performed on each data segment to align it with the corresponding class label; ensure that each window segment contains sufficient information for training the model and is consistent with its true class label, thereby constructing the input dataset for the model; after a series of preprocessing steps, the data is successfully input into the model, providing rich input information for the model.

[0088] Table 2 Comparison of the number of data points in the sliding window

[0089]

[0090] Table 3 Comparison of different overlap rates of the sliding window

[0091]

[0092] During the model training process, the following parameter settings are adopted to optimize performance and stability: 150 samples are processed per batch, and 50 training epochs are carried out; The Nadam stochastic optimization algorithm is used for model optimization. Its key parameters include the learning rate set to 0.001, and the exponential decay rates of the first and second momentum estimates are b1 = 0.9 and b2 = 0.999 respectively; It aims to balance the stability of gradient updates and the learning speed, and improve the training efficiency; To further enhance the performance of the model and reduce the risk of overfitting, a dropout rate of 0.1 is introduced; Dropout can effectively randomly discard the outputs of some neurons, forcing the model to learn more robust features during training; In addition, an early stopping strategy is also implemented to monitor the progress of model training in real time and prevent overtraining. The Parkinson's diagnosis and prediction models are respectively subjected to 10 - fold random cross - validation and 10 - fold cross - validation; It can comprehensively and accurately evaluate the performance of the model on different tasks and ensure the effectiveness and reliability of the final results.

[0093] To evaluate the effectiveness of the deep neural network model of the present invention, its performance is compared with existing Parkinson's disease (PD) diagnosis methods; During the evaluation process, first, the performance of the model on all data segments is obtained, and then the results of the subjects are synthesized by means of majority voting, and finally the evaluation results of the model for each individual are obtained; Specifically, when performing majority voting on the results of the subjects, the diagnosis results of multiple data segments are combined to obtain a comprehensive evaluation of each individual; The results are shown in Table 4. According to the data in Table 4, the model of the present invention performs better than other methods in Parkinson's disease diagnosis; Specifically, the model of the present invention achieves an accuracy of 97.0% at the individual subject level, reflecting its excellent effect in identifying PD patients; In addition, the model of the present invention also shows satisfactory results in terms of sensitivity and specificity; High sensitivity ensures that the model can accurately identify the vast majority of PD patients, while high specificity reduces the possibility of misdiagnosis, thereby improving the reliability of the overall diagnosis; Such a comprehensive performance indicates that the model of the present invention has great potential in the early detection and diagnosis of Parkinson's disease and can effectively support decision - making in clinical practice.

[0094] Table 4 Comparison of Parkinson's diagnosis algorithms

[0095]

[0096] As can be seen from Table 4, although Khoury and Only the accuracy of the algorithms was reported in the studies, but these accuracies were significantly lower than those of the model of the present invention. This difference may be due to their reliance on manually extracted features, which failed to fully capture the subtle changes in the original data, resulting in a lack of diagnostic performance. Compared with the detection algorithm proposed by Nguyen et al., the model of the present invention has significantly improved in both classification accuracy and sensitivity. This improvement is mainly attributed to the in-depth mining of key features in the modeling process of the model of the present invention, especially the effective utilization of frequency-domain information. By combining time-domain and frequency-domain features, the model of the present invention can more comprehensively capture the potential patterns in gait signals, thereby improving the accuracy of diagnosis. Although the diagnostic model proposed by Zhao et al. achieved good results in terms of accuracy and sensitivity, its specificity was significantly lower than that of the model of the present invention. The lack of specificity may lead to a high misdiagnosis rate and reduce the reliability of the model in clinical applications. In contrast, the model of the present invention not only maintains high accuracy and sensitivity but also performs well in terms of specificity, further enhancing the comprehensiveness and stability of diagnosis. Dong's research results are similar to those of the present invention. Although it is slightly higher than the present invention in terms of sensitivity, it has decreased in the other two indicators. In addition, Dong's method mainly targets gait data, while the algorithm of the present invention can independently process one-dimensional input signals, making it more flexible and versatile. This advantage makes the model of the present invention easier to be extended to other experimental settings, especially in different clinical gait research scenarios, and can also maintain high adaptability and diagnostic performance.

[0097] The model of the present invention has achieved good results in the diagnosis of Parkinson's disease. In order to achieve personalized treatment for different Parkinson's disease patients and monitor their rehabilitation progress, the present invention has carefully evaluated the severity of patients based on the Unified Parkinson's Disease Rating Scale (UPDRS) score. For this task, the present invention retains the network structure of the diagnostic model and modifies the last layer on this basis, adding a fully connected layer containing five neurons and introducing a softmax activation layer to adapt to the multi-class severity classification requirements. In order to verify the effectiveness of the model of the present invention, it was compared with existing related algorithms, and the results are shown in Table 4. The results show that the model of the present invention performs well in multiple indicators such as accuracy, sensitivity, and specificity, further proving its potential and practicality in clinical applications. In particular, the model demonstrated superior performance in the classification task of UPDRS scores, which provides a reliable tool for the comprehensive management and treatment of Parkinson's disease patients.

[0098] Table 5 Comparison of Parkinson's severity prediction algorithms

[0099]

[0100] From the results in Table 5, it can be seen that the model of the present invention performs well in classifying different severity levels of Parkinson's disease, with an accuracy rate of 88.4%. This shows that the method of the present invention can effectively utilize the designed model structure to achieve accurate classification of Parkinson's disease; in contrast, Naimi et al. used a time and space encoder to jointly model the time and space features of the data in their study. Although multi-dimensional features were considered, the classification accuracy of their model was only 81.3%. This may be because in the feature extraction process, the model over-processed the complexity of the data. Although comprehensive modeling was performed in the time and space dimensions, the most effective features for the classification task were not fully extracted, thus affecting the accuracy of the final result; El Maachi et al. used a one-dimensional convolutional network structure and achieved an accuracy rate of 85.3% in the same classification task; but due to the shortcomings of convolutional neural networks in processing time series data, the characteristics of time series data cannot be fully utilized. This may result in the model failing to achieve the best classification effect when faced with complex Parkinson's disease data, thereby limiting the breadth and depth of its practical application; in addition, these research results also demonstrate the innovative contribution of the present invention in model design and feature extraction strategy, and provide new ideas and methods for the automated classification of Parkinson's disease; this not only broadens the application prospects of related research, but also provides more reliable technical support for clinical diagnosis and treatment.

[0101] Ablation experiment, the present invention conducts a detailed ablation study on the prediction of Parkinson's severity by comparing and analyzing the proposed model with the designs of several related models to identify and evaluate the role of each key component; the method of gradually eliminating or adjusting specific modules of the model is adopted to systematically explore the impact of each component on the overall performance of the model; this method not only helps to deeply understand the contribution of each part of the model, but also reveals the possible room for improvement in the optimization process; through this detailed analysis, the actual effect of each design choice can be more accurately evaluated, thereby providing a strong basis for further optimization and promotion of the model; Table 6 shows the main components used in each model in the ablation study.

[0102] Table 6 Summary of model ablation study

[0103]

[0104] In model A, only the time series correlation feature encoder is used for experiments. This setting helps to understand the contribution of time series features to model performance. In model B, other modules are removed and only the parallel multi-layer convolutional feature extractor is retained, which can evaluate the independent effect of convolutional feature extraction. Model C combines the time series correlation feature encoder and the parallel multi-layer convolutional feature extractor to explore the performance improvement of the combination of the two. Finally, based on model C, the frequency domain feature fusion module is added to form the final model of the present invention.

[0105] By comparing models composed of different modules, it is found that the final model performs best in terms of performance; specifically, when only using Model A, the obtained accuracy is the lowest, which indicates that using only the time series feature encoder alone cannot fully exploit the potential information in the data; when using Model B, although the performance of the model is improved, it is still lower than that of Model C; this is because Model C combines the advantages of the time series feature encoder and the parallel multi-layer convolutional feature extractor, which can not only capture the subtle features in the data but also fully model the time series relationship of the data; therefore, the performance of Model C is better than that of Model B; finally, a frequency domain feature fusion module is further added to Model C to explore the potential frequency domain features of the data; this additional module not only enhances the model's processing ability for time series features but also improves the capture of frequency domain features, thus enabling the final model of the present invention to achieve the best performance level; this result shows that by carefully designing and combining different model components, the comprehensive performance of the model can be significantly improved, laying a solid foundation for further optimization and practical applications.

[0106] Finally, a confusion matrix is drawn to further verify the effect achieved by the final model of the present invention; the confusion matrix provides the detailed prediction results of the classification model for each category, and such detailed classification results help to comprehensively evaluate the model performance. By analyzing the confusion matrix, it is possible to identify in which categories the model is prone to errors, understand the limitations and biases of the model, so as to provide further guidance for model optimization; from Figure 9 it can be observed that the model of the present invention performs excellently in Categories 0 and 1 with relatively high accuracies.

[0107] Taking the ideal embodiments of the present invention described above as inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. The gait recognition method based on the PFTgait model is characterized by: The following steps are involved: Step 1: Collect plantar pressure data and pre-process the pressure data; Step 2: Construct a PFTgait model, including: cascading several multi-layer convolution feature extractors, frequency domain feature fusion modules and time series correlation feature encoders; extracting gait features using multi-layer convolution feature extractors; extracting frequency domain features and converting frequency-time features of gait features using frequency domain feature fusion modules; and capturing the global dependency of gait using time series correlation feature encoders; Step 3: Use the classifier to output the gait category.

2. The gait recognition method based on the PFTgait model according to claim 1, characterized in that: The multi-layer convolution feature extractor includes: a first convolution layer, a first activation function, a second convolution layer, a second activation function, a maximum pooling layer, a third convolution layer and a third activation function; wherein the first to third convolution layers are 1-dimensional convolutions, and the first to third activation functions are selu functions.

3. The gait recognition method based on the PFTgait model according to claim 1, characterized in that: The frequency domain feature fusion module includes: First, the feature map is split into several subgroups along the channel dimension [v0,v1,v2,···,v n-1 ],v i =R 1×L And n = N v ; Secondly, each subgroup outputs a frequency vector [Freq 0 ,Freq 1 ,···,Freq n-1 ] and stack all frequency vectors through stack operation; Secondly, the stacked frequency vector is input into the fully connected layer; Finally, the frequency domain features output by the fully connected layer are inversely transformed.

4. The gait recognition method based on the PFTgait model according to claim 1, characterized in that: The temporal correlation feature encoder includes: Time step data X t For L x ×d model , L x represents the sequence length, d model is the dimension of the model; X t After the multi-head attention mechanism, it is processed through residual connection and layer normalization and then sent to the feedforward neural network; Data compression is then performed through convolution and pooling.

5. The gait recognition method based on the PFTgait model according to claim 1, characterized in that: The classifier is a fully connected layer with 1 neuron and 5 neurons.

6. The gait recognition method based on the PFTgait model according to claim 2, characterized in that: The number of multi-layer convolutional feature extractors is 18.

7. The gait recognition method based on the PFTgait model according to claim 1, characterized in that: Preprocessing includes: Eliminate missing values, perform window processing on data, and set overlap rate.

8. The gait recognition method based on the PFTgait model according to claim 4, characterized in that: The number of temporal correlation feature encoders is 3.

9. The gait recognition system based on the PFTgait model is characterized by: include: a memory for storing instructions executable by a processor; A processor, configured to execute instructions to implement the gait recognition method based on the PFTgait model as described in any one of claims 1 to 8.

10. A computer readable medium storing computer program code, characterized in that: When the computer program code is executed by a processor, the computer program code implements the gait recognition method based on the PFTgait model as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Optical fiber intelligent carpet gait recognition system based on strain-contour bimodal network

    CN121570168A

  • Fiber-optic smart carpet gait recognition system based on strain-profile dual-modal network

    CN121570168B