Abnormal gait recognition method based on multi-scale convolution and cross attention fusion

By adopting multi-scale convolution and cross-attention fusion methods in abnormal gait recognition, the problem of insufficient feature extraction and fusion in the prior art is solved, and the precise recognition accuracy and generalization ability of multiple abnormal gait patterns are achieved.

CN119989271APending Publication Date: 2025-05-13ZHONGYUAN ENGINEERING COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076059.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to fully capture high-dimensional dynamic features in gait data in abnormal gait recognition, and the multi-scale feature extraction and feature fusion mechanisms are insufficient, which affects the recognition accuracy and generalization ability.

Method used

The abnormal gait recognition method based on multi-scale convolution and cross-attention fusion is adopted. Time features are extracted through multi-scale convolution modules, and important features are dynamically paid attention to by combining channel and spatial attention mechanisms, and feature fusion is carried out through cross-attention modules to enhance feature representation capabilities.

Benefits of technology

It realizes accurate identification of multiple abnormal gait patterns, improves recognition accuracy and generalization ability, is highly adaptable, and can dynamically pay attention to the time changes of gait data and pressure changes in different areas of the foot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989271A_ABST
    Figure CN119989271A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal gait recognition method based on multi-scale convolution and cross attention fusion. The method comprises the steps that an abnormal gait data set is acquired; performing data preprocessing according to the abnormal gait data set and the disclosed Parkinson's disease gait data set to obtain preprocessed data; an abnormal gait recognition model is constructed, the abnormal gait recognition model is optimized according to the preprocessed data, and an optimized abnormal gait recognition model is obtained; gait data to be detected are obtained and input into the abnormal gait recognition model, and an abnormal gait recognition result is obtained. According to the invention, various abnormal gait modes can be accurately recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of abnormal gait recognition, and in particular relates to an abnormal gait recognition method based on multi-scale convolution and cross-attention fusion. Background Art

[0002] Gait is a basic human activity mode, and its specific gait pattern reflects the coordinated movement of the limbs during walking. However, diseases of the nervous system and musculoskeletal system, such as Parkinson's Disease (PD), often lead to abnormal gait patterns. This abnormal gait pattern manifests as changes in gait dynamics. For example, in the early stages of Parkinson's disease, patients may experience symptoms such as mild bradykinesia and reduced arm swing. As the disease progresses, conditions such as shortened stride length, forward posture, and gait freezing become more obvious. In the late stage of the disease, extremely short stride length, frequent gait freezing, and severe postural instability can significantly affect the patient's mobility and daily life functions. Therefore, accurate identification of abnormal gait patterns is of great significance for early diagnosis, treatment monitoring, and rehabilitation evaluation of the disease, which helps to improve the treatment effect and quality of life of patients.

[0003] Traditional gait analysis research mainly relies on data analysis and machine learning techniques. These methods usually extract basic features such as step length, gait speed and gait cycle, and apply algorithms such as support vector machines (SVM), decision trees (DT) and random forests (RF) for gait recognition. However, these traditional methods rely heavily on hand-designed features and cannot fully capture the rich information of high-dimensional dynamic features in gait data, thus limiting the further development of data potential. In recent years, deep learning technology has made significant progress in the field of abnormal gait recognition. Methods such as convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM) and attention mechanisms have demonstrated powerful capabilities. For example, hybrid models such as CNN-LSTM can effectively combine the spatial feature extraction capabilities of CNN and the temporal feature modeling capabilities of LSTM, thereby comprehensively extracting spatiotemporal features in gait data and improving recognition accuracy and generalization ability. However, these methods still face major challenges in capturing the complexity of gait abnormalities. Abnormal gait patterns may manifest themselves at different time scales, which makes multi-scale feature extraction necessary. At the same time, plantar pressure data contains rich spatial and channel information, requiring complex feature selection and fusion mechanisms for accurate recognition.

[0004] Multi-scale convolution techniques, inspired by the Inception module, have been shown to be able to extract features at different time scales and capture the complexity of abnormal gait patterns. In addition, attention mechanisms such as the Convolutional Block Attention Module (CBAM) further optimize feature selection by dynamically focusing on the most relevant spatial and channel information. However, simple feature fusion methods (such as concatenation or summation) are often used in existing studies, which fail to fully exploit the interactions between temporal, spatial, and channel features. The cross-attention mechanism is an attention mechanism for multi-input sequence models, which focuses on calculating attention across sequences rather than within a single sequence. This mechanism is particularly suitable for modeling plantar pressure data because the data contains key information from multiple sensor channels. By combining the channel attention mechanism and the spatial attention mechanism, the model can dynamically focus on the temporal changes of the gait data and the pressure changes in different regions of the foot, thereby adaptively emphasizing the most informative features and enhancing sensitivity to abnormal gait.

[0005] In model design, the lightweight architecture of the network must also be considered to balance performance and deployment feasibility on resource-constrained devices (such as Raspberry Pi). Therefore, developing an abnormal gait recognition method that can effectively extract multi-scale features, dynamically fuse time, space and channel features, and has the advantage of lightweight has important research and application value. Summary of the invention

[0006] To solve the above technical problems, the present invention proposes an abnormal gait recognition method based on multi-scale convolution and cross-attention fusion, which can accurately identify a variety of abnormal gait patterns.

[0007] To achieve the above object, the present invention provides an abnormal gait recognition method based on multi-scale convolution and cross attention fusion, comprising:

[0008] Obtain abnormal gait dataset;

[0009] Performing data preprocessing according to the abnormal gait dataset and a public Parkinson's disease gait dataset to obtain preprocessed data;

[0010] Constructing an abnormal gait recognition model, and optimizing the abnormal gait recognition model according to the preprocessed data to obtain an optimized abnormal gait recognition model;

[0011] The gait data to be tested is obtained, input into the abnormal gait recognition model, and the abnormal gait recognition result is obtained.

[0012] Optionally, obtaining an abnormal gait dataset includes:

[0013] Gait data from several participants were obtained using pressure insoles;

[0014] Preset labels according to normal gait data and abnormal gait data;

[0015] According to the acquired gait data and preset labels, an abnormal gait dataset is obtained.

[0016] Optionally, obtaining preprocessed data includes:

[0017] Normalizing the abnormal gait dataset and the public Parkinson's disease gait dataset to obtain normalized data;

[0018] The normalized data is segmented according to a preset size to obtain preset data.

[0019] Optionally, constructing an abnormal gait recognition model includes a multi-scale convolution module, a cross-attention fusion module and an abnormal gait recognition module;

[0020] The multi-scale convolution module extracts the temporal feature map using convolution layers with convolution kernels of different sizes, and concatenates the extracted temporal feature maps to obtain a multi-scale feature map;

[0021] The cross attention fusion module is used to further integrate and extract a first intermediate feature map using a pixel-level convolutional layer according to the multi-scale feature map, extract feature maps from the channel dimension and the spatial dimension based on the first intermediate feature map, and cross-fuse the first intermediate feature map and the feature maps extracted in the channel dimension and the spatial dimension to obtain a cross-fused feature map;

[0022] The abnormal gait recognition module is used to perform gait recognition on the fused feature graph to obtain an abnormal gait recognition result.

[0023] Optionally, the cross attention fusion module includes a channel attention unit, a spatial attention unit and a cross attention unit;

[0024] The channel attention unit is used to extract a channel feature map from a channel dimension according to the first intermediate feature map;

[0025] The spatial attention unit is used to extract a spatial feature map from a spatial dimension according to the first intermediate feature map;

[0026] The cross attention unit is used to fuse the first intermediate feature map, the channel feature map and the spatial feature map to obtain a fused feature map.

[0027] Optionally, extracting a channel feature map from a channel dimension according to the first intermediate feature map includes:

[0028] Performing global average pooling processing on the width dimension of the first intermediate feature map to obtain a channel descriptor;

[0029] The channel descriptor is processed through a convolutional layer and a ReLU activation function layer to obtain an intermediate feature map of channel attention;

[0030] Use convolutional layers to restore the original channels;

[0031] Use sigmoid activation function to generate attention weights;

[0032] The first intermediate feature map is reweighted by multiplying the produced attention weights element-wise to generate a channel feature map.

[0033] Optionally, extracting a spatial feature map from a spatial dimension according to the first intermediate feature map includes:

[0034] Performing global average pooling and maximum pooling along the channel dimension according to the intermediate feature map to obtain two spatial descriptors;

[0035] The two spatial descriptors are concatenated and passed through a convolution layer, a batch normalization layer, and a ReLU activation function layer to generate a second intermediate feature map;

[0036] Generate a spatial attention map according to the second intermediate feature map through a Sigmoid function;

[0037] The first intermediate feature map is multiplied element-by-element by the spatial attention map to obtain a spatial feature map.

[0038] Optionally, fusing the first intermediate feature map, the channel feature map, and the spatial feature map to obtain a fused feature map includes:

[0039] Permuting the time and space information of the first intermediate feature map, the channel feature map, and the spatial feature map to generate three corresponding permuted feature maps;

[0040] Perform linear transformation based on the three generated permutation feature maps to extract query, key and value vectors;

[0041] In each branch, the query vector in one feature map interacts with the key vector in another feature map and calculates the attention score through the scaled dot product attention mechanism;

[0042] Calculate cross attention between each pair of branches and generate several cross attention output results;

[0043] The generated cross-attention output results are connected along the time dimension and fused through cross-attention to obtain a fused feature map.

[0044] Optionally, performing abnormal gait recognition on the fused feature graph to obtain an abnormal gait recognition result includes:

[0045] Performing feature conversion on the fused feature map to generate a feature conversion result;

[0046] According to the feature conversion result, feature refinement is performed through a convolution layer with a batch normalization layer and a ReLU activation function, and an abnormal gait recognition result is generated through a fully connected layer and a softmax layer.

[0047] Technical effect of the invention: The present invention discloses an abnormal gait recognition method based on multi-scale convolution and cross-attention fusion, which combines a multi-scale convolution module with a channel and spatial attention mechanism to effectively capture features in time, channel and spatial dimensions. The cross-attention fusion module further enhances the feature representation, enabling the network to accurately identify a variety of abnormal gait patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0049] Figure 1 A schematic diagram of a flow chart of an abnormal gait recognition method based on multi-scale convolution and cross-attention fusion according to an embodiment of the present invention;

[0050] Figure 2 The sensor layout of the pressure insole when collecting the PIAG data set according to the embodiment of the present invention;

[0051] Figure 3 It is the overall framework of the abnormal gait recognition method MSCAF-Gait according to an embodiment of the present invention;

[0052] Figure 4 Schematic diagram of different dimensional features of foot pressure data across time, channel and space domains according to an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the structure of the channel attention module according to an embodiment of the present invention;

[0054] Figure 6 Schematic diagram of the structure of the spatial attention module of an embodiment of the present invention;

[0055] Figure 7 Schematic diagram of the structure of the cross attention module according to an embodiment of the present invention;

[0056] Figure 8Schematic diagram of the confusion matrix of MSCAF-Gait on the GaitinPD dataset according to an embodiment of the present invention, wherein (a) is the confusion matrix obtained for a two-classification task, and (b) is the confusion matrix obtained for a four-classification task based on Hoehn-Yahr classification;

[0057] Fig. 9 This is a schematic diagram of deploying MSCAF-Gait on a Raspberry Pi 4B for real-time abnormal gait recognition according to an embodiment of the present invention;

[0058] Fig.10 Schematic diagram of the inference time of 200 random samples on Raspberry Pi according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0060] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0061] like Figure 1 As shown, this embodiment provides an abnormal gait recognition method based on multi-scale convolution and cross attention fusion, including:

[0062] Obtain abnormal gait dataset;

[0063] Performing data preprocessing according to the abnormal gait dataset and a public Parkinson's disease gait dataset to obtain preprocessed data;

[0064] Constructing an abnormal gait recognition model, and optimizing the abnormal gait recognition model according to the preprocessed data to obtain an optimized abnormal gait recognition model;

[0065] The gait data to be tested is obtained, input into the abnormal gait recognition model, and the abnormal gait recognition result is obtained.

[0066] Furthermore, obtaining abnormal gait data sets includes:

[0067] Gait data from several participants were obtained using pressure insoles;

[0068] Preset labels according to normal gait data and abnormal gait data;

[0069] According to the acquired gait data and preset labels, an abnormal gait dataset is obtained.

[0070] Specifically, in this embodiment, the OpenGo insole developed by Motion is used to collect abnormal gait data to build the abnormal gait data set PIAG. Each insole is embedded with 16 pressure sensors and a 6-axis inertial sensor, and the data collection frequency is 100Hz. The layout of the pressure sensors in PIAG is as follows: Figure 2 As shown in the figure, the pressure sensors on each insole are numbered from 1 to 16. The areas covered by the sixteen sensors include the toes, soles, arches, and heel areas. PIAG was collected from 12 healthy adult participants, including gait data of normal and abnormal gait patterns. The normal gait data are labeled 0-3, including slow walking, normal walking, fast walking, and running, while the abnormal gait is labeled 4-10, including pigeon-toed gait, pigeon-toed gait, right leg pain gait, left leg pain gait, magnetic gait, cross-domain gait, and gluteus medius gait. For each gait type, the data collection time is 2 minutes.

[0071] Furthermore, the preprocessed data includes:

[0072] Normalizing the abnormal gait dataset and the public Parkinson's disease gait dataset to obtain normalized data;

[0073] The normalized data is segmented according to a preset size to obtain preset data.

[0074] Specifically, in addition to the collected abnormal gait dataset PIAG, a public dataset, the Parkinson's disease gait dataset GaitinPD, was selected for the experiment. The GaitinPD dataset can be downloaded at https: / / physionet.org / content / gaitpdb / 1.0.0 / and includes gait data of 93 patients with idiopathic Parkinson's disease and 73 healthy controls. GaitinPD records the vertical ground reaction force of participants walking at a normal speed on a flat road for about two minutes. The data is collected from eight sensors under each foot with a sampling rate of 100Hz.

[0075] When using plantar pressure sensors for gait analysis, the raw data is usually collected from multiple sensors located in different areas of the plantar. In order to facilitate efficient processing, the sliding window technique is used to segment the original abnormal gait dataset PIAG and the Parkinson's disease gait dataset GaitinPD. A sliding window of size T and 50% overlap is applied along the time axis to effectively capture the temporal pattern within the gait cycle. Each pressure sensor is a separate channel. For example, in the PIAG dataset, the data from 32 different pressure points are organized into 32 channels, each of which is in the format of 1×T. Therefore, the input sample X is represented as a tensor of dimension 32×1×T. In order to speed up convergence and enhance feature extraction, Min-Max normalization is applied to the raw data before segmentation. This preprocessing strategy enables the network to more effectively capture the dynamic changes of plantar pressure throughout the gait cycle, thereby improving the accuracy of abnormal gait recognition.

[0076] Furthermore, an abnormal gait recognition model is constructed including a multi-scale convolution module, a cross-attention fusion module and an abnormal gait recognition module;

[0077] The multi-scale convolution module extracts the temporal feature map using convolution layers with convolution kernels of different sizes, and concatenates the extracted temporal feature maps to obtain a multi-scale feature map;

[0078] The cross attention fusion module is used to further integrate and extract a first intermediate feature map using a pixel-level convolutional layer according to the multi-scale feature map, extract feature maps from the channel dimension and the spatial dimension based on the first intermediate feature map, and cross-fuse the first intermediate feature map and the feature maps extracted in the channel dimension and the spatial dimension to obtain a cross-fused feature map;

[0079] The abnormal gait recognition module is used to perform gait recognition on the fused feature graph to obtain an abnormal gait recognition result.

[0080] Specifically, the overall framework of the abnormal gait recognition method MSCAF-Gait is as follows Figure 3 As shown in the figure, temporal features are first extracted from pressure data through multi-scale convolution modules with different temporal receptive fields, so that the network can capture patterns on different time scales. Then, the channel attention and spatial attention mechanisms are used to adjust the weights on the channel and spatial dimensions to enhance feature extraction. The channel attention mechanism focuses on capturing temporal features related to overall pressure changes, while the spatial attention mechanism focuses on identifying pressure change patterns in different areas of the foot. Figure 4As shown in the figure, foot pressure gait data contains multi-dimensional features, including the temporal features of a single pressure sensor changing over time, the channel features of the interaction between multiple sensors, and the spatial features of the overall pressure distribution of both feet. Therefore, a multi-scale convolution module, channel attention, and spatial attention are designed to extract features of these three dimensions respectively. Finally, the cross-attention module is used to fuse these features, enrich the feature representation, and enhance the performance of the network in abnormal gait recognition.

[0081] like Figure 3 As shown in Figure 2, the multi-scale convolution module is specially designed to extract local temporal features from foot pressure sensor data using convolution kernels of different sizes. These convolution kernels with different receptive fields enable the network to capture short-term and long-term dependencies across time, which is crucial for detecting complex patterns in gait sequences. Specifically, a convolution kernel of size k is used. i ∈{1,3,5,7} convolution kernel to process input data Among them C in is the number of input channels, and T is the time length of each window. The output feature map F of each convolutional layer i The calculation of is shown in formula (1), where X represents the preprocessed input data, Indicates that the size is 1×k i Conv2d represents a two-dimensional convolutional neural network. This formulation combines batch normalization (BN), activation function (ReLU), and max pooling layer (MP), and padding is consistently applied in all convolutional layers to ensure the consistency of output dimensions. This consistency helps to efficiently connect output feature maps across different convolutional scales.

[0082]

[0083] After multi-scale convolution, the obtained feature maps F1, F3, F5, and F7 are concatenated along the channel dimension to obtain a multi-scale feature map F multiscale , as shown in formula (2), where C out Represents the aggregated output channel of each convolutional block, and Concat represents the concatenation of feature maps. Multi-scale convolution is used to generate a rich feature matrix that can capture local temporal patterns across multiple scales, thereby significantly enhancing the network's ability to effectively identify various gait patterns.

[0084] F multiscale =Concat(F1,F3,F5,F7) (2).

[0085] Furthermore, the cross attention fusion module includes a channel attention unit, a spatial attention unit and a cross attention unit;

[0086] The channel attention unit is used to extract a channel feature map from a channel dimension according to the first intermediate feature map;

[0087] The spatial attention unit is used to extract a spatial feature map from a spatial dimension according to the first intermediate feature map;

[0088] The cross attention unit is used to fuse the first intermediate feature map, the channel feature map and the spatial feature map to obtain a fused feature map.

[0089] Specifically, Figure 3 As shown, when generating a multi-scale feature map F multiscale After that, a pixel-level convolutional layer is first applied to further integrate the extracted features to generate an intermediate feature map MC. The intermediate feature map MC is then sent to the two parallel branches of channel attention and spatial attention respectively.

[0090] Further, extracting a channel feature map from a channel dimension according to the first intermediate feature map includes:

[0091] Performing global average pooling processing on the width dimension of the first intermediate feature map to obtain a channel descriptor;

[0092] The channel descriptor is processed through a convolutional layer and a ReLU activation function layer to obtain an intermediate feature map of channel attention;

[0093] Use convolutional layers to restore the original channels;

[0094] Use sigmoid activation function to generate attention weights;

[0095] The first intermediate feature map is reweighted by multiplying the produced attention weights element-wise to generate a channel feature map.

[0096] Specifically, Figure 5 As shown in Figure 1, the established channel attention mechanism can highlight important channels in the input feature map, enabling the network to emphasize key information in different feature channels. This technique allows the network to adaptively reweight feature channels according to their relevance to the task. It calculates channel attention weights and then applies them to the input feature map through element-wise multiplication.

[0097] Channel attention mechanism emphasizes feature maps The most important channel in . First, apply the global average pooling operation on the width dimension of the feature map MC to obtain the channel descriptor This descriptor aggregates the average value of each channel, compressing the temporal information into a single value for each channel. The channel descriptor Y is then processed through a 1×1 convolution layer and a ReLU activation function layer. pool , get the intermediate feature map of channel attention Then the convolutional layer is used again to restore the original channel size to Then, the sigmoid activation function σ is used to generate the attention weight. Finally, the original feature map MC is re-weighted by element-by-element multiplication with the generated attention weight to generate the output of the channel attention. The whole process is shown in the following formula, where MC c,1,i Represents the input feature map, Conv2d represents a two-dimensional convolutional layer with a convolution kernel size of 1×1, and ⊙ is element-by-element multiplication.

[0098]

[0099] Y o =Conv2d(ReLU(Conv2d(Y pool ))) (4);

[0100] CA=σ(Y o )⊙MC (5).

[0101] Further, extracting a spatial feature map from a spatial dimension according to the first intermediate feature map includes:

[0102] Performing global average pooling and maximum pooling along the channel dimension according to the intermediate feature map to obtain two spatial descriptors;

[0103] The two spatial descriptors are concatenated and passed through a convolution layer, a batch normalization layer, and a ReLU activation function layer to generate a second intermediate feature map;

[0104] Generate a spatial attention map according to the second intermediate feature map through a Sigmoid function;

[0105] The first intermediate feature map is multiplied element-by-element by the spatial attention map to obtain a spatial feature map.

[0106] Specifically, the spatial attention mechanism focuses on identifying important spatial regions by analyzing the foot pressure distribution. Figure 6 As shown in the figure, the spatial attention module mainly uses the global average pooling layer and the maximum pooling layer to The channel dimension is aggregated to capture important spatial information.

[0107] First, global average pooling and maximum pooling are performed along the channel dimension to obtain two spatial descriptors D avg and D max , where D avg , These descriptors are concatenated and passed through a 1×1 convolution layer, a batch normalization layer (BN), and a ReLU activation function layer to generate an intermediate feature map. Then apply the Sigmoid function (σ) to generate the spatial attention map Finally, the original feature map is multiplied element-wise with the spatial attention map S' to obtain the output of the spatial attention module The process is shown in formulas (6)-(9), where M i is the input feature map, C is the total number of channels of the input feature map, and Concat represents the concatenation of two spatial descriptors D along the channel dimension. avg and D max Concatenate, Conv2d represents a two-dimensional convolution layer with a convolution kernel size of 1×1.

[0108]

[0109] D O =ReLU(BN(Conv2d(Concat(D avg ,D max )))) (7);

[0110] S'=σ(D O ) (8);

[0111] s=S'⊙M (9).

[0112] Further, fusing the first intermediate feature map, the channel feature map, and the spatial feature map to obtain a fused feature map includes:

[0113] Permuting the time and space information of the first intermediate feature map, the channel feature map, and the spatial feature map to generate three corresponding permuted feature maps;

[0114] Perform linear transformation based on the three generated permutation feature maps to extract query, key and value vectors;

[0115] In each branch, the query vector in one feature map interacts with the key vector in another feature map and calculates the attention score through the scaled dot product attention mechanism;

[0116] Calculate cross attention between each pair of branches and generate several cross attention output results;

[0117] The generated cross-attention output results are connected along the time dimension and fused through cross-attention to obtain a fused feature map.

[0118] Specifically, Figure 7The cross-attention module shown aims to enhance the fusion of temporal and spatial information while preserving the integrity of the original feature map. The module has three inputs: the output M of the previous convolutional layer, the channel attention output H, and the spatial attention output S, all of size C×1×W.

[0119] The cross-attention module combines the temporal and spatial information of the input feature maps M, H, S. First, the temporal and spatial information of the input feature maps M, H, S are permuted to generate M', H', S', with the size of W×C. Then a linear transformation is applied to extract the query (q), key (k), and value (v) vectors from each input (M', H', S'), and the formulas are q = XW Q , k = XW k , v = XW V , where W Q , W k , W V is the learned weight matrix. The vector obtained based on M', H', S' is shown in formula (10), where q1, q2 and q3 are query vectors, k1, k2 and k3 are key vectors, and v1, v2 and v3 are value vectors.

[0120] q1,k1,v1←M′; q2,k2,v2←H′; q3,k3,v3←S′ (10);

[0121] In each branch, the query vector in one feature map interacts with the key vector in another feature map and the attention score is calculated using the scaled dot product attention mechanism as shown in the following formula (11), where d k is the dimension of the key vector, which is used to scale the dot product and stabilize the softmax function when processing high-dimensional key vectors. Q, K, and V are the query vector, key vector, and value vector in formula (10), respectively.

[0122]

[0123] Formula (11) is used to calculate the cross attention between each pair of branches, generating 6 cross attention outputs (O c1 , O c2 , O c3 , O c4 , O c5 and O c6 ), as shown in formula (12).

[0124]

[0125] These outputs are then concatenated along the temporal dimension to fuse the cross-attention outputs to form O C , followed by feature transformation to generate the final output O f, as shown in formula (13) and formula (14), where the superscript T indicates the transposition of the six cross-attention outputs, and Concatenate indicates the concatenation of the six outputs along the time dimension.

[0126]

[0127] Final output O f It contains temporal and spatial information extracted by cross-attention across multiple branches. Figure 3 The output O of the criss-cross attention module is shown f Feature refinement is performed through a convolutional layer with batch normalization (BN) and ReLU activation function. Then, the final classification result is generated through a fully connected layer (FC) and a softmax layer to distinguish different gait patterns in abnormal gait recognition.

[0128] Furthermore, the MSCAF-Gait method is tested and evaluated.

[0129] The MSCAF-Gait method was tested and evaluated based on the Parkinson's disease gait dataset GaitinPD and the abnormal gait dataset PIAG.

[0130] Experimental setup: All experiments were implemented using the PyTorch framework on a laptop with an NVIDIA GTX1650 GPU. The dataset was divided into training, validation, and test sets in a ratio of 6:2:2. The Adam optimizer and cross entropy loss function were used to update parameters and minimize the loss during training, and the maximum number of iterations was set to 100. Optuna was used to automatically adjust the learning rate and batch size for hyperparameter optimization to ensure efficient training and enhance performance. The search range for the learning rate was set between 0.00001 and 0.01, and the candidate batch sizes included 32, 64, and 128. After 20 optimization trials, the optimal hyperparameter learning rate was determined to be 0.00019 and the batch size was 128.

[0131] Evaluation indicators: In order to comprehensively evaluate the performance of the MSCAF-Gait method, four commonly used indicators are used: precision, recall, F1-score, and accuracy. The formulas for these indicators are defined as shown in formulas (15) to (18), where TP, FN, FP, and TN represent true positive, false negative, false positive, and true negative, respectively. In order to further measure the computational efficiency of the network, the number of parameters and computational cost (FLOPs) are also used as key evaluation indicators.

[0132]

[0133] Experimental results on the GaitinPD dataset: Two sets of experiments were conducted on the GaitinPD dataset. The first set of experiments was a binary classification task to identify healthy gait (Co) and Parkinson's gait. The second set of experiments was a four-level classification based on the Hoehn-Yahr grading scale, which was divided into level 0 (Co), level 2, level 2.5, and level 3 according to the severity of the disease. In both sets of experiments, the proposed MSCAF-Gait achieved excellent performance. For the binary classification task, the network achieved an accuracy of 99.61%. For the four-classification task of assessing the severity of Parkinson's disease, MSCAF-Gait achieved an accuracy of 98.88%.

[0134] Figure 8 (a)-(b) show the confusion matrices of the MSCAF-Gait network on the GaitinPD dataset for two tasks. Figure 8 (b) It can be seen that the number of test samples for each label in the second set of experiments is as follows: 2327 (level 0), 2730 (level 2), 1866 (level 2.5), and 607 (level 3). MSCAF-Gait has a slight decrease in classification accuracy when distinguishing between gaits with severity levels 2 and 3. The sample size of level 3 gait is significantly smaller than that of level 2 gait, which may be a factor leading to the decrease in classification performance in this category.

[0135] Experimental results on the PIAG dataset: On this dataset, several convolutional neural network (CNN)-based hybrid networks (CNN-LSTM, CNN-SENet, CNN-CBAM, CNN-Self-Attention) and the proposed MSCAF-Gait method are compared. Among them, CNN-LSTM consists of two convolutional layers and a long short-term memory network (LSTM) layer, while CNN-SENet and CNN-CBAM integrate SENet and CBAM modules into two consecutive convolutional layers, respectively. CNN-Self-Attention combines a single convolutional layer with a self-attention mechanism. As shown in Table 1, CNN-LSTM achieved an accuracy of 95.90%, while the model with attention mechanism always outperformed CNN-LSTM. The proposed MSCAF-Gait integrates channel and spatial attention in a parallel structure and fuses cross attention, achieving the highest accuracy of 99.42%.

[0136] Table 1

[0137]

[0138]

[0139] Ablation Experiment Results: Through ablation studies, we can evaluate the contribution of different components in the MSCAF-Gait method to its overall performance, as shown in Table 2. Experiments were conducted on the GaitinPD dataset and the PIAG dataset by gradually adding or removing key components (including multi-scale convolution, channel attention, spatial attention, and cross attention). The experimental configurations (I-VII) represent specific combinations of these components that are selectively enabled or disabled to analyze the impact of different components on classification metrics (accuracy, precision, recall, and F1 score).

[0140] Table 2

[0141]

[0142] As shown in Table 2, the comparison of Experiments I, II, and III shows that adding channel attention or spatial attention can enhance network performance, and the effect of channel attention is slightly better than spatial attention. The comparison of Experiment IV with the full method MSCAF-Gait (Experiment VII) shows that adding cross attention can significantly improve the overall performance. Similarly, comparing Experiment V with the full method MSCAF-Gait (Experiment VII), it can be found that adding multi-scale convolution can improve the accuracy of GaitinPD and PIAG datasets by 1.07% and 0.96%, respectively. In addition, the comparison of Experiment VI with Experiment VII shows that the parallel structure of channel and spatial attention further enhances the network performance. In summary, the full method MSCAF-Gait (Experiment VII) containing all components achieves the highest performance indicators on both datasets - 99.61% accuracy on GaitinPD dataset and 99.42% accuracy on PIAG dataset. These findings emphasize the key role of multi-scale convolution and attention mechanisms (especially the parallel structure of channel and spatial attention combined with cross attention fusion) in optimizing performance.

[0143] Model deployment: To evaluate the real-time deployment performance of MSCAF-Gait, experiments were conducted on the Raspberry Pi 4B platform. After training on the PIAG dataset, the trained model was deployed on the Raspberry Pi4B. The MSCAF-Gait method has 220,000 parameters and 15.83 million FLOPs, which is highly optimized for resource-constrained devices. Fig. 9 As shown, the interface on the lower right shows the output categories of the method, which can identify the gait type and determine whether the gait is abnormal. Fig.10The inference time for 200 randomly selected test samples was recorded, showing an average inference time of approximately 41.24 milliseconds on a Raspberry Pi. These results demonstrate the usefulness of the proposed network for real-time gait analysis, ensuring a balance between accuracy and performance.

[0144] The present invention is specifically designed for abnormal gait recognition using plantar pressure sensors. MSCAF-Gait combines multi-scale convolution modules with channel and spatial attention mechanisms to effectively capture features in temporal, channel and spatial dimensions. An innovative cross-attention fusion module further enhances the feature representation, allowing the network to accurately identify multiple abnormal gait patterns. To facilitate this research, the Pressure Insole Abnormal Gait (PIAG) dataset is introduced, which contains gait data related to common neurological and musculoskeletal abnormalities. Experimental evaluations on the publicly available Parkinson's Disease Gait (GaitinPD) dataset and the PIAG dataset show that MSCAF-Gait achieves 99.61% accuracy in Parkinson's gait detection and 98.88% accuracy in Parkinson's disease severity classification. In the abnormal gait recognition task on the PIAG dataset, MSCAF-Gait achieves 99.42% accuracy. In addition, MSCAF-Gait shows superior performance with low floating point operations (FLOPs) and parameter count while maintaining low computational complexity. The feasibility of deployment on Raspberry Pi verifies its potential for practical applications. These results highlight the advantages of MSCAF-Gait as an efficient and accurate solution for abnormal gait recognition.

[0145] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An abnormal gait recognition method based on multi-scale convolution and cross attention fusion, characterized in that: The following steps are involved: Obtain abnormal gait dataset; Performing data preprocessing according to the abnormal gait dataset and a public Parkinson's disease gait dataset to obtain preprocessed data; Constructing an abnormal gait recognition model, and optimizing the abnormal gait recognition model according to the preprocessed data to obtain an optimized abnormal gait recognition model; The gait data to be tested is obtained, input into the abnormal gait recognition model, and the abnormal gait recognition result is obtained.

2. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 1, characterized in that: Acquiring abnormal gait datasets includes: Gait data from several participants were obtained using pressure insoles; Preset labels according to normal gait data and abnormal gait data; According to the acquired gait data and preset labels, an abnormal gait dataset is obtained.

3. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 1, characterized in that: The preprocessed data include: Normalizing the abnormal gait dataset and the public Parkinson's disease gait dataset to obtain normalized data; The normalized data is segmented according to a preset size to obtain preset data.

4. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 1, characterized in that: The abnormal gait recognition model is constructed, including a multi-scale convolution module, a cross-attention fusion module and an abnormal gait recognition module; The multi-scale convolution module extracts the temporal feature map using convolution layers with convolution kernels of different sizes, and concatenates the extracted temporal feature maps to obtain a multi-scale feature map; The cross attention fusion module is used to further integrate and extract a first intermediate feature map using a pixel-level convolutional layer according to the multi-scale feature map, extract feature maps from the channel dimension and the spatial dimension based on the first intermediate feature map, and cross-fuse the first intermediate feature map and the feature maps extracted in the channel dimension and the spatial dimension to obtain a cross-fused feature map; The abnormal gait recognition module is used to perform gait recognition on the fused feature graph to obtain an abnormal gait recognition result.

5. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 4, characterized in that: The cross attention fusion module includes a channel attention unit, a spatial attention unit and a cross attention unit; The channel attention unit is used to extract a channel feature map from a channel dimension according to the first intermediate feature map; The spatial attention unit is used to extract a spatial feature map from a spatial dimension according to the first intermediate feature map; The cross attention unit is used to fuse the first intermediate feature map, the channel feature map and the spatial feature map to obtain a fused feature map.

6. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 5, characterized in that: Extracting a channel feature map from the channel dimension according to the first intermediate feature map includes: Performing global average pooling processing on the width dimension of the first intermediate feature map to obtain a channel descriptor; The channel descriptor is processed through a convolutional layer and a ReLU activation function layer to obtain an intermediate feature map of channel attention; Use convolutional layers to restore the original channels; Use sigmoid activation function to generate attention weights; The first intermediate feature map is reweighted by multiplying the produced attention weights element-wise to generate a channel feature map.

7. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 5, characterized in that: Extracting a spatial feature map from a spatial dimension according to the first intermediate feature map includes: Performing global average pooling and maximum pooling along the channel dimension according to the intermediate feature map to obtain two spatial descriptors; The two spatial descriptors are concatenated and passed through a convolution layer, a batch normalization layer, and a ReLU activation function layer to generate a second intermediate feature map; Generate a spatial attention map according to the second intermediate feature map through a Sigmoid function; The first intermediate feature map is multiplied element-by-element by the spatial attention map to obtain a spatial feature map.

8. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 5, characterized in that: Fusing the first intermediate feature map, the channel feature map, and the spatial feature map to obtain a fused feature map includes: Permuting the time and space information of the first intermediate feature map, the channel feature map, and the spatial feature map to generate three corresponding permuted feature maps; Perform linear transformation based on the three generated permutation feature maps to extract query, key and value vectors; In each branch, the query vector in one feature map interacts with the key vector in another feature map and calculates the attention score through the scaled dot product attention mechanism; Calculate cross attention between each pair of branches and generate several cross attention output results; The generated cross-attention output results are connected along the time dimension and fused through cross-attention to obtain a fused feature map.

9. The abnormal gait recognition method based on multi-scale convolution and cross attention fusion as claimed in claim 4, characterized in that: Performing abnormal gait recognition on the fused feature graph to obtain an abnormal gait recognition result includes: Performing feature conversion on the fused feature map to generate a feature conversion result; According to the feature conversion result, feature refinement is performed through a convolution layer with a batch normalization layer and a ReLU activation function, and an abnormal gait recognition result is generated through a fully connected layer and a softmax layer.