Gait recognition method and system based on local and global feature fusion

By combining multi-scale CNN and Transformer models, the local and global features are integrated, and the existing gait recognition methods are solved in terms of accuracy and robustness, achieving higher gait recognition accuracy and processing efficiency.

CN119939345APending Publication Date: 2025-05-06SHANDONG MANAGEMENT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510023283.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing gait recognition methods are difficult to accurately capture subtle differences in complex gait data, and feature extraction based on a single model cannot fully characterize gait features, resulting in insufficient accuracy of the recognition results.

Method used

The gait recognition method based on local and global feature fusion is adopted, and the acceleration and angular velocity information is collected through the inertial measurement unit, combined with the local feature extraction capability of multi-scale CNN and the global feature extraction capability of Transformer, the full extraction of gait features is achieved.

Benefits of technology

It effectively ensures the accuracy in the gait recognition task, improves processing efficiency, and enhances the spatial invariance and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939345A_ABST
    Figure CN119939345A_ABST
Patent Text Reader

Abstract

The invention provides a gait recognition method and system based on local and global feature fusion. The method comprises the following steps: acquiring gait data of an individual to be recognized; wherein the gait data comprises acceleration and angular velocity; the gait data serve as a pre-trained gait recognition model based on deep learning, a gait recognition result is obtained, the gait recognition model comprises a first deep learning model used for extracting multi-scale local features and a second deep learning model used for extracting global features, and a gait recognition result is obtained; the gait recognition specifically executes the following processing processes: taking gait data as input of a first deep learning model, and obtaining multi-scale local gait features; taking the gait data as input of a second deep learning model to obtain global gait features; performing shape and semantic alignment on the local gait features subjected to down-sampling processing and the global special gait features subjected to pooling processing, and performing feature fusion to obtain fused gait features; and inputting the fused gait features into a preset classifier to obtain a gait recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gait recognition, and in particular relates to a gait recognition method and system based on the fusion of local and global features. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In today's society, gait recognition, as a non-invasive biometric recognition technology, has received widespread attention. Compared with other biometric recognition technologies (such as fingerprint, iris or face recognition), gait recognition has the advantages of long-distance recognition and difficulty in disguise, which makes it show great application potential in monitoring, security authentication, and health monitoring of the elderly. However, the complexity and diversity of gait data pose challenges to accurate recognition. Gait is affected by many factors, such as individual physiological structure, walking speed, clothing, etc. These factors will affect gait data, increasing the difficulty of gait recognition.

[0004] The inventors found that traditional gait recognition methods mainly rely on manually extracted features, such as gait energy graph (GEI), gait manifold, etc. These methods can reflect the characteristics of gait to a certain extent, but manually extracted features are difficult to capture subtle differences in complex gait data, limiting the accuracy and robustness of gait recognition; secondly, with the development of deep learning technology, some gait feature extraction methods based on deep learning models have emerged in gait recognition, but these methods often use a single model, and a single model can only collect local or global features of gait, which makes it impossible to accurately characterize the characteristics of gait with the obtained features, which in turn leads to insufficient accuracy of gait recognition results and cannot meet actual needs. Summary of the invention

[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a gait recognition method and system based on the fusion of local and global features. The scheme is based on the acceleration and angular velocity information collected by the inertial measurement unit, combined with the local feature extraction capability of the multi-scale CNN and the global feature extraction capability of the Transformer, to achieve full extraction of gait features, which can not only capture fine-grained local information, but also retain important global information, effectively ensuring the accuracy in the gait recognition task.

[0006] According to a first aspect of an embodiment of the present invention, a gait recognition method based on local and global feature fusion is provided, comprising:

[0007] Acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity;

[0008] The gait data is used as a pre-trained deep learning-based gait recognition model to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features. The gait recognition specifically performs the following processing:

[0009] The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

[0010] Furthermore, the first deep learning model adopts a CNN network model, wherein the CNN network model includes several convolution blocks and combines a spatial variable scale pooling strategy to obtain local gait features from different scales.

[0011] Furthermore, the spatial variable scale pooling strategy is specifically as follows: applying pooling operations of different scales to the feature map output by the convolutional layer of the CNN network model, performing the pooling operations with different window sizes and step sizes, and then concatenating the pooling results of different scales to form a vector of fixed length.

[0012] Furthermore, the feature fusion is specifically as follows: downsampling the local gait features respectively, performing adaptive average pooling on the global gait features, realizing shape and semantic alignment of the downsampled local gait features and the global gait features processed by adaptive average pooling through the reshape function and the fully connected layer respectively, and fusing the aligned local gait features and the global gait features to obtain a fused gait feature.

[0013] Furthermore, the feature fusion specifically adopts a method of adding local gait features and global gait features, and the preset classifier adopts a Softmax classifier.

[0014] Furthermore, the second deep learning model adopts the Transformer model.

[0015] Furthermore, the gait data is obtained based on an inertial measurement unit provided on the individual to be measured.

[0016] According to a second aspect of an embodiment of the present invention, a gait recognition system based on local and global feature fusion is provided, comprising:

[0017] A data acquisition unit, which is used to acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity;

[0018] A gait recognition unit is used to use the gait data as a pre-trained gait recognition model based on deep learning to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features, and the gait recognition specifically performs the following processing:

[0019] The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

[0020] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the method for gait recognition based on fusion of local and global features is implemented.

[0021] According to a fourth aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the gait recognition method based on the fusion of local and global features is implemented.

[0022] One or more of the above technical solutions have the following beneficial effects:

[0023] (1) The present invention provides a gait recognition method and system based on the fusion of local and global features. The scheme is based on the acceleration and angular velocity information collected by the inertial measurement unit, combined with the local feature extraction capability of the multi-scale CNN and the global feature extraction capability of the Transformer, to achieve full extraction of gait features, which can capture fine-grained local information while retaining important global information, effectively ensuring the accuracy in the gait recognition task.

[0024] (2) The scheme described in the present invention abandons the traditional gait recognition scheme that relies on gait energy graph and gait manifold, and replaces the traditional data in gait recognition with the acceleration and angular velocity collected by the inertial measurement unit, thereby effectively improving the processing efficiency of gait recognition while ensuring the accuracy of gait recognition.

[0025] (3) The solution described in the present invention adopts spatial variable scale pooling technology to improve the traditional CNN. The above improvement can realize the processing of input images of any size, enhance the spatial invariance of the model, and improve the computational efficiency; at the same time, it can effectively improve the spatial coverage of features and has strong adaptability.

[0026] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0028] Figure 1 is a flow chart of a gait recognition method based on local and global feature fusion described in an embodiment of the present invention;

[0029] Figure 2 Schematic diagram of the feature extraction process described in an embodiment of the present invention;

[0030] Figure 3 Schematic diagram of the feature fusion process described in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0032] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0033] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0034] Embodiment 1

[0035] The purpose of this embodiment is to provide a gait recognition method based on local and global feature fusion, including:

[0036] Acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity;

[0037] The gait data is used as a pre-trained deep learning-based gait recognition model to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features. The gait recognition specifically performs the following processing:

[0038] The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

[0039] In a specific implementation, the first deep learning model adopts a CNN network model, wherein the CNN network model includes several convolution blocks and combines a spatial variable scale pooling strategy to obtain local gait features from different scales.

[0040] In a specific implementation, the spatial variable-scale pooling strategy is specifically as follows: applying pooling operations of different scales to the feature maps output by the convolutional layer of the CNN network model, performing the pooling operations with different window sizes and step sizes, and then concatenating the pooling results of different scales to form a vector of fixed length.

[0041] In a specific implementation, the feature fusion is specifically as follows: down-sampling the local gait features respectively, performing adaptive average pooling on the global gait features, realizing shape and semantic alignment of the down-sampled local gait features and the global gait features processed by adaptive average pooling respectively through the reshape function and the fully connected layer, and fusing the aligned local gait features and the global gait features to obtain a fused gait feature.

[0042] In a specific implementation, the feature fusion specifically adopts a method of adding local gait features and global gait features, and the preset classifier adopts a Softmax classifier.

[0043] In a specific implementation, the second deep learning model adopts a Transformer model.

[0044] In a specific implementation, the gait data is obtained based on an inertial measurement unit disposed on the individual to be measured.

[0045] Specifically, for ease of understanding, the solution described in this embodiment is described in detail below with reference to the accompanying drawings:

[0046] In order to solve the problems existing in the prior art, such as Figure 1As shown, this embodiment provides a gait recognition method based on the fusion of local and global features, and the solution specifically includes the following processing procedures:

[0047] Step 1: Data Collection

[0048] In this embodiment, an IMU device is used to collect gait data from a walking individual. The IMU device can provide rich information about the individual's movement, including acceleration, angular velocity, etc., which is used as the basis for gait recognition.

[0049] Step 2: Data Preprocessing

[0050] The collected raw IMU data is preprocessed, including filtering, denoising, and standardization steps, to improve the data quality and the accuracy of subsequent processing.

[0051] Step 3: Feature Extraction

[0052] This step is divided into two sub-steps, using spatial multi-scale convolutional neural network (CNN) and Transformer model to extract local and global features of gait respectively:

[0053] Spatial multi-scale convolutional neural network feature extraction: Through the designed CNN architecture with multiple convolution blocks, the spatial variable scale pooling technology is used to extract local features of gait from different scales. Each convolution block extracts features by applying multiple one-dimensional convolution kernels, and gradually reduces the feature length through spatial variable scale pooling, while increasing the number of channels to compensate for the information loss that may be caused by feature dimensionality reduction.

[0054] In the specific implementation, this embodiment adopts a CNN architecture consisting of five convolution blocks to extract local features of gait. Multiple one-dimensional convolution kernels of size 3 are used in each convolution block for feature extraction to retain important local information of gait data. By introducing spatial variable scale pooling technology, the network can perform feature pooling at different scales and effectively capture multi-scale information of gait features. As the network level deepens, the scale of the feature gradually decreases, and the number of feature channels increases accordingly to compensate for the information loss that may be caused by feature downsampling.

[0055] Spatial Pyramid Pooling (SPP) is a technique used in image processing and computer vision tasks, especially in Convolutional Neural Networks (CNNs). It generates a fixed-length output by pooling feature maps at different scales, which solves the limitation of traditional convolutional neural networks in processing input images of different sizes.

[0056] Spatial variable scale pooling is achieved by applying multiple scale pooling operations on the feature map output by the convolution layer. These pooling operations are performed with different window sizes and step sizes, and finally the pooling results of different scales are concatenated to form a vector of fixed length. The specific steps are as follows:

[0057] (1) Generation of feature maps

[0058] Assume that the input image is I, and the feature map after the convolution layer is F, whose size is H×W×C, where H and W are the height and width of the feature map respectively, and C is the number of channels. The feature map is represented as:

[0059] F=f(I;W)

[0060] Among them, f represents the convolution operation and W represents the convolution kernel parameters.

[0061] (2) Spatial pooling operation

[0062] SPP performs pooling operations on feature maps at different scales. Assume that there are N different pooling scales {n1,n2,...,n N}. For each scale n i , divide the feature map into n i ×n i sub-regions, and perform a pooling operation (such as maximum pooling or average pooling) on ​​each sub-region. i The pooling result can be expressed as:

[0063] P i =Pool(F,n i )

[0064] Among them, Pool represents a pooling operation (such as maximum pooling).

[0065] Assume that the size of feature map F is H×W and the pooling window size is h i × i ,in:

[0066]

[0067] The pooling operation performs pooling calculation on each sub-region, and the obtained pooling value constitutes the pooling result P i The pooling calculation formula for each sub-region is:

[0068] P i [k]=max{F[m,n]|(m,n)∈R k}

[0069] Among them, R k represents the pixel set of the kth sub-region,

[0070] (3) Concatenation of feature vectors

[0071] The pooling results of all scales are concatenated together to form a feature vector V of fixed length.

[0072] V=[P1,P2,...,P N ]

[0073] Among them, P i The length is Therefore, the length of the final feature vector V is:

[0074]

[0075] In this way, SPP can process input images of any size and generate output feature vectors of fixed length, effectively improving the robustness and adaptability of the model.

[0076] The above improvement strategy has the following benefits:

[0077] (1) Processing input images of any size: Traditional CNN models usually require input images to have a fixed size, such as 224x224 pixels. Spatial variable-scale pooling technology allows CNN to process input images of any size and scale, because the spatial variable-scale pooling technology (SPP) layer can convert feature maps of different sizes from the previous layer into fixed-length outputs. This capability makes the model more flexible and powerful in dealing with a variety of input sizes in practical applications.

[0078] (2) Enhance the spatial invariance of the model: Spatial scale-variant pooling helps capture the representation of objects in the image at different sizes and perspectives by pooling features from different scales (i.e., resolutions). This makes the model more robust when recognizing objects that are deformed or of different scales.

[0079] (3) Improve computational efficiency: Using the SPP layer can avoid redesigning or fine-tuning the CNN structure every time a new image size is input. Because the output of the SPP layer is fixed, it means that the subsequent fully connected layers do not need to change due to changes in the size of the input image, thereby reducing the complexity of model adjustment and the computational cost of training.

[0080] (4) Improved spatial coverage of features: The SPP layer can better cover features in different regions through its multi-level pooling design and provide richer spatial information. This is particularly helpful for tasks where structural information is particularly important (such as scene parsing, object detection, etc.).

[0081] (5) Strong adaptability: Since its design allows processing inputs of different sizes, SPP technology can be easily integrated into various types of CNN architectures, enhancing the capabilities of existing models without requiring major modifications to the original architecture.

[0082] Transformer-based global feature extraction: An architecture consisting of multiple Transformer Encoder layers is adopted, each of which extracts global features from gait data through a two-headed self-attention mechanism to capture long-range dependencies.

[0083] Specifically, in order to capture long-distance dependencies in gait data, the solution described in this embodiment introduces a Transformer-based global feature extraction branch, which is composed of multiple Transformer Encoder layers, each of which contains a two-headed self-attention mechanism, so that the model can effectively extract the global gait features of the input data.

[0084] In the specific implementation, a spatial multi-scale convolutional neural network (CNN) (hereinafter referred to as the first branch) and a Transformer model (hereinafter referred to as the second branch) are used to extract the local and global features of gait.

[0085] The first branch is based on the local feature extraction branch of CNN (such as Figure 2 As shown in the figure, it consists of 5 convolution blocks. Each convolution block uses multiple one-dimensional convolution kernels of size 3 for feature extraction. As the network model deepens, the length of the feature is gradually reduced to half of the previous block through spatial variable scale pooling operations, and the number of channels is increased to twice that of the previous block to compensate for the loss of gait information that may be caused by feature downsampling.

[0086] The second branch is a Transformer-based global feature extraction branch, which consists of four TransformerEncoder layers. Each TransformerEncoder layer contains a two-head self-attention mechanism, such as Figure 2 As shown in Figure 2, the dataset D is used as input and the global gait features of long-distance dependencies in the input data can be obtained through layer-by-layer Transformer Encoder calculation (Trans embedding).

[0087] In order to solve the problem of feature semantic size misalignment between local gait feature expression (CNN feature) and global gait feature expression (Trans embedding) in the dual-branch algorithm structure, this study proposes a branch feature fusion module, such as Figure 3 This module intends to fuse gait features of different semantic sizes.

[0088] When the features extracted by the two branches are simultaneously input into the branch feature fusion module, the length of the Trans embedding is firstly made consistent with the number of channels of the CNN feature through the downsampling operation. Then, the length of the CNN feature is reduced by the adaptive average pooling (Avgpool) operation. Next, the shapes of the Trans embedding and CNN features are semantically aligned through the reshape operation and the fully connected layer. Finally, the two feature vectors are combined into a composite feature vector through the add operation. The obtained fused feature vector will be input into the Softmax gait classification layer for gait classification.

[0089] In summary, the solution described in this embodiment adopts a dual-branch gait recognition method based on spatial variable scale pooling, combining the local feature extraction capability of CNN and the global feature extraction capability of Transformer. By fusing gait features of different scales, this method can effectively extract fine-grained local features and retain global information, thereby realizing gait recognition.

[0090] Step 4: Feature Fusion

[0091] The features extracted by CNN and Transformer are combined through a specific fusion strategy to take advantage of the complementary advantages of local and global features. First, the dimensions of the two types of features are adjusted to match through downsampling and adaptive average pooling operations, then the shape and semantic alignment are achieved through reshape and fully connected layers, and finally the features are merged through addition operations.

[0092] Specifically, when the features extracted from the dual branches are ready to be fused, the length of the Transformer encoding is first adjusted through a downsampling operation to match the number of channels of the CNN features. Subsequently, adaptive average pooling is used to reduce the dimensionality of the length of the CNN features to facilitate feature fusion. The shape alignment and semantic alignment between the Transformer encoding and the CNN features are achieved through the reshape operation and the fully connected layer. Finally, the two parts of the features are merged through an addition operation to form a fused composite feature vector. The composite feature vector is then input into the Softmax classifier for gait classification.

[0093] The solution described in this embodiment achieves effective fusion of gait features by combining the local feature extraction capability of CNN and the global feature extraction capability of Transformer, which can capture fine-grained local information while retaining important global information, thereby achieving higher accuracy in gait recognition tasks.

[0094] Step 5: Gait Recognition

[0095] The fused feature vector is input into the Softmax classifier for gait classification. The classifier can map the feature vector to the predefined gait category to achieve accurate recognition of individual gait.

[0096] In summary, the solution described in this embodiment realizes an efficient and accurate gait recognition method through IMU data acquisition and preprocessing, feature extraction based on CNN and Transformer, and feature fusion and classification. It can not only capture the detailed features of gait data, but also understand the global structure of the data, thereby providing reliable gait recognition services in multiple application scenarios.

[0097] Embodiment 2

[0098] The purpose of this embodiment is to provide a gait recognition system based on the fusion of local and global features.

[0099] A gait recognition system based on local and global feature fusion, comprising:

[0100] A data acquisition unit, which is used to acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity;

[0101] A gait recognition unit is used to use the gait data as a pre-trained gait recognition model based on deep learning to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features, and the gait recognition specifically performs the following processing:

[0102] The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

[0103] It should be noted here that each module in this embodiment corresponds to each step in Example 1 one by one, and the specific implementation process is the same, which will not be repeated here.

[0104] In further embodiments, there is also provided:

[0105] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, no further description is given here.

[0106] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0107] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0108] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method described in embodiment 1 is completed.

[0109] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0110] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in the present embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A gait recognition method based on local and global feature fusion, characterized in that: include: Acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity; The gait data is used as a pre-trained deep learning-based gait recognition model to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features. The gait recognition specifically performs the following processing: The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

2. A gait recognition method based on local and global feature fusion as claimed in claim 1, characterized in that: The first deep learning model adopts a CNN network model, wherein the CNN network model includes a plurality of convolution blocks and combines a spatial variable scale pooling strategy to obtain local gait features from different scales.

3. A gait recognition method based on local and global feature fusion as claimed in claim 2, characterized in that: The spatial variable scale pooling strategy is specifically: applying pooling operations of different scales to the feature map output by the convolutional layer of the CNN network model, performing the pooling operations with different window sizes and step sizes, and then splicing the pooling results of different scales to form a vector of fixed length.

4. A gait recognition method based on local and global feature fusion as claimed in claim 1, characterized in that: The feature fusion is specifically as follows: down-sampling the local gait features respectively, performing adaptive average pooling on the global gait features, realizing shape and semantic alignment of the down-sampled local gait features and the global gait features processed by adaptive average pooling respectively through a reshape function and a fully connected layer, fusing the aligned local gait features and the global gait features to obtain a fused gait feature.

5. A gait recognition method based on local and global feature fusion as claimed in claim 1, characterized in that: The feature fusion specifically adopts a method of adding local gait features and global gait features, and the preset classifier adopts a Softmax classifier.

6. A gait recognition method based on local and global feature fusion as claimed in claim 1, characterized in that: The second deep learning model adopts the Transformer model.

7. A gait recognition method based on local and global feature fusion as claimed in claim 1, characterized in that: The gait data is obtained based on an inertial measurement unit arranged on the individual to be measured.

8. A gait recognition system based on local and global feature fusion, characterized in that: include: A data acquisition unit, which is used to acquire gait data of the individual to be identified; wherein the gait data includes acceleration and angular velocity; A gait recognition unit is used to use the gait data as a pre-trained gait recognition model based on deep learning to obtain a gait recognition result, wherein the gait recognition model includes a first deep learning model for extracting multi-scale local features and a second deep learning model for extracting global features, and the gait recognition specifically performs the following processing: The gait data is used as the input of the first deep learning model to obtain multi-scale local gait features; the gait data is used as the input of the second deep learning model to obtain global gait features; the local gait features processed by downsampling and the global gait features processed by pooling are respectively aligned in shape and semantics, and then feature fusion is performed to obtain fused gait features; the fused gait features are input into a preset classifier to obtain gait recognition results.

9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the gait recognition method based on local and global feature fusion as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a gait recognition method based on local and global feature fusion as described in any one of claims 1 to 7 is implemented.