Multi-modal gait recognition method based on space-time semantic modeling and cross-modal cooperation
By employing a multimodal gait recognition method that combines spatiotemporal semantic modeling with cross-modal collaboration, and utilizing a multi-branch structure and feature interaction enhancement module, this method addresses the problem of existing technologies failing to fully utilize spatial and temporal features, and achieves efficient gait recognition in complex environments.
Patent Information
- Application Number
- CN202511035789.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-31
AI Technical Summary
Existing multimodal gait recognition methods fail to fully utilize spatial, temporal, and temporal-spatial features, resulting in insufficient recognition performance and robustness in complex external environments.
A multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration is adopted. The skeleton energy map, contour energy map, contour map and skeleton map are respectively input into four feature extraction branches with the same structure. The lightweight dual-branch attention module and multi-branch spatiotemporal modeling module are used to extract features, and cross-modal fusion is performed through a two-stage feature interaction enhancement module. Finally, mapping and pooling processing are performed in the classification head.
It effectively focuses on key features under different perspectives and walking conditions, preserves the integrity of local channel information, improves gait recognition performance under multiple perspectives and different walking conditions, and enhances recognition accuracy and robustness.
Smart Images

Figure CN120877381A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gait recognition technology and relates to a multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration. Background Technology
[0002] Gait recognition, as an emerging biometric identification method, has broad application prospects, covering multiple fields such as medical rehabilitation, security monitoring, and human-computer interaction. It also demonstrates significant advantages in security, applicability, and user experience. As a unique biometric feature, gait recognition is non-invasive, long-distance, easy to collect, has low requirements for image quality, and is difficult to steal. However, in real-world scenarios, pedestrians typically appear from different perspectives, and images of pedestrians from different perspectives show significant differences. Furthermore, the conditions under which a pedestrian walks, such as wearing a coat or carrying a backpack, are also important factors affecting the recognition rate. Therefore, there is an urgent need to improve the performance of gait recognition models in complex external environments.
[0003] Current gait recognition methods mostly employ single-modal data, with a few multimodal methods performing fusion only once during feature extraction, thus failing to fully leverage the complementary advantages of multimodal approaches. Furthermore, the variations and diversity of gait features across multiple viewpoints significantly impact recognition accuracy and robustness.
[0004] The data used in multimodal gait recognition mainly includes inertial sensor data, contours, skeletons, optical flow, and energy maps. Yi, Sijia et al. combined multimodal data from wearable sensors (such as accelerometers and gyroscopes) and used attention neural networks to extract spatiotemporal features, significantly improving the robustness of gait recognition. Khan et al. combined raw video frames with optical flow and significantly improved recognition accuracy through Bayesian optimization and improved feature selection algorithms. Castro FM et al. used a convolutional neural network (CNN)-based method, fusing grayscale pixels, optical flow, and depth... Figure 3 Gait recognition utilizes the characteristics of pedestrians. GaitCode uses an autoencoder network to fuse acceleration and ground contact force data. Zou et al. proposed a multimodal fusion strategy that fuses contour sequences and skeleton sequences, significantly improving gait recognition performance in complex scenes. Mehdi et al. fused RGB images and depth data using multimodal data fusion to improve accuracy. Transgait combines contour and pose heatmaps to mine pedestrian gait features. These studies demonstrate the feasibility of using multimodal gait recognition.
[0005] Although many researchers have attempted to use multimodal methods for more complex gait problems, existing multimodal fusion methods often neglect the importance of spatial, temporal, and temporal-spatial features for gait integrity, thus limiting the overall performance and robustness of gait recognition. Summary of the Invention
[0006] To address the problems existing in the traditional methods mentioned above, this invention proposes a multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration, which can improve the performance of gait recognition and make it more reliable and effective in practical applications.
[0007] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions: On the one hand, a multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration is provided, the method including the following steps: The skeleton energy map, contour energy map, contour map, and skeleton map are input into four structurally identical feature extraction branches to obtain four modal features and two-stage features for each modality. The feature extraction branches include convolutional pooling modules and dual-path feature enhancement modules. The dual-path feature enhancement modules include an upper branch, a lower branch, and a concatenation operation. The upper branch includes three convolutional modules. The lower branch is used to extract and enhance the semantic information of the data using LDBAM and MBSTM modules. The LDBAM module is used to extract features from different perspectives by using branches that introduce residual connections and depthwise separable convolutions and branches that introduce attention mechanisms. The MBSTM module is used to extract the spatiotemporal changes of gait from different perspectives at different spatial scales through a multi-head self-attention mechanism and a multi-branch hybrid architecture of convolutions.
[0008] The two-stage features of the three cross-modal combinations are respectively input into three two-stage feature interaction enhancement modules to obtain three cross-modal fusion features. The two-stage feature interaction enhancement module is used to fuse the first-stage and second-stage features of the cross-modality respectively, and the fusion results of the two stages are spliced together to obtain the cross-modal fusion features. The first cross-modal combination includes skeleton energy mode and contour energy mode, the second cross-modal combination includes contour energy mode and contour mode, and the third cross-modal combination includes contour mode and skeleton mode.
[0009] The four modal features and three cross-modal fusion features are mapped and pooled before being input into the classification head to obtain the multimodal gait recognition results.
[0010] One of the above technical solutions has the following advantages and beneficial effects: The aforementioned multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration inputs skeleton energy map, contour energy map, contour map, and skeleton map into four structurally identical feature extraction branches to obtain four modal features and two-stage features for each modality. The two-stage features of each cross-modal combination are then input into a two-stage feature interaction enhancement module to obtain cross-modal fusion features. All modal features and cross-modal fusion features are mapped and pooled before being input into a classification head to obtain the multimodal gait recognition result. This method effectively focuses on key features under different viewpoints and walking conditions, preserves the integrity of local channel information, models spatiotemporal information through multiple branches, extracts high-level semantic features, accurately captures the spatiotemporal changes of gait under different viewpoints and walking conditions, and fuses features from different modal data, thereby improving the performance of gait recognition under multiple viewpoints and different walking conditions. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration in one embodiment. Figure 2 This is a schematic diagram of a skeleton energy map, a contour energy map, a contour map, and a skeleton map in one embodiment; Figure 3 This is a general framework diagram of a multimodal gait recognition model based on spatiotemporal semantic modeling and cross-modal collaboration in one embodiment; Figure 4 Here is a block diagram of a lightweight dual-branch attention module in one embodiment; Figure 5 This is a general block diagram of the Multi-Branch Spatiotemporal Modeling Module (MBSTM) in one embodiment; Figure 6 This is an overall block diagram of the multimodal feature fusion module (MFFM) in one embodiment. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0015] It should be noted that, in this document, the reference to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The presentation of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments. The term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, including such combinations.
[0016] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0017] In one embodiment, such as Figure 1 As shown, a multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration is provided, which may include the following processing steps 100 to 104: Step 100: Input the skeleton energy map, contour energy map, contour map, and skeleton map into four feature extraction branches with the same structure to obtain four modal features and two stage features for each modality; the feature extraction branches include a convolutional pooling module and a dual-path feature enhancement module; the dual-path feature enhancement module includes an upper branch, a lower branch, and a concatenation operation; the upper branch includes three convolutional modules; the lower branch is used to extract and enhance the semantic information of the data using the LDBAM module and the MBSTM module; the LDBAM module is used to extract features from different perspectives by using a branch that introduces residual connections and depthwise separable convolutions and a branch that introduces an attention mechanism; the MBSTM module is used to extract the spatiotemporal changes of gait from different perspectives at different spatial scales through a multi-head self-attention mechanism and a multi-branch hybrid architecture of convolutions.
[0018] Specifically, the silhouette map Silhouette, skeleton map Skeleton, silhouette energy map GEI, and skeleton energy map SEI are all derived from the CASIA-B dataset.
[0019] The contour maps were obtained directly from http: / / www.cbsr.ia.ac.cn / english / Gait%20Databases.asp. The contour energy maps were obtained using the method described in reference 1 (Sheth, Abhishek, et al. "Gait Recognition Using Convolutional Neural Network." International Journal of Online & Biomedical Engineering 19.1 (2023)). Multiple contour maps were summed and averaged to obtain the contour energy map for each gait. Figure 2 The second row shown is an example of the synthesized contour energy map. The specific formula for energy map synthesis is as follows:
[0020] Where N represents the number of frames in multiple gait cycles, B i (x,y) represents the contour image.
[0021] The skeleton images are first extracted frame by frame from each video and saved to their respective folders. Then, the OpenPose human pose estimation algorithm is used to extract the human skeleton and save it to the corresponding folder. The skeleton energy map is obtained using the same method as the contour energy map. Each energy map is synthesized from the contour and skeleton maps with a fixed number of frames. This ensures that the four types of data input into the model are one-to-one corresponding in the time dimension, thus enhancing the semantic information of the data during data fusion.
[0022] like Figure 2 The image data shown consists of four types after processing. From top to bottom, the first to fourth rows represent the contour map, contour energy map, skeleton map, and skeleton energy map, respectively.
[0023] Silhouette preserves the spatial distribution information of the moving subject through binarized silhouette; GEI aggregates gait cycle features in a time-accumulated manner, enhancing the representation ability of movement patterns; Skeleton extracts kinematic features based on the topological structure of human key points; and SEI enhances dynamic features through spatiotemporal coding of joint movement trajectories.
[0024] The four feature extraction branches have the same structure. In each feature extraction branch, the input data undergoes convolution and pooling operations to extract features and reduce spatial dimensionality. Then, it enters the dual-path feature enhancement module. The upper branch uses conventional convolution for feature extraction, while the lower branch uses LDBAM and MBSTM modules to extract and enhance the semantic information of the data. After that, the features from the two branches are concatenated by a cat operation to obtain the corresponding modal features. The first-stage features of each feature extraction branch are the features extracted by the first convolution module of the upper branch of its dual-path feature enhancement module. The second-stage features are the features extracted by the second convolution module of the upper branch of its dual-path feature enhancement module and the features extracted by the MBSTM module of its lower branch.
[0025] Lightweight Bi-branch Attention Module (LDBAM): By introducing residual connections and depthwise separable convolutions, the number of model parameters is significantly reduced. This not only alleviates the computational burden but also mitigates the vanishing gradient problem, enabling the model to accurately capture key features from different perspectives while maintaining the integrity of local channel information. This design outperforms existing techniques in both efficiency and performance.
[0026] Multi-branch Spatiotemporal Modeling Module (MBSTM): Employing a multi-head self-attention mechanism and a multi-branch hybrid architecture of convolutions, this module extracts global spatiotemporal features. The multi-branch structure accurately captures spatiotemporal variations of gait across different spatial scales. This multi-branch design allows the model to more comprehensively understand gait features compared to traditional single-view or simpler models. Step 102: Input the two-stage features of the three cross-modal combinations into the three dual-stage feature interaction enhancement modules to obtain three cross-modal fusion features; the dual-stage feature interaction enhancement module is used to fuse the first-stage and second-stage features of the cross-modality respectively, and splice the fusion results of the two stages to obtain the cross-modal fusion features; the first cross-modal combination includes skeleton energy mode and contour energy mode, the second cross-modal combination includes contour energy mode and contour mode, and the third cross-modal combination includes contour mode and skeleton mode.
[0027] Specifically, for each cross-modal combination, the first-stage and second-stage features extracted from the feature extraction branches corresponding to the two modalities are respectively subjected to first-stage cross-modal feature fusion and second-stage cross-modal feature fusion. Then, the resulting first-stage and second-stage cross-modal fused features are concatenated to obtain the cross-modal fused feature. Specifically, in the two-stage cross-modal fusion, the first stage uses the `cat` operation, ECA attention, and CBR module to achieve primary feature interaction and enhancement; the second stage uses the MFFM module and MBSTM module to extract deep features from the fused data. Afterward, the two fused data are concatenated using the `cat` operation.
[0028] This two-stage cross-modal fusion strategy fully leverages the complementary advantages of human contour maps, skeleton maps, contour energy maps, and skeleton energy maps. This method effectively improves the accuracy and robustness of gait recognition, especially under multi-view and multi-walking conditions, overcoming the shortcomings of existing technologies in complex environments.
[0029] Step 104: After mapping and pooling the four modal features and three cross-modal fusion features, input them into the classification head to obtain the multimodal gait recognition results.
[0030] Four modality features and three cross-modal fusion features are horizontally pyramidalized to divide the resulting feature map into several parts, each of which is pooled into a feature vector. Accordingly, several feature vectors are obtained, and these are further mapped to a metric space using a separate fully connected layer. Finally, the loss is calculated using a triplet loss function.
[0031] The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration proposed in this application mainly includes: four feature extraction branches, three two-stage feature interaction enhancement modules, and a classification head.
[0032] The structural block diagram of the multimodal gait recognition model based on spatiotemporal semantic modeling and cross-modal collaboration is as follows: Figure 3 As shown.
[0033] The aforementioned multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration inputs skeleton energy map, contour energy map, contour map, and skeleton map into four structurally identical feature extraction branches to obtain four modal features and two-stage features for each modality. The two-stage features of each cross-modal combination are then input into a two-stage feature interaction enhancement module to obtain cross-modal fusion features. All modal features and cross-modal fusion features are mapped and pooled before being input into a classification head to obtain the multimodal gait recognition result. This method effectively focuses on key features under different viewpoints and walking conditions, preserves the integrity of local channel information, models spatiotemporal information through multiple branches, extracts high-level semantic features, accurately captures the spatiotemporal changes of gait under different viewpoints and walking conditions, and fuses features from different modal data, thereby improving the performance of gait recognition under multiple viewpoints and different walking conditions.
[0034] In one embodiment, in the first feature extraction branch of step 100: the contour map is input into the convolutional pooling module to obtain convolutional pooling features; the convolutional pooling features are input into the upper branch of the dual-path feature enhancement module, the convolutional pooling features are processed by the first convolutional module to obtain the first-stage features of the first modality, the first-stage features of the first modality are processed by the second convolutional module and added to the features output by the MBSTM module of the lower branch of the dual-path feature enhancement module to obtain the second-stage features of the first modality; the second-stage features of the first modality are processed by the third convolutional module to obtain the third-stage features of the first modality; the convolutional pooling features are input into the lower branch of the dual-path feature enhancement module, the convolutional pooling features are processed by the LDBAM module and added to the convolutional pooling features to obtain the first intermediate features, the first intermediate features are processed by the MBSTM module to obtain spatiotemporal features, and the spatiotemporal features are processed by the convolutional module to obtain the lower branch features; the third-stage features and the lower branch features are concatenated to obtain the contour modality features.
[0035] In one embodiment, the LDBAM module includes: a depthwise separable convolutional branch and an attention branch; in the LDBAM module: convolutional pooling features are input into the depthwise separable convolutional branch of the LDBAM module to obtain depthwise convolutional features as follows:
[0036] in, X These are convolutional pooling features. This is a depthwise separable convolution operation; The convolutional pooling features are input into the attention branch of the LDBAM module to obtain the attention features:
[0037] Where M represents the attention feature. X The features are derived from convolutional pooling, where σ is the sigmoid activation function and G is the max pooling operation. It is a two-dimensional convolution.
[0038] The output features of the LDBAM module are obtained by adding the depthwise convolution features and the attention features.
[0039] Specifically, the Lightweight Bi-branch Attention Module (LDBAM) enhances data features by reducing the number of model parameters. The overall block diagram of the Lightweight Bi-branch Attention Module (LDBAM) is shown below. Figure 4 As shown.
[0040] The Lightweight Dual-Branch Attention Module (LDBAM) adopts a dual-branch structure. The upper branch uses residual connections and depthwise separable convolutions to effectively reduce the number of parameters, alleviate the gradient vanishing problem, and preserve the integrity of local channel information. The lower branch uses an attention mechanism to enable the model to better focus on important features in the input data and accurately capture key features from different perspectives.
[0041]
[0042]
[0043]
[0044]
[0045] Where X is the input feature map, DsConv is the depthwise separable convolution, σ is the Sigmoid function, and G is the max pooling function.
[0046] Finally, add Z and M together to get the final result, which is:
[0047] Where Z represents the final result.
[0048] In one embodiment, the MBSTM module includes: a multi-head attention branch, a first convolutional branch, and a second convolutional branch; processing the first intermediate feature through the MBSTM module to obtain spatiotemporal features includes: inputting the first intermediate feature into the multi-head attention branch, the first convolutional branch, and the second convolutional branch respectively to obtain a first feature, a second feature, and a third feature; adding the first feature and the second feature to obtain a first fused feature; adding the second feature and the third feature to obtain a second fused feature; concatenating the first fused feature, the second fused feature, and the first intermediate feature to obtain a concatenated feature; processing the concatenated feature using a convolutional layer and then adding it to the first intermediate feature to obtain the output feature of the MBSTM module.
[0049] In one embodiment, the first feature is:
[0050]
[0051]
[0052] in, As the first feature, X 1 is the first intermediate feature. This is the first splicing feature. This is a feature of multi-head attention. As a multi-head attention mechanism, This is a two-dimensional convolution operation. This is for splicing operations.
[0053] In one embodiment, the first convolutional branch and the second convolutional branch have the same structure, both including three convolutional layers. In the first convolutional branch: the first intermediate feature is input into the first convolutional branch and processed by the first convolutional layer to obtain the first convolutional feature of the first convolutional branch; the first convolutional feature, the first intermediate feature, and the second convolutional feature of the second convolutional branch are added together to obtain the first convolutional fusion feature of the first convolutional branch; the first convolutional fusion feature of the first convolutional branch is processed by the second convolutional layer of the first convolutional branch to obtain the second convolutional feature of the first convolutional branch; the second convolutional feature, the first convolutional fusion feature, and the second convolutional feature of the second convolutional branch are added together and processed by the third convolutional layer of the first convolutional branch to obtain the second feature.
[0054] Specifically, the Multi-Branch Spatiotemporal Modeling Module (MBSTM) accurately captures the spatiotemporal variations of gait from different perspectives at different spatial scales, such as... Figure 5 As shown.
[0055] For the input data X1, it is fed into three branches in parallel. Each branch first performs a convolution operation to extract the features of the input data.
[0056] Multi-head attention branch: The result obtained after the first convolutional layer is F 1_1 Then, perform a cat operation between the obtained result and the original input:
[0057] The results are then used to capture spatial information using multi-head self-attention. In the attention module, multi-head attention computation is first performed on the query vector Q, key vector K, and value vector V. The computation process follows the formula:
[0058]
[0059] In this calculation process, d k As a key vector dimension, it influences the expression and correlation of feature vectors across different dimensions; the number of heads represented by h determines how many different subspace perspectives the model can capture the correlations between gait features; W0, as the output weight matrix, performs further linear transformation and integration on the multi-head attention calculation results, and the final result is denoted as:
[0060] Next, global spatial attention is obtained by performing a convolution operation using a 2D convolution kernel of size 1 and the sigmoid function. The result is then multiplied by F1, and finally, after another convolution layer, the final result of the first branch is obtained, i.e.:
[0061] First and second convolution branches: The results obtained after the first convolution are F2 and F3, which are then added together.
[0062]
[0063]
[0064]
[0065] Then perform a convolution operation on the two branches, followed by an addition operation, that is:
[0066]
[0067]
[0068]
[0069] The results obtained above are then used for feature extraction through convolution. Finally, the results from the multi-head attention branch, the first convolutional branch, and the second convolutional branch are summed pairwise.
[0070]
[0071]
[0072]
[0073] Finally, the results are concatenated and then convolved. The convolution result is added to feature X1 to obtain the output features of the Multi-Branch Spatiotemporal Modeling Module (MBSTM). The specific operation is as follows:
[0074] in, Output features of the Multi-Branch Spatiotemporal Modeling Module (MBSTM).
[0075] In one embodiment, the two-stage feature interaction enhancement module includes: a first-stage feature interaction enhancement module and a second-stage feature interaction enhancement module; the first-stage feature interaction enhancement module includes: a concatenation operation, an ECA attention mechanism, and a CBR module, the CBR module including a convolutional layer, a batch normalization layer, and a ReLU activation function; the second-stage feature interaction enhancement module includes an MFFM module and an MBSTM module; the MFFM module is used to perform cross-modal feature interaction using lightweight channel attention (ECA), combined with batch normalization and the ReLU activation function, to achieve deep spatiotemporal fusion of multi-view features; in the first two-stage feature interaction enhancement module: the first-stage features of the first set of cross-modal combinations are input into the first-stage feature interaction enhancement module, and after processing by the concatenation, ECA attention mechanism, and CBR module, the first-stage fused features are obtained; the second-stage features of the first set of cross-modal combinations are input into the second-stage feature interaction enhancement module, and the MFFM module and MBSTM module are used to perform deep feature extraction of the fused data to obtain the second-stage fused features; the first-stage fused features and the second-stage fused features are concatenated to obtain the first cross-modal fused features.
[0076] In one embodiment, in the MFFM module of the second-stage feature interaction enhancement module of the first two-stage feature interaction enhancement module: the second-stage features of the first group of cross-modal combinations are concatenated and processed by the CBR module to obtain intermediate features; the intermediate features are processed by the ECA attention mechanism and multiplied with the intermediate features to obtain the first intermediate fusion feature; the intermediate fusion feature, the second-stage features of the contour modality, and the second-stage features of the skeleton modality are added together to obtain the second intermediate fusion feature; the second intermediate fusion feature is processed by convolution and batch normalization and added together with the second intermediate fusion feature, and the added feature is activated by the ReLU function and added together with the second-stage features of the contour modality to obtain the output feature of the MFFM module.
[0077] Specifically, the Multimodal Feature Fusion (MFFM) module effectively enhances feature representation capabilities by combining lightweight channel attention (ECA) with batch normalization (BN) and the ReLU activation function. Through innovative cross-modal feature interaction, it achieves deep spatiotemporal fusion of multi-view features, significantly improving gait recognition performance under multi-view and multi-walking conditions. This module's design addresses the shortcomings of existing technologies in feature fusion, improving overall recognition performance.
[0078] High-level semantic features play a crucial role in gait recognition tasks. The Multimodal Feature Fusion (MFFM) module aims to better fuse features from two types of data to obtain deeper semantic information, achieving deep spatiotemporal fusion of multi-view features, such as... Figure 6 As shown.
[0079] First, the two input features (first feature and second feature) are concatenated and summed. Then, convolution is used for feature extraction to obtain intermediate features. To maintain information integrity, ECA attention is also used. The result after attention is multiplied by the intermediate features, and then added to the two original input features to obtain the second intermediate fused feature. Next, a 3×3 convolution is used to extract features from the second intermediate fused feature and perform batch standardization to obtain the second intermediate feature. To reduce information loss, residual connections are used to add the second intermediate feature to the second intermediate fused feature before the ReLU operation. Then, the ReLU activation function is used to obtain the third intermediate feature. Finally, the first feature and the third intermediate feature are added to obtain the final result. The specific operation is as follows:
[0080]
[0081] Where F a and F b The input features are two, and Conv, BN, and ReLU represent the convolutional layer, batch normalization, and ReLU activation function, respectively. The CBR module includes Conv, BN, and ReLU.
[0082] It is worth noting that while existing research also uses multimodal data fusion for gait recognition, most only use two modalities. This paper uses four modalities. The convolutional neural network method used in this paper can be replaced by other deep learning or machine learning methods, such as graph convolutional networks, support vector machines, etc.
[0083] In some implementations, experimental examples are also provided. In this experiment, the proposed network model is evaluated using two publicly available gait datasets: CASIA-B and CASIA-C. Table 1 provides details of the CASIA-B dataset used.
[0084] The CASIA-B dataset contains 124 subjects (001-124), each with three walking conditions (NM01-NM06, carrying a bag BG01-BG02, and wearing a coat CL01-CL02), totaling 10 sequences. Each sequence has 11 viewpoints (0°-180°, 18° intervals), meaning each subject has 110 sequences. Since this dataset does not provide official training and testing set partitioning, this paper experiments with three popular partitioning methods currently in the literature: small sample training (ST), medium sample training (MT), and large sample training (LT). In ST, we used the first 24 subjects for training and the last 100 subjects for testing. In MT, we used the first 62 subjects for training and the last 62 subjects for testing. In LT, we used the first 74 subjects for training and the last 50 subjects for testing. During evaluation, the first four normal walking sequences for each subject were considered as galleries, and the rest as probes.
[0085] CASIA-C is a large-scale image database captured by nighttime infrared (thermal imaging) cameras. This dataset contains infrared images and silhouettes of 153 subjects and is a publicly available single-view (90-degree) gait dataset. The dataset covers four walking states: normal walking (fn), slow walking (fs), fast walking (fq), and walking with a backpack (fb). This embodiment uses the data from the first 100 subjects for model training, and the data from the last 53 subjects for testing. During the testing phase, data numbered fn#01-02 are considered galleries, and data numbered fn#03-04, fs#01-02, fq#01-02, and fb#01-02 are considered probes.
[0086] Implementation details: In this embodiment, the contour data, skeleton data, and energy maps (GEI, SEI) used are all converted to 64×44 using the method in
[52] . For the CASIA-B dataset, during training, the batch size of the model is set to 32 (p×k=4×8), the maximum number of iterations is set to 150K, the model uses the triplet loss function, its boundary parameter is set to 0.2, the optimizer is Adam, and the learning rate is set to 1e-4. For the CASIA-C dataset, during training, the batch size is set to 32 (p×k=4×8), the number of iterations is 80K, and the initial learning rate is set to 1e-4. Regarding the experimental environment and computer hardware configuration, the operating system used is Windows 11, the Python version is 3.7, the PyTorch version is 2.0.1, and the training is performed on an Intel i5-12500 CPU and an RTX 3090 GPU.
[0087] Table 1 shows the number of identities (#ID) and sequences (#Seq) included in the dataset used.
[0088] To comprehensively evaluate the effectiveness of this method, comparative experiments were conducted using GEINet, CNN-LB, GaitSet, GaitPart, CapsNet, MvGAN, GaitNet, and GaitMGL. Following the settings in Reference 2 (Chao, Hanqing, et al. "Gaitset: Regarding gait as a set for cross-view gait recognition." Proceedings of the AAAI conference on artificial intelligence. Vol. 33. No.01. 2019.), AE, GaitSet, MT3D, MGAN, and GaitSlice methods were selected for comparison in medium-sample scenarios. GaitSlice, GaitSet, and MT3D methods were selected for comparison in small-sample scenarios.
[0089] (1) Comparison experiment of CASIA-B dataset In the experiment, small sample (ST), medium sample (MT), and large sample (LT) tests were performed on the CASIA-B dataset, and the dataset splitting followed the currently popular method.
[0090] Tables 2 to 4 show that our method exhibits high recognition accuracy under the large-sample (LT, 74) partitioning of the CASIA-B dataset. Under the three conditions of NM, BG, and CL, the accuracy reaches 96.78%, 92.63%, and 76.28%, respectively. Compared with the MvGAN model, our method improves the recognition rate by 0.71% and 0.73% under the BG and CL conditions, respectively. Compared with the GaitNet model, our method shows a more significant advantage in recognition rate under the three conditions, exceeding them by 4.53%, 3.69%, and 13.99%, respectively.
[0091] Table 2. Rank-1 accuracy (NM condition) on the CASIA-B dataset (large sample LT, 74)
[0092] Table 3. Rank-1 accuracy on the CASIA-B dataset (BG condition) (Large sample LT, 74)
[0093] Table 4. Rank-1 accuracy (CL condition) on the CASIA-B dataset (Large sample LT, 74)
[0094] It is worth noting that the data in Tables 2 to 4 exclude cases with the same perspective.
[0095] The experimental data in Table 5 show that our method exhibits significant advantages in the medium-sample (MT, 62) setting of the CASIA-B dataset. Specifically, in the normal walking (NM) category, our method leads all comparison methods with an average accuracy of 95.62%; it also maintains an excellent performance of 90.08% in the carrying backpack (BG) scenario. Particularly noteworthy is that our method significantly outperforms benchmark methods such as AE and GaitSet in accuracy at multiple key viewpoints (e.g., 18°, 36°, 90°), with a peak accuracy of 99.42% at the NM-36° viewpoint. Although the performance of all methods decreases in the wearing coat (CL) scenario, our method still maintains a relative advantage. These results fully validate the robustness and generalization ability of our method under medium-sample conditions, especially its stability in handling complex scenarios such as viewpoint changes and carrying a backpack.
[0096] Table 5. Rank-1 accuracy on the CASIA-B dataset (MT, 62)
[0097] It is worth noting that the data in Table 5 excludes cases with the same perspective.
[0098] The experimental results in Table 6 show that our method achieves significant progress in a small sample (ST, 24) setting on the CASIA-B dataset. In the normal walking (NM) scenario, our method significantly outperforms the comparison methods (GaitSlice 84.1%, GaitSet 79.5%) with an average accuracy of 86.6%, especially achieving a peak accuracy of 93.7% at a 36° viewpoint. In the carrying backpack (BG) scenario, it also maintains a robust performance of 78.7%, an improvement of 10.1% over GaitSet. However, in the wearing coat (CL) scenario, although the average accuracy of 52.2% is a significant improvement over GaitSet (40.9%), it still lags behind MT3D (56.6%), especially with an accuracy of 50.3% at a 90° viewpoint, indicating that the method's adaptability to complex wearing coat scenarios still needs improvement.
[0099] Table 6. Rank-1 accuracy on the CASIA-B dataset (small sample ST, 24)
[0100] It is worth noting that the data in Table 6 excludes cases with the same perspective.
[0101] (2) Experiments on different frames of the CASIA-B dataset In this implementation, the impact of different sequence lengths on the model's recognition accuracy was evaluated through ablation experiments. The sequence length was gradually increased from 8 frames to 32 frames, and the parameter settings were consistent with the large sample (LT) data in the CASIA-B dataset. Table 7 shows that normal walking (NM) exhibited high robustness across different frame counts, particularly at 32 frames, where the accuracy was highest (96.78%). For carrying a backpack (BG) and wearing a jacket (CL), the recognition rate gradually improved with increasing frame counts, reaching 92.63% and 76.28% respectively at 32 frames. It is noteworthy that the states of carrying a backpack and wearing a jacket are more sensitive to changes in input length, especially performing poorly with fewer frames (8f or 16f). Therefore, this study ultimately set the input sequence length of the model to 32 frames to achieve more stable and superior recognition performance under various conditions.
[0102] Table 7. Average Rank-1 accuracy for different frames
[0103] (3) Experiments on different fusion methods of CASIA-B dataset Traditional methods for multimodal data fusion include the `add` and `cat` operations. However, using `add` may lead to the loss of feature information, while `cat` increases the feature dimensionality, thus increasing computational resources. This method replaces the MFFM fusion module with `add` and `cat` operations respectively, and evaluates its impact on model recognition performance while keeping other network structures unchanged. Table 8 shows the results comparing the proposed Multimodal Feature Fusion Module (MFFM) with `add` (addition) and `cat` (concatenation). The table shows that this method achieves significant improvements, increasing performance by 2.05% and 1.37% respectively.
[0104] Table 8 Comparison of different fusion modules
[0105] (4) Comparison experiment of CASIA-C dataset Table 9 shows the performance comparison of gait recognition on the CASIA-C dataset. Our method significantly outperforms existing methods with an average accuracy of 99.3%, achieving 99.0% accuracy on the most challenging frame-by-frame (FB) task, validating the model's robustness in complex scenarios. Compared to Hanif (95.0%) and DDSTPDN (98.6%), our method achieves a 1.4% performance improvement through innovative design, maintaining stable performance, especially in challenging scenarios such as cross-viewpoint and changing walking conditions, demonstrating its advantages in gait feature extraction and matching.
[0106] Table 9. Accuracy of average Rank-1 on the CASIA-C dataset
[0107] (5) Ablation test To demonstrate the effectiveness of the proposed modules, ablation experiments were conducted on each module. Table 10 clearly shows the performance of each module, improving the accuracy from 85.29% to 88.56%, effectively illustrating the effectiveness of each module. The LDBAM module, while maintaining a lightweight design, effectively focuses on the features of the input feature map, providing a more refined feature base for subsequent processing and effectively improving the recognition rate by 1.08%. The MBSTM module, by capturing gait features in depth at different branches and scales, provides a more comprehensive understanding of the complex dynamics of gait sequences. Finally, the MFFM fusion module integrates features from different processing stages and branches, significantly enhancing the model's resistance to appearance changes and resulting in a significant improvement of 0.81%.
[0108] Table 10 Ablation experimental results on the CASIA-B dataset
[0109] It is worth noting that the Rank-1 accuracy in this table excludes cases with the same viewpoint.
[0110] It should be understood that, although the above process Figure 1 The steps in the diagram are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these steps are executed; they can be performed in other orders. Furthermore, the above process... Figure 1 At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0111] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0112] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and all such modifications and improvements fall within the scope of protection of this application.
Claims
1. A multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration, characterized in that, Including the following steps: The skeleton energy map, contour energy map, contour map, and skeleton map are input into four structurally identical feature extraction branches to obtain four modal features and two-stage features for each modality. Each feature extraction branch includes a convolutional pooling module and a dual-path feature enhancement module. The dual-path feature enhancement module includes an upper branch, a lower branch, and a concatenation operation. The upper branch includes three convolutional modules. The lower branch is used to extract and enhance the semantic information of the data using the LDBAM and MBSTM modules. The LDBAM module is used to extract features from different perspectives using branches that introduce residual connections and depthwise separable convolutions, and branches that introduce attention mechanisms. The MBSTM module is used to extract the spatiotemporal variations of gait from different perspectives at different spatial scales through a multi-head self-attention mechanism and a multi-branch hybrid architecture of convolutions. The two-stage features of the three cross-modal combinations are respectively input into three two-stage feature interaction enhancement modules to obtain three cross-modal fusion features; The dual-stage feature interaction enhancement module is used to fuse the first-stage and second-stage features of cross-modality respectively, and then splice the fusion results of the two stages to obtain cross-modal fused features. The first group of cross-modal combinations includes the skeleton energy mode and the profile energy mode; the second group of cross-modal combinations includes the profile energy mode and the profile mode; and the third group of cross-modal combinations includes the profile mode and the skeleton mode. The four modal features and three cross-modal fusion features are mapped and pooled before being input into the classification head to obtain the multimodal gait recognition results.
2. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 1, characterized in that, In the first feature extraction branch: The contour map is input into the convolutional pooling module to obtain convolutional pooling features; The convolutional pooling features are input into the upper branch of the dual-path feature enhancement module. The convolutional pooling features are processed by the first convolution module to obtain the first-stage features of the first modality. The first-stage features of the first modality are processed by the second convolution module and then added to the features output by the MBSTM module in the lower branch of the dual-path feature enhancement module to obtain the second-stage features of the first modality. The second-stage features of the first modality are processed by the third convolution module to obtain the third-stage features of the first modality. The convolutional pooling features are input into the lower branch of the dual-path feature enhancement module. The convolutional pooling features are processed by the LDBAM module and then added to the convolutional pooling features to obtain the first intermediate features. The first intermediate features are processed by the MBSTM module to obtain the spatiotemporal features. The spatiotemporal features are then processed by the convolution module to obtain the lower branch features. The third-stage features and the lower branch features are concatenated to obtain the contour modal features.
3. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 2, characterized in that, The LDBAM module includes: a depthwise separable convolutional branch and an attention branch; In the LDBAM module: The convolutional pooling features are input into the depthwise separable convolutional branch of the LDBAM module to obtain the depthwise convolutional features: in, X These are convolutional pooling features. This is a depthwise separable convolution operation; The convolutional pooling features are input into the attention branch of the LDBAM module to obtain the attention features as follows: Where M represents the attention feature. X The features are derived from convolutional pooling, where σ is the sigmoid activation function and G is the max pooling operation. It is a two-dimensional convolution; The depthwise convolutional features and the attention features are added together to obtain the output features of the LDBAM module.
4. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 2, characterized in that, The MBSTM module includes: a multi-head attention branch, a first convolutional branch, and a second convolutional branch; The first intermediate feature is processed by the MBSTM module to obtain spatiotemporal features, including: The first intermediate feature is input into the multi-head attention branch, the first convolution branch, and the second convolution branch respectively to obtain the first feature, the second feature, and the third feature; The first feature and the second feature are added together to obtain the first fused feature; The second feature and the third feature are added together to obtain the second fused feature; The first fusion feature, the second fusion feature, and the first intermediate feature are concatenated to obtain the concatenated feature; The spliced features are processed by a convolutional layer and then added to the first intermediate features to obtain the output features of the MBSTM module.
5. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 4, characterized in that, The first feature is: in, As the first feature, X 1 is the first intermediate feature. This is the first splicing feature. This is a feature of multi-head attention. As a multi-head attention mechanism, This is a two-dimensional convolution operation. This is for splicing operations.
6. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 4, characterized in that, The first and second convolutional branches have the same structure, and both the first and second convolutional branches include three convolutional layers. In the first convolution branch: The first intermediate features are input into the first convolutional branch, and after processing by the first convolutional layer, the first convolutional features of the first convolutional branch are obtained. The first convolutional feature of the first convolutional branch, the first intermediate feature, and the second convolutional feature of the second convolutional branch are added together to obtain the first convolutional fusion feature of the first convolutional branch. The first convolutional fusion feature of the first convolutional branch is processed through the second convolutional layer of the first convolutional branch to obtain the second convolutional feature of the first convolutional branch; The second convolutional feature of the first convolutional branch, the first convolutional fusion feature of the first convolutional branch, and the second convolutional feature of the second convolutional branch are added together and then processed through the third convolutional layer of the first convolutional branch to obtain the second feature.
7. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 1, characterized in that, The dual-stage feature interaction enhancement module includes: a first-stage feature interaction enhancement module and a second-stage feature interaction enhancement module; the first-stage feature interaction enhancement module includes: a concatenation operation, an ECA attention mechanism, and a CBR module, the CBR module including a convolutional layer, a batch normalization layer, and a ReLU activation function; the second-stage feature interaction enhancement module includes an MFFM module and an MBSTM module; the MFFM module is used to perform cross-modal feature interaction by employing lightweight channel attention (ECA), combined with batch normalization and the ReLU activation function, to achieve deep spatiotemporal fusion of multi-view features; In the first two-stage feature interaction enhancement module: The first-stage features of the first group of cross-modal combinations are input into the first-stage feature interaction enhancement module. After concatenation, ECA attention mechanism and CBR module processing, the first-stage fused features are obtained. The second-stage features of the first group of cross-modal combinations are input into the second-stage feature interaction enhancement module. The MFFM module and MBSTM module are used to extract deep features from the fused data to obtain the second-stage fused features. The first-stage fusion feature and the second-stage fusion feature are concatenated to obtain the first cross-modal fusion feature.
8. The multimodal gait recognition method based on spatiotemporal semantic modeling and cross-modal collaboration according to claim 7, characterized in that, In the MFFM module within the second-stage feature interaction enhancement module of the first two-stage feature interaction enhancement module: The second-stage features of the first group of cross-modal combinations are concatenated and then processed using the CBR module to obtain intermediate features; The intermediate feature is processed by the ECA attention mechanism and then multiplied with the intermediate feature to obtain the first intermediate fusion feature; The second intermediate fusion feature is obtained by adding the intermediate fusion feature, the second-stage feature of the contour mode, and the second-stage feature of the skeleton mode. The second intermediate fusion feature is processed by convolution and batch normalization, then added to the second intermediate fusion feature. After the addition is activated by the ReLU function, it is added to the second-stage feature of the contour modality to obtain the output feature of the MFFM module.
Citation Information
Cited By
Adaptive gait recognition method based on multi-mode dynamic complementation
CN121527852A
An adaptive gait recognition method based on multi-modal dynamic complement
CN121527852B