Multi-modal gait recognition method and system
By combining the multimodal gait recognition method with gait profile map and skeleton Gaussian heat map, deep learning technology is used to extract high-precision gait features, solving the problem of low accuracy of gait recognition in complex environments, and achieving high-precision and robust gait recognition.
Patent Information
- Application Number
- CN202510470927.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
The existing gait recognition technology has low recognition accuracy in complex environments, a shallow feature extraction network, and it is difficult to capture deep features. In addition, the data source in multimodal recognition is single, and time modeling is ignored.
The multimodal gait recognition method is adopted, combining the gait profile map and the skeleton Gaussian heat map, shallow features are extracted through the residual-connected multi-scale pseudo-3D convolution, and feature fusion is performed using the cross-modal attention fusion module, and deep gait features are extracted through the sliding window Transformer module, combining triplets and cross-entropy loss function optimization model.
It significantly improves the accuracy and robustness of gait recognition in complex scenarios, especially in occlusion and complex backgrounds, and improves the recognition performance across perspectives and across scenes.
Smart Images

Figure CN120340134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning, and particularly relates to a gait recognition method and system based on deep learning, which is particularly suitable for identity recognition and behavior analysis in complex environments. Background Art
[0002] Gait is the change in posture during a person's walking. Different from human face, fingerprint, iris, etc., gait is the only biometric feature that can be obtained under long-distance and uncontrolled conditions. Psychological evidence shows that there are certain differences in the gait of each person, so it can be used for identity discrimination; gait recognition refers to the technology of using gait information to identify a person's identity, and has been widely studied and applied in many fields. Although current biometric technologies such as face recognition, fingerprint recognition, and iris recognition have reached high accuracy, most of them have defects such as short recognition distance, limited recognition angle, requirement for active and contact recognition, etc., and are limited in application in some extreme scenarios. In contrast, gait recognition has unique advantages such as non-contact, non-invasive, long-distance recognition, and difficulty in disguise, and can achieve identity recognition at a farther distance and in a more complex environment, showing great application potential in fields such as social security, video surveillance, and biometric authentication.
[0003] Currently, gait recognition methods are divided into two categories based on whether there is a model: model-based gait recognition and model-free gait recognition. The former generally has a large number of model parameters and is difficult to deploy, so the current research focus at home and abroad is on model-free gait recognition.
[0004] The essence of model-free gait recognition is to capture information such as the human body contour and structure in the gait sequence as discriminative gait features to achieve recognition. Traditional model-free gait recognition methods first use principal component analysis or moment feature segmentation method, etc. to extract features from the gait energy map, and then input the feature information into different classification systems for classification, and finally achieve the purpose of gait recognition. With the rapid development of deep learning, the model-free gait recognition method based on deep learning has become one of the mainstream methods of gait recognition. Compared with traditional model-free gait recognition, the algorithm structure based on deep learning not only effectively improves the recognition accuracy of the target person, but also is simpler and more applicable. Hanqing Chao et al. proposed a new perspective of regarding gait as a set of frames, and extracted spatio-temporal information through set pooling and horizontal pyramid projection (HPP), significantly improving the robustness and accuracy of cross-view gait recognition. [1] On this basis, for gait features of different granularities, Zhen Huang et al. proposed a 3D local convolutional neural network (3D Local CNN), which adaptively extracts local 3D volume features and learns spatio-temporal patterns of different body parts. [3]; Beibei Lin et al. proposed the GaitGL network, which extracts global and local features simultaneously through global-local convolutional layers (GLCL) and introduces a masking strategy to enhance local feature learning. [8] In addition, to address the issue of the sensitivity of silhouette images to viewpoints, Tianhuan Huang et al. proposed a cross-view gait recognition method based on adversarial domain adaptation, which outputs gait features robust to viewpoint changes through a generative adversarial network (GAN). [9] ; Shuai Zhang and Chibiao Liu further proposed a viewpoint information elimination mechanism (VIEM), which strips viewpoint information from features through adversarial learning to achieve viewpoint-invariant gait feature learning and significantly improve cross-view gait recognition performance.
[10] Fan Chao et al. revisited existing methods, identified their limitations on outdoor datasets, and proposed a baseline model GaitBase with a simple structure and powerful performance, providing a new benchmark for practical applications. [5] Meanwhile, researchers have also actively explored the recognition effects of different modalities. Jinkai Zheng et al. proposed gait parsing sequences (GPS) to replace traditional silhouette images, and combined with the ParsingGait framework to learn gait features with high information entropy, significantly improving the recognition performance in complex scenarios.
[11] ; Torben Teepe et al. proposed the GaitGraph method, which extracts pure gait features by combining skeleton pose estimation and graph convolutional networks (GCN), avoiding the limitations of traditional silhouette extraction.
[12] On this basis, Fan Chao et al. further proposed a skeleton map representation method, which converts joint coordinates into Gaussian heatmaps, retains structural information and eliminates body shape interference, achieving a performance breakthrough.
[13] .
[0005] However, in order to extract more discriminative features to improve the recognition accuracy, the following problems still need to be solved urgently in the existing technologies: (1) The source of gait data is relatively single. Despite attempts at multi-modal recognition, most models still use human silhouette images or gait energy maps as the only modality, resulting in a significant reduction in accuracy when parts of the human body are occluded. (2) The feature extraction network is relatively shallow. It is difficult to extract deep features, resulting in poor recognition performance in open and complex scenarios (complex backgrounds, extreme viewpoints, poor lighting, etc.). (3) Ignoring temporal modeling. Most gait is regarded as a set of unordered images input into the network, ignoring the distinct features of gait as a periodic time series in the time dimension at different scales.
[0006] The relevant references are as follows:
[0007] [1]Chao H, He Y, Zhang J, et al. Gaitset: Regarding gait as a set for cross - view gait recognition[C] / / Proceedings of the AAAI conference on artificial intelligence. 2019, 33(01): 8126 - 8133.
[0008] [2]Fan C, Peng Y, Cao C, et al. Gaitpart: Temporal part - based model for gait recognition[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 14225 - 14233.
[0009] [3]Huang Z, Xue D, Shen X, et al. 3d local convolutional neural networks for gait recognition[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 14920 - 14929.
[0010] [4]Huang X, Zhu D, Wang H, et al. Context - sensitive temporal feature learning for gait recognition[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021: 12909 - 12918.
[0011] [5] Fan C, Liang J, Shen C, et al. Opengait: Revisiting gait recognition towards better practicality[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023:9707-9716.
[0012] [6] Ge Z, Liu S, Wang F, et al. Yolox: Exceeding yolo series in 2021[J]. arXiv preprint arXiv:2107.08430, 2021.
[0013] [7] Sun R, Zhang Q, Luo C, et al. Human action recognition using a convolutional neural network based on skeleton heatmaps from two-stage pose estimation[J]. Biomimetic Intelligence and Robotics, 2022, 2(3):100062.
[0014] [8] Lin B, Zhang S, Wang M, et al. Gaitgl: Learning discriminative global-local feature representations for gait recognition[J]. arXiv preprint arXiv:2208.01380, 2022. [9] Huang T, Ben X, Gong C, et al. Gaitdan: Cross-view gait recognition via adversarial domain adaptation[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024.
[0015]
[10] Zhang S, Liu C. Cross-view Gait Recognition via View Information Elimination Mechanism[J]. IEEE Access, 2024.
[0016]
[11] Zheng J, Liu X, Wang S, et al. Parsing is all you need for accurate gait recognition in the wild[C] / / Proceedings of the 31st ACM International Conference on Multimedia. 2023:116 - 124.
[0017]
[12] Teepe T, Khan A, Gilg J, et al. Gaitgraph: Graph convolutional network for skeleton - based gait recognition[C] / / 2021 IEEE international conference on image processing(ICIP). IEEE, 2021:2314 - 2318.
[0018]
[13] Fan C, Ma J, Jin D, et al. Skeletongait: Gait recognition using skeleton maps[C] / / Proceedings of the AAAI conference on artificial intelligence. 2024, 38(2):1662 - 1669. Summary of the Invention
[0019] In view of the deficiencies of the prior art, the present invention proposes a multi - modal gait recognition method and system. A shallow feature extraction module based on residual connection is used to perform multi - scale pseudo - 3D convolution on the input gait contour map and skeleton Gaussian heat map. A feature fusion module based on the attention mechanism is used to effectively fuse the shallow features of the two modalities respectively. A sliding window Transformer module is used to extract the deep gait information of the fused features, realizing high - resolution gait feature extraction, so as to achieve high - precision gait recognition in complex scenarios.
[0020] The technical solution adopted by the present invention is as follows: A new multi-modal gait recognition method. First, obtain the gait images of pedestrians in RGB in the environment, and then use the YOLO algorithm to extract high-precision gait contour maps and skeleton maps in complex environments. [6] , and further generate a skeleton Gaussian heat map. The gait contour map and the skeleton Gaussian heat map are respectively used as the inputs of the two branches of the multi-modal gait feature extraction network. Secondly, we make full use of the prior human features to construct a multi-scale convolutional feature extraction module MSConv based on residual links to extract shallow gait features from the input data of the two branches respectively. Then, we construct a cross-modal attention fusion module MFF to perform weight assignment and frame-level feature fusion on the shallow gait features of the two branches. After that, we introduce a sliding window Transformer module to extract deep gait features from the mixed feature representation, model the long-range spatio-temporal dependence of gait information, and obtain the final gait feature vector. On the basis of constructing the above multi-modal gait feature extraction network, we design a hybrid loss function suitable for this network and train it on the CASIA-B and Gait3D datasets respectively. Finally, according to the trained gait recognition model, feature extraction is performed on the input query gait sequence, the obtained query feature vector is calculated for similarity with all feature vectors in the gait identity library, and the similarity ranking of all feature vectors and the obtained feature vector is output to achieve gait recognition. The method includes the following steps:
[0021] Step 1, obtain gait information of the contour and skeleton modalities: Obtain the RGB image of the target person in the target scene, use the YOLO image segmentation and joint prediction algorithm to obtain a series of gait contour maps and the positions of skeleton points, and input the positions of skeleton points into the skeleton Gaussian heat map generation algorithm to obtain the gait skeleton Gaussian heat map;
[0022] Step 2, construct a multi-modal gait feature extraction network for gait recognition. The gait feature extraction network includes a two-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module; input the gait information of the two modalities into the two-branch shallow gait feature extraction module to initially extract shallow gait features, fuse the extracted features through the cross-modal attention fusion module, and then input the fused features into the deep gait feature extraction module to extract discriminative deep gait features;
[0023] Step 3, combine with the loss function, and use the data augmentation strategy to train the multi-modal gait feature extraction network to obtain a high-precision gait feature extraction model in complex scenarios.
[0024] Step 4: For the query gait sequence, input it into the trained gait feature extraction model to obtain the query gait feature vector. Calculate the similarity between all feature vectors in the gait identity library and the query gait feature vector, and output the similarity ranking to achieve gait recognition.
[0025] Furthermore, the shallow feature extraction module consists of two shallow feature encoders, which are used to initially extract different shallow gait features from two modalities. Except for the different input channels, the structures and parameters of the two encoders are the same. Each encoder includes a global feature sub-encoder, a local feature sub-encoder, and a spatio-temporal feature fusion device. For the gait information representation input to the shallow feature encoder, the global feature sub-encoder captures the global gait features of the gait sequence, and the local feature sub-encoder captures the local gait features; the obtained global features and local features are further input into the spatio-temporal feature fusion device to fuse the global-local spatial features and the short-range temporal features between frames, generating a two-branch shallow spatio-temporal feature representation.
[0026] Furthermore, the specific processing process of the shallow encoder is as follows;
[0027] For a given sequence of gait silhouette images or sequence of skeleton heat maps (The 1st, 2nd, 3rd, and 4th dimensions respectively represent the time dimension, channel dimension, height dimension, and width dimension of the input sequence), parallelly send it into the global feature sub-encoder and the local feature sub-encoder to extract global and local feature information. The global feature auto-encoder directly performs several groups of pseudo-3D convolution operations on the feature sequence map, and the local feature sub-encoder uses human prior knowledge to divide the feature sequence into several horizontal strips, and then independently performs several groups of local pseudo-3D convolution-concatenation operations on each strip. The data of the two modalities are encoded by the two sub-encoders to obtain four feature tensors Sil gl 、Sil lc 、Ske gl and Sil lc :
[0028] Sil gl =Gl_P3DConv n (Sil)
[0029] Sil lc =Cat(Lc_P3DConv n (Sil))
[0030] Ske gl =Gl_P3DConv n (Ske)
[0031] Ske lc = Cat(Lc_P3DConv n (Ske))
[0032] where Sil gl , Sil lc , Ske gl and Sil lc represent the global contour feature, local contour feature, global skeleton heatmap feature, and local skeleton heatmap feature respectively; Gl_P3DConv n (·) represents performing n groups of global pseudo-3D convolution operations, and Lc_P3DConv n (·) represents performing n groups of local pseudo-3D convolution operations; Cat(·) represents concatenation along the height dimension.
[0033] The pseudo-3D convolution operation group introduced in this method replaces the traditional m×n×n convolution layer, achieving a balance between information throughput and network parameters. The pseudo-3D convolution operation group consists of two 1×n×n convolutions and one m×1×1 convolution. The input feature vector F in is first output as F1 by the first 1×n×n convolution, and F1 is then converted to F2 by an m×1×1 convolution. After F2 is input to the Leaky ReLU activation function, it is connected with F1 by residual connection to obtain F3. F3 is then passed through the second 1×n×n convolution to obtain F4. Finally, F4 is connected with the original input F in by residual connection and passed through the Leaky ReLU activation function to obtain the final output F out :
[0034] F1 = Conv 1×n×n (F in )
[0035] F2 = Conv m×1×1 (F1)
[0036]
[0037] F4 = Conv 1×n×n (F3)
[0038]
[0039] where represents residual connection. For the Gl_P3DConv(·) and Lc_P3DConv(·) operations, in order to capture gait information at different levels with a more diverse spatial receptive field, in this network, n of Gl_P3DConv(·) is set to 5, and n of Lc_P3DConv(·) is set to 3; m of both methods is set to 3.
[0040] The four obtained feature tensors are input into a spatio-temporal feature fuser for spatial feature fusion and temporal feature aggregation:
[0041]
[0042] Sil st = GeM 3×1×1 (Sil sp )
[0043] Ske st = GeM 3×1×1 (Sil sp )
[0044] where Sil mix and Ske mix represent the spatial features after global-local feature fusion of two modalities, represents element-wise addition; Sil st and Ske st represent the spatial feature fusion and temporal feature aggregation of two modalities, and GeM 3×1×1 (·) represents generalized mean pooling in the time dimension.
[0045] Furthermore, the cross-modal attention fusion module is used to fuse the shallow features extracted from the gait contour map and the skeleton heat map. The specific processing process is as follows:
[0046] The input feature maps of the two input branches and are first concatenated along the channel dimension, and then the cross-modal semantic information is extracted through a small convolutional network by performing 1×1×1, P3D, and 1×1×1 convolutional operations in sequence Then, each pixel value of the feature map is converted into a probability distribution through a Softmax layer to generate the inter-modal attention weights W sol and W ske , and finally, the feature information of the two modalities is weighted and summed according to the weights to output the fused feature map The generation of the weights and the fused feature map can be described by the following formula:
[0047]
[0048] F mix = Sil st W sil + Ske st W ske
[0049] where, represents the cross-modal semantic information F attThe data in the spatio-temporal dimension of (t, i, j) is summed according to Sil st in the channel where it is located; represents the cross-modal semantic information F att The data in the spatio-temporal dimension of (t, i, j) is summed according to Ske st in the channel where it is located.
[0050] Furthermore, the deep feature extraction module is used to extract deep gait features from the contour-skeleton fusion feature map, and includes a pre-processor and several Transformer layers. The specific processing process is as follows:
[0051] To match the sliding window partitioning mechanism, we preprocess the output fusion feature map of the shallow feature extraction module First, perform bilinear interpolation on F mix to facilitate window partitioning; secondly, F mix is input into the patch partition layer to partition image patches and sliding windows. The above partitioned feature tensors are then fed into several Swin Transformer layers for deep feature extraction. A Swin Transformer layer is composed of two Swin Transformer sub-blocks in series, and the basic structures of the sub-blocks are roughly the same: the first sub-block within the group first normalizes the input features through an LN layer, then uses the window self-attention mechanism to perform sliding window Transformer operations on the normalized features, then adds the output of the Transformer self-attention to the original features through a residual connection, and then normalizes and non-linearly transforms the added features through an LN layer and an MLP. Finally, the transformed features are input into the second sub-block for similar feature extraction operations. It should be noted that before the second sub-block performs the window self-attention operation, the self-attention window is slid a certain distance along the length and width dimensions to achieve feature interaction between windows. The features processed by the second sub-block are input into the next Swin Transformer layer for feature extraction. The features processed by the last Swin Transformer layer are used as the output of this module.
[0052] Furthermore, the calculation process of the window self-attention mechanism is specifically as follows: the input feature tensor The i-th feature window of can be expressed as For the first sub-block of each Swin Transformer layer, it calculates self-attention for each feature window:
[0053] Q wi = F wi WQi , K wi = F wi W Ki , V wi = F wi W Vi
[0054]
[0055] A' wi (j, k) = Softmax(A wi (j, k))
[0056] F out_i = A' wi V wi
[0057] where Q wi , K wi , V wi are query, key, and value matrices obtained by linear transformation from F wi , W Qi , W Ki , W Vi are learnable transformation matrices; A wi , A' wi represent the attention scores before and after normalization; F out_i is the output of the sub-block.
[0058] For the second sub-block of each Swin Transformer layer, before calculating the multi-head self-attention for each feature window, window sliding is performed:
[0059]
[0060] where is the i-th feature window after sliding, and M is the sliding step. Using the obtained as the input, window self-attention calculation is performed as described for the first sub-block.
[0061] Furthermore, the following loss function is used as the guidance for network optimization:
[0062]
[0063] where, is the total loss, γ1 and γ2 are two hyperparameters for adjusting different losses, is the triplet loss of the final output of the network. The triplet loss is a metric learning loss function used to optimize the feature space so that samples of the same class are closer in the feature space and samples of different classes are farther apart. Its expression is:
[0064]
[0065] Where a represents the anchor sample, p represents the positive sample of the same class as the anchor, n represents the negative sample of the same class as the anchor, d(·,·) represents the distance between features, and α represents the interval parameter that controls the distance between positive and negative samples.
[0066] is the cross-entropy loss, which is a classification loss function used to measure the difference between the probability distribution predicted by the model and the true label. Its expression is:
[0067]
[0068] where y i is the true label represented in one-hot code, indicating that the sample belongs to the i-th class; p i represents the probability that the sample predicted by the model belongs to the i-th class; N represents the total number of classes.
[0069] Furthermore, the specific implementation of step 1 includes the following sub-steps:
[0070] Step 1.1, continuously capture through the camera to obtain a series of n RGB image gait sequences Im of the target person in the target scene n .
[0071] Step 1.2, respectively use the YOLO image segmentation algorithm and the joint prediction algorithm to perform inference on Im n to obtain the gait contour and the position of the skeleton points of each frame of the RGB image gait sequence:
[0072] Sil i = YOLO_S(Im i )
[0073] where Sil i is the gait contour of the i-th frame image, is the coordinate of the k-th joint point in the plane coordinate system of the i-th frame image, N is the total number of human joint points set, and YOLO_S(·) and YOLO_J(·) represent the inference processes of the forward propagation of the YOLO image segmentation neural network and the joint prediction neural network.
[0074] Step 1.3, generate the skeleton Gaussian heatmap of the gait sequence according to the existing skeleton point positions The following takes the processing of one frame of image as an example. First, perform centering preprocessing on the original joint coordinates (x k , y k ) and scale the human skeleton to the same size to eliminate the interference of gait-irrelevant information such as walking trajectory and shooting distance.
[0075] x' k = x k - x core + W / 2
[0076] y' k = y k - y core + H / 2
[0077] where (x core , y core ) represents the central points of the two hip joints of the human body, and their centers can be regarded as the center of the human body; W and H represent the width and length of the input image respectively.
[0078]
[0079] where, (y max , y min ) represents the maximum and minimum values among the coordinates of all joint points of the human body, and R represents the size of the target area for skeleton scaling. After the above steps, the joint coordinates are scaled into the R×R area at the center of the image, and the aligned joint coordinates (x aligned , y aligned ) are obtained.
[0080] Then, the K-Gaussian method is used to create a skeleton heat map. Create a blank map with a size of R×R. For each pixel (i, j) on the map, calculate the Gaussian distribution according to its distance from all K skeleton joint points:
[0081]
[0082] where, J (i,j) is the Gaussian value of the joint at this point; σ is the standard deviation of the Gaussian distribution and is an adjustable parameter.
[0083] Define the line connecting two naturally adjacent joint coordinate points of the human body as a limb. Using a similar method, the limb Gaussian value L (i,j) of each pixel (i, j) on the map can be calculated as follows:
[0084]
[0085] where S[- - , n + represents the nth limb of the human body, and n - and n + are respectively defined as the two end joints determining a limb; the function D((i, j), S[n - , n + ) represents the Euclidean distance between a point on the map and the nth limb; N represents the total number of human limbs.
[0086] Finally, fill J (i,j) and L (i,j) into two different channels of the R×R blank map to draw the skeleton heat map.
[0087] The present invention also provides a multi-modal gait recognition-based system, including the following units:
[0088] A gait information acquisition unit for acquiring gait information of the contour and skeleton modalities;
[0089] A multi-modal gait feature extraction network construction unit for constructing a multi-modal gait feature extraction network. The gait feature extraction network includes a two-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module. Input the gait information of the two modalities into the two-branch shallow gait feature extraction module to initially extract shallow gait features, fuse the extracted features through the cross-modal attention fusion module, and then input the fused features into the deep gait feature extraction module to extract discriminative deep gait features;
[0090] A training unit for training the multi-modal gait feature extraction network in combination with a loss function by adopting a data augmentation strategy to obtain a high-precision gait feature extraction model in complex scenarios;
[0091] An identification unit for inputting a query gait sequence into the trained gait feature extraction model to obtain a query gait feature vector, calculating the similarity between all feature vectors in the gait identity library and the query gait feature vector, and outputting a similarity ranking to achieve gait recognition.
[0092] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: The present invention proposes a method and system based on multimodal gait recognition, which effectively combines two modal information, gait contour map and skeleton Gaussian heat map, and significantly improves the gait recognition accuracy in complex scenes. By constructing a dual-branch shallow gait feature extraction module, P3D convolution is used to extract global and local features from the gait contour map and skeleton Gaussian heat map respectively, and the feature fusion is combined with the spatiotemporal feature fusion device to fully capture the shallow spatiotemporal features of the gait, and balance the computational complexity and feature representation ability. The cross-modal attention fusion module adaptively weights and fuses the features of the two modalities through the attention mechanism, and complementarily utilizes the different information of the two modalities. The deep gait feature extraction module introduces a sliding window Transformer module to model the long-range spatiotemporal dependency of gait information, extract deep gait features, and enhance the model's ability to capture gait periodicity and time series features. Experimental results show that the fusion of multimodal features significantly improves the robustness and accuracy of gait recognition, especially in the Gait3D dataset with complex scenes such as occlusion and complex background. In addition, by introducing a hybrid loss function combining triplet loss and cross entropy loss, the feature space is further optimized, making the model perform better in cross-view and cross-scenario gait recognition tasks. The method of the present invention outperforms the existing deep learning-based gait recognition methods on public datasets, providing a new solution for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 It is a schematic diagram of the original RGB gait sequence, the contour gait sequence, and the skeleton heat map gait sequence.
[0094] Figure 2 This is the structural diagram of the shallow feature extraction module.
[0095] Figure 3 This is the structural diagram of the feature proofreading module.
[0096] Figure 4 This is the structural diagram of the deep feature extraction module.
[0097] Figure 5 It is a diagram of the multimodal gait recognition network structure. DETAILED DESCRIPTION
[0098] In order to facilitate ordinary technicians in the field to understand and implement the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0099] This invention mainly targets the need for high-precision gait recognition in an open environment. We propose a multimodal gait recognition method and system. Figure 1As shown in the figure, by combining the multimodal information of the gait contour map and the skeleton Gaussian heat map, the deep learning technology is used to achieve high-precision gait recognition in complex scenarios. The YOLO algorithm is used to extract high-precision gait contour maps and skeleton maps, and further generate skeleton Gaussian heat maps. The multi-scale convolutional feature extraction module based on residual connections is used to extract shallow gait features from the input data of the two modalities respectively, as Figure 2 shown. The cross-modal attention fusion module (MFF) is used to assign weights and fuse frame-level features of the shallow gait features of the two branches, as Figure 3 shown. The sliding window Transformer module is used to extract deep gait features from the fused features, model the long-range spatio-temporal dependence relationship of gait information, and obtain the final gait feature vector, as Figure 4 shown. On the basis of constructing a multi-modal gait feature extraction network, we designed a hybrid loss function suitable for this network and trained it on the CASIA-B and Gait3D datasets. Finally, the trained gait recognition model is used to extract features from the input query gait sequence, output the similarity ranking, and achieve gait recognition. The method includes the following steps:
[0100] Step 1, obtain gait information of the contour and skeleton modalities: Obtain the RGB image of the target person in the target scene, use the YOLO image segmentation and joint prediction algorithm to obtain a series of gait contour maps and the positions of skeleton points, and input the positions of skeleton points into the skeleton Gaussian heat map generation algorithm to obtain the gait skeleton Gaussian heat map;
[0101] Step 2, construct a multi-modal gait feature extraction network for gait recognition. The gait feature extraction network includes a two-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module; Input the gait information of the two modalities into the two-branch shallow gait feature extraction module to initially extract shallow gait features, fuse the extracted features through the cross-modal attention fusion module, and then input the fused features into the deep gait feature extraction module to extract discriminative deep gait features;
[0102] Step 3, combine the loss function, use the data augmentation strategy to train the multi-modal gait feature extraction network, and obtain a high-precision gait feature extraction model in complex scenarios;
[0103] Step 4, for the query gait sequence, input it into the trained gait feature extraction model to obtain the query gait feature vector. Calculate the similarity between all feature vectors in the gait identity library and the query gait feature vector, output the similarity ranking, and achieve gait recognition.
[0104] Furthermore, the specific implementation of Step 1 includes the following sub-steps,
[0105] Step 1.1, Continuously capture through a camera to obtain a series of n RGB image gait sequences Im of the target person in the target scene. n . To ensure that the gait sequence covers the entire gait cycle, it is necessary to satisfy n≥30.
[0106] Step 1.2, Respectively adopt the YOLO image segmentation algorithm and the joint prediction algorithm to perform inference on Im n to obtain the gait contour and the position of the skeleton points of each frame of the RGB image gait sequence:
[0107] Sil i =YOLO_S(Im i )
[0108] where Sil i is the gait contour of the i-th frame image, is the coordinate of the k-th joint point in the plane coordinate system of the i-th frame image, N is the total number of human joint points set, and YOLO_S(·) and YOLO_J(·) represent the inference processes of the forward propagation of the YOLO image segmentation neural network and the joint prediction neural network.
[0109] Step 1.3, Generate the skeleton Gaussian heat map of the gait sequence according to the existing skeleton point positions . Take the processing of one frame of image as an example. First, perform centering preprocessing on the original joint coordinates (x k ,y k ) and scale the human skeleton to the same size to eliminate the interference of gait-independent information such as walking trajectory and shooting distance.
[0110] x′ k =x k -x core +W / 2
[0111] y′ k =y j -y core +H / 2
[0112] where (x core ,y core ) represents the center point of the two hip joints of the human body, and their center can be regarded as the center of the human body; W and H represent the width and length of the input image respectively.
[0113]
[0114] where, (y max ,y min) represents the maximum and minimum values among all joint point coordinates of the human body, and R represents the size of the target area for skeleton scaling. After the above steps, the joint coordinates are scaled into the R×R area at the center of the image, obtaining the aligned joint coordinates (x aligned , y aligned ).
[0115] Then, the K-Gaussian method is used to create the skeleton heat map. Create a blank map with a size of R×R. For each pixel (i, j) on the map, calculate the Gaussian distribution according to its distance from all K skeleton joint points:
[0116]
[0117] where J (i,j) is the Gaussian value of the joint at this point; σ is the standard deviation of the Gaussian distribution, which is an adjustable parameter.
[0118] Define the line connecting two naturally adjacent joint coordinate points of the human body as a limb. Using a similar method, the limb Gaussian value L (i,j) for each pixel (i, j) on the map can be calculated as follows:
[0119]
[0120] where S[n - , n + represents the nth limb of the human body, and n - and n + are respectively defined as the two end joints determining a limb; the function D((i, j), S[n - , n + ) represents the Euclidean distance between a point on the map and the nth limb; N represents the total number of human limbs.
[0121] Finally, fill J (i,j) and L (i,j) into two different channels of the R×R blank map to draw the skeleton heat map.
[0122] Furthermore, the specific implementation of step 2 includes the following sub-steps:
[0123] Step 2.1, construct a shallow feature extraction module. For input data of different modalities, we construct a dual-branch shallow gait feature extraction module based on pseudo-3D convolution to preliminarily extract the spatio-temporal features of the gait sequence at multiple scales. The schematic diagram of the module is as shown in the appendix Figure 3 .
[0124] For a given sequence of gait silhouette maps or a sequence of skeleton heat maps (The 1st, 2nd, 3rd, and 4th dimensions represent the time dimension, channel dimension, height dimension, and width dimension of the input sequence, respectively) and send them to the global feature sub-encoder and the local feature sub-encoder in parallel to extract global and local feature information. We choose the residual network ResNet as the backbone of the two encoders with similar structures. Each encoder has two layers, and the layers are connected by a downsampling module (Downsample) to downsample the feature map to an appropriate size to enhance the representation ability of the features. Specifically, the dimension of the feature map input to the first layer is (T×1×64×44) or (T×2×64×44), and the dimension of the feature map input to the second layer after the downsampling module is (T×64×32×22), and the number of residual blocks in each layer is (m1,m2)=(1,2). Among them, the global feature autoencoder directly extracts features from the feature sequence map, and the local feature sub-encoder uses the human body prior knowledge to divide the feature sequence into several horizontal strips, and then independently performs local pseudo-3D convolution-splicing operations on each strip to extract features. Specifically, the number of strips i=4 and the horizontal segmentation ratio k1:k2:k3:k4=1:3:3:1 are used to extract the gait features of the head, torso, legs, and feet of the human body separately. The data of the two modalities are encoded by two sub-encoders to obtain four feature tensors Sil gl Sil lc ,Ske gl With Sil lc :
[0125] Sil gl =Gl_P3DConv n (Sil)
[0126] Sil lc =Cat(Lc_P3DConv n (Sil)
[0127] Ske gl =Gl_P3DConv n (Ske)
[0128] Ske lc =Cat(Lc_P3DConv n (Ske)
[0129] Among them, Sil gl Sil lc ,Ske gl With Sil lc Respectively represent contour global features, contour local features, skeleton heat map global features, and skeleton heat map local features; Gl_P3DConv n (·) represents the execution of n global pseudo 3D convolution operations, Lc_P3DConv n(·) represents the execution of n local pseudo-3D convolution operation groups; Cat(·) represents concatenation along the height dimension.
[0130] The pseudo-3D convolution operation group introduced in this method replaces the traditional m×n×n convolution layer, achieving a balance between information throughput and network parameters. The pseudo-3D convolution operation group consists of two 1×n×n convolutions and one m×1×1 convolution. The input feature vector F in First, it is output as F1 through the first 1×n×n convolution. F1 is then converted to F2 through an m×1×1 convolution. After F2 is input into the Leaky ReLU activation function, it is connected with F1 through a residual connection to obtain F3. F3 then passes through the second 1×n×n convolution to obtain F4. Finally, F4 is in connected with the original input F through a residual connection and passes through the Leaky ReLU activation function to obtain the final output F out :
[0131] F1 = Conv 1×n×n (F in )
[0132] F2 = Conv m×1×1 (F1)
[0133]
[0134] F4 = Conv 1×n×n (F3)
[0135]
[0136] where represents the residual connection. For the Gl_P3DConv(·) and Lc_P3DConv(·) operations, in order to capture gait information at different levels with a more diverse spatial receptive field, in this network, n of Gl_P3DConv(·) is set to 5, and n of Lc_P3DConv(·) is set to 3; m of both methods is set to 3.
[0137] The four obtained feature tensors are input into the spatio-temporal feature fusion device for spatial feature fusion and temporal feature aggregation:
[0138]
[0139] Sil st = GeM 3×1×1 (Sil sp )
[0140] Ske st = GeM 3×1×1 (Sil sp )
[0141] Among them, Sil mix and Ske mix represent the spatial features after the global-local feature fusion of two modalities, represents element-wise addition; Sil st and Ske st represent the feature representations after the fusion of the spatial features of two modalities and the aggregation of the temporal features. GeM 3×1×1 (·) represents generalized mean pooling in the time dimension.
[0142] Step 2.2, construct a cross-modal attention fusion module. Our goal is to perform identity recognition based on the combination of the information of the gait silhouette map and the skeleton Gaussian heat map. For this purpose, we propose an attention-based feature fusion module to adaptively and dynamically aggregate the bimodal features. The schematic diagram of the module is shown in the appendix Figure 2 as follows.
[0143] The input feature maps and of the two input branches are first concatenated along the channel dimension, and then 1×1×1, P3D, and 1×1×1 convolutional operations are sequentially performed through a small convolutional network to extract the cross-modal semantic information Then, each pixel value of the feature map is converted into a probability distribution through a Softmax layer to generate the inter-modal attention weights W sil and W ske , and finally, the feature information of the two modalities is weighted and summed according to the weights to output the fused feature map The generation of the weights and the fused feature map can be described by the following formula:
[0144]
[0145] F mix = Sil st W sil + Ske st W ske
[0146] where, represents the sum of the data of the spatio-temporal dimensions of the cross-modal semantic information F att at (t, i, j) along the channel where Sil st is located; represents the sum of the data of the spatio-temporal dimensions of the cross-modal semantic information F att at (t, i, j) along the channel where Ske st is located.
[0147] Step 2.3, construct a deep feature extraction module. To capture more subtle and complex gait patterns in an open environment, we introduce a sliding window Transformer module for deep gait feature extraction from the contour-skeleton fusion feature map, which includes a preprocessor and several Transformer layers. The schematic diagram of the module is shown in Appendix Figure 4 as follows.
[0148] To match the sliding window partitioning mechanism, we preprocess the output fusion feature map of the shallow feature extraction module . First, perform bilinear interpolation on F mix to change its width and height dimensions from 32×22 to 32×24, facilitating window partitioning; second, F mix is input into the PatchPartition layer. For the convenience of Transformer calculation, we regard each 3D block with a size of 2C×2×2 in the tensor as a token, and take [T, W, H] = [3, 4, 4] as a 3D window (window). The feature tensor after the above partitioning is then fed into 6 Swin Transformer layers for deep feature extraction. A Swin Transformer layer is composed of two Swin Transformer sub-blocks connected in series, and the basic structures of the sub-blocks are roughly the same: the first sub-block within the group first normalizes the input features through an LN layer, then performs a sliding window Transformer operation on the normalized features using the window self-attention mechanism, then adds the output of the Transformer self-attention to the original features through a residual connection, and then normalizes and non-linearly transforms the added features through an LN layer and an MLP. Finally, the transformed features are input into the second sub-block for a similar feature extraction operation. It should be noted that before performing the window self-attention operation, the second sub-block first slides the self-attention window a certain distance along the length and width dimensions to achieve feature interaction between windows. The features processed by the second sub-block are input into the next Swin Transformer layer for feature extraction. After the 4th layer of Transformer operation, a Linear Embedding layer is added to map the feature map (C = 64) to a high-dimensional channel space (C = 128) using a linear transformation, and then input into the next Transformer layer. The features processed by the last Swin Transformer layer are used as the output of this module.
[0149] Furthermore, the calculation process of the window self-attention mechanism is as follows: the i-th feature window of the input feature tensor can be expressed as For the first sub-block of each Swin Transformer layer, it calculates self-attention for each feature window:
[0150] Q wi = F wi W Qi ,K wi = F wi W Ki ,V wi = F wi W Vi
[0151]
[0152] A′ wi (j,k) = Softmax(A wi (j,k))
[0153] F out_i = A′ wi V wi
[0154] where Q wi 、K wi 、V wi are the query, key, and value matrices obtained by linear transformation from F wi ; W Qi 、W Ki 、W Vi are learnable transformation matrices; A wi 、A’ wi represent the attention scores before and after normalization; F out_i is the output of the sub-block.
[0155] For the second sub-block of each Swin Transformer layer, before calculating self-attention for each feature window, window sliding is performed:
[0156]
[0157] where is the i-th feature window after sliding, and M is the sliding step. Using the obtained as the input, window self-attention calculation is performed as described in the first sub-block.
[0158] Furthermore, the specific implementation of step 3 includes the following sub-steps:
[0159] Step 3.1, design the following loss function as the guidance for network optimization:
[0160]
[0161] where, is the total loss, and γ1, γ2 are three hyperparameters for adjusting different losses, with default values of 0.5 in the example. is the triplet loss finally output by the network. The triplet loss is a metric learning loss function used to optimize the feature space so that samples of the same class are closer in the feature space and samples of different classes are farther apart. Its expression is:
[0162]
[0163] where a represents the anchor sample, p represents the positive sample of the same class as the anchor, n represents the negative sample of the same class as the anchor, d(·,·) represents the distance between features, and α represents the margin parameter for controlling the distance between positive and negative samples.
[0164] is the cross-entropy loss, which is a classification loss function used to measure the difference between the probability distribution predicted by the model and the true labels. Its expression is:
[0165]
[0166] where y i is the true label represented in one-hot code, indicating that the sample belongs to the i-th class; p i represents the probability that the sample predicted by the model belongs to the i-th class; N represents the total number of classes.
[0167] Step 3.2: Adopt data augmentation strategies to train the multi-modal gait feature extraction network on the CASIA-B and Gait3D datasets. To prevent the model from overfitting and improve the generalization ability of the model, we have adopted a series of data augmentation strategies, such as horizontal flipping, small-angle rotation, perspective transformation, etc. For the occlusion problem that may be encountered in real recognition applications, we have also adopted the random erasing method, that is, randomly select a rectangular area of appropriate size in the image and set its pixel values to zero to simulate the situation where the pedestrian is partially occluded. After training, a high-precision gait recognition model in an open scene and complex environment is obtained.
[0168] Furthermore, the specific implementation of Step 4 includes the following sub-steps:
[0169] Step 4.1: For the query gait sequence, input the trained gait feature extraction model, and obtain the query gait feature vector q through the forward propagation inference of the model.
[0170] Step 4.2: Calculate the similarity between all feature vectors in the gait identity library and the query gait feature vector. Before identifying q, input the gait sequences of different subjects into the same gait feature extraction model to obtain a gait identity library Q = {q 1, q2, …}. Calculate the Euclidean distances between q and all elements in Q, sort the distances from small to large to obtain the similarity ranking, which is the gait recognition result. The higher the similarity ranking of the gait feature vector q1 and q, the higher the similarity of the identities of the subjects corresponding to q1 and q, that is, the more likely they are the same person.
[0171] Based on the above steps, the gait recognition results on the laboratory dataset and the open environment dataset are obtained. To compare with other methods, we use GaitSet (Chao H, He Y, Zhang J, et al. Gaitset: Regarding gait as a set for cross-view gait recognition[C] / / Proceedings of the AAAI Conference on artificial intelligence. 2019, 33(01): 8126-8133.), Gait Part (Fan C, Peng Y, Cao C, et al. Gaitpart: Temporal part-based model for gait recognition[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 14225-14233.), 3D Local (Huang Z, Xue D, Shen X, et al. 3d local convolutional neural networks for gait recognition[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 14920-14929.), CSTL (Huang X, Zhu D, Wang H, et al. Context-sensitive temporal feature learning for gait recognition[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021: 12909-12918.), GaitBase (Fan C, Liang J, Shen C, et al. Opengait: Revisiting gait recognition towards better practicality[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition.(2023: 9707-9716, etc.) Several methods were compared with our method on laboratory datasets and open environment datasets. To quantitatively evaluate the gait recognition performance, we selected the following indicators to evaluate the recognition performance of the model: Camera Positions, Rank-1, Rank-5, and Rank-10 accuracy (Rank-x Accuracy), mean of Rank-1 accuracy, standard deviation of Rank-1 accuracy Std, mean average precision mAP (mean average precision), and mean Inverse Negative Penalty mINP as evaluation indicators. The smaller the standard deviation value, the better the recognition stability, and the larger the other indicator values, the better the recognition effect. In the quantitative comparison results on the laboratory dataset CASIA-B, different Camera Positions represent the angle between the camera and the human body when collecting gait sequences, and the data below represent the Rank-1 accuracy at this camera view; NM, BG, and CL represent three walking conditions: normal walking, carrying a backpack, and wearing a coat, respectively.
[0172] Table 1 Quantitative analysis of different gait recognition methods on laboratory datasets
[0173]
[0174] The quantitative comparison results on the open environment dataset Gait3D are as follows:
[0175] Table 2 Quantitative analysis of different gait recognition methods on open environment datasets
[0176]
[0177] The results of the quantitative indicators show that the reconstruction results obtained by the method proposed in the present invention are better than the existing methods on both laboratory datasets and open environment datasets, can recognize the target with high accuracy at different perspectives, and have stronger perspective robustness.
[0178] On the other hand, the present invention also provides a multi-modal gait recognition-based system, including the following units:
[0179] A gait information acquisition unit for acquiring gait information in the contour and skeleton modalities;
[0180] The multi-modal gait feature extraction network construction unit is used to construct a multi-modal gait feature extraction network. The gait feature extraction network includes a dual-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module. The gait information of two modalities is input into the dual-branch shallow gait feature extraction module to initially extract shallow gait features. The extracted features are fused through the cross-modal attention fusion module, and then the fused features are input into the deep gait feature extraction module to extract discriminative deep gait features.
[0181] The training unit is used to train the multi-modal gait feature extraction network by combining a loss function and adopting a data augmentation strategy to obtain a gait feature extraction model with high accuracy in complex scenarios.
[0182] The recognition unit is used to input a query gait sequence into the trained gait feature extraction model to obtain a query gait feature vector, calculate the similarity between all feature vectors in the gait identity library and the query gait feature vector, and output a similarity ranking to achieve gait recognition.
[0183] The specific implementation methods of each module are the same as each step, and the present invention will not describe them.
[0184] It should be understood that the parts not elaborated in detail in this specification belong to the prior art.
[0185] It should be understood that the above description of the embodiments is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.
Claims
1. A multi-modal gait recognition method, characterized in that, It includes the following steps: Step 1: Obtain gait information of the contour and skeleton modalities; Step 2: Construct a multi-modal gait feature extraction network for gait recognition. The gait feature extraction network includes a dual-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module. Input the gait information of the two modalities into the dual-branch shallow gait feature extraction module to initially extract shallow gait features, fuse the extracted features through the cross-modal attention fusion module, and then input the fused features into the deep gait feature extraction module to extract discriminative deep gait features; Step 3: Combine the loss function and use the data augmentation strategy to train the multi-modal gait feature extraction network to obtain a high-precision gait feature extraction model in complex scenarios; Step 4: For the query gait sequence, input it into the trained gait feature extraction model to obtain the query gait feature vector, calculate the similarity between all feature vectors in the gait identity database and the query gait feature vector, and output the similarity ranking to achieve gait recognition.
2. The multimodal gait recognition method according to claim 1, characterized in that: Step 1 includes obtaining the RGB image of the target person in the target scene, using the YOLO image segmentation and joint prediction algorithm to obtain a series of gait contour maps and the positions of skeleton points, and inputting the positions of skeleton points into the skeleton Gaussian heat map generation algorithm to obtain the gait skeleton Gaussian heat map.
3. A multimodal gait recognition method according to claim 1 or 2, characterized in that: The specific implementation of Step 1 includes the following sub-steps: Step 1.1, continuously capture through a camera to obtain a series of n RGB image gait sequences Im of the target person in the target scene n ; Step 1.2, respectively use the YOLO image segmentation algorithm and the joint prediction algorithm to perform inference on Im n to obtain the gait contour and the skeleton point positions of the gait sequence of each frame of RGB image: Among them, Sil i is the gait contour of the i-th frame image, is the coordinate of the j-th joint point in the plane coordinate system of the i-th frame image. N is the total number of human joint points set. YOLO_S(·) and YOLO_J(·) represent the inference processes of the forward propagation of the YOLO image segmentation neural network and the joint prediction neural network; Step 1.3, according to the positions of existing skeleton points Generate the skeleton Gaussian heatmap of the gait sequence: First, perform centering preprocessing on the original coordinates to center-align each joint coordinate (x k , y k ) within one gait cycle: x′ k = x k - x core + W / 2 y' k = y k - y core + H / 2 where (x core , y cere ) represents the center points of the two hip joints of the human body; W and H respectively represent the width and length of the input image; Among them, (y max , y min ) represents the maximum and minimum values among the coordinates of all joint points of the human body, and R represents the size of the target area for skeleton scaling; the joint coordinates are scaled into the R×R area at the center position of the image to obtain the aligned joint coordinates (x aligned , y aligned ). Then, use the K-Gaussian method to create a skeleton heat map, create a blank map of size R×R. For each pixel (i, j) on the map, calculate the Gaussian distribution according to its distance from all K skeleton joint points: J (i,j) is the joint Gaussian value at this point; σ is the standard deviation of the Gaussian distribution and is an adjustable parameter; Define the line connecting the coordinate points of two naturally adjacent joints of the human body as a limb. Using the same method, the limb Gaussian value L of each pixel (i, j) on the graph can be calculated. (i,j) : where S[n - ,n + represents the nth limb of the human body, and n - and n + are respectively defined as the two endpoint joints that determine a limb; the function D((i,j), S[n - ,n + ) represents the Euclidean distance between a point on the graph and the nth limb; N represents the total number of human limbs; Finally, fill J (i,j) and L (i,j) into two different channels of the R×R blank map to draw the skeleton heat map.
4. The multimodal gait recognition method according to claim 1, characterized in that: The dual-branch shallow gait feature extraction module consists of two shallow feature encoders, which are used to initially extract different shallow gait features from the two modalities. The two shallow feature encoders have the same structure and parameters except for the input channels. Each shallow feature encoder includes a global feature sub-encoder, a local feature sub-encoder, and a spatio-temporal feature fuser. For the gait information representation input to the shallow feature encoder, the global feature sub-encoder captures the global gait features of the gait sequence, and the local feature sub-encoder captures the local gait features. The obtained global features and local features are further input into the spatio-temporal feature fuser to fuse the global-local spatial features and the short-range inter-frame temporal features, generating a dual-branch shallow spatio-temporal feature representation.
5. The multimodal gait recognition method according to claim 1, wherein: The specific processing process of the shallow encoder is as follows; For a given gait profile sequence Sil or skeleton heat map sequence Ske, it is sent to the global feature sub-encoder and the local feature sub-encoder in parallel to extract global and local feature information; the global feature autoencoder directly performs several sets of pseudo 3D convolution operations on the feature sequence graph, and the local feature sub-encoder uses the human body prior knowledge to divide the feature sequence into several horizontal strips, and then performs several sets of local pseudo 3D convolution-splicing operations on each strip independently; the data of the two modes are encoded by the two sub-encoders to obtain four feature tensors Sil gl Sil lc ,Ske gl With Sil lc : Sil gl = Gl_P3DConv n (Sil) Sil lc = Cat(Lc_P3DConv n (Sil)) Ske gl = Gl_P3DConv n (Ske) Ske lc = Cat(Lc_P3DConv n (Ske)) Among them, Sil gl , Sil lc , Ske gl and Sil lc represent the global contour feature, the local contour feature, the global skeleton heatmap feature, and the local skeleton heatmap feature respectively; Gl_P3DConv n represents the execution of n groups of pseudo-3D convolution operations, and Lc_P3DConv n (·) represents the execution of n groups of local pseudo-3D convolution operations; Cat(·) represents concatenation along the width dimension; the pseudo-3D convolution operation group consists of two 1×n×n convolutions and one m×1×1 convolution. The input feature vector F in is first output as F1 by the first 1×n×n convolution. F1 is then converted to F2 by a m×1×1 convolution. After F2 is input into the Leaky ReLU activation function, it is residually connected with F1 to obtain F3. F3 is then passed through the second 1×n×n convolution to obtain F4. Finally, F4 is residually connected with the original input F in and, after passing through the Leaky ReLU activation function, the final output F out is obtained: F1 = Conv 1×n×n (F in ) F2 = Conv m×1×1 (F1) F4 = Conv 1×n×n (F3) Among them represents a residual connection; the four obtained feature tensors are input into a spatio-temporal feature fusion device for spatial feature fusion and temporal feature aggregation: Sil st = GeM 3×1×1 (Sil sp ) Ske st = GeM 3×1×1 (Sil sp ) Among them, Sil mix and Ske mix represent the spatial features after global-local feature fusion of two modalities, represents element-wise addition; Sil st and Ske st represent the fusion of spatial features of two modalities and the aggregation of temporal features, and GeM 3×1×1 (·) represents generalized mean pooling in the temporal dimension.
6. The multimodal gait recognition method according to claim 1, characterized in that: The cross-modal attention fusion module is used to fuse the shallow features extracted from the gait contour map and the skeleton heat map. The specific processing process is as follows: The input feature maps Sil of the two input branches st and Ske st are first concatenated along the channel dimension, and then the cross-modal semantic information F is extracted through a small convolutional network by performing convolutional operations of 1×1×1, P3D, and 1×1×1 in sequence att . Then, each pixel value of the feature map is converted into a probability distribution through a Softmax layer to generate the inter-modal attention weights W sil and W ske . Finally, the feature information of the two modalities is weighted and summed according to the weights to output the fused feature map F mix ; The generation of the weights and the fused feature map is described by the following formula: Among them, represents the cross-modal semantic information F att The data in the spatio-temporal dimension of (t, i, j) is summed according to Sil st in the channel where it is located; represents the cross-modal semantic information F att The data in the spatio-temporal dimension of (t, i, j) is summed according to Ske st in the channel where it is located.
7. The multimodal gait recognition method according to claim 1, wherein: The deep feature extraction module is used to extract deep gait features from the contour-skeleton fusion feature map, and includes a pre-processor and several Transformer layers. The specific processing process is as follows: To match the sliding window partitioning mechanism, the output fused feature map F of the shallow feature extraction module mix is preprocessed. First, bilinear interpolation is performed on F mix to facilitate window partitioning; second, F mix is input to the chunking layer to partition image chunks and sliding windows. The above partitioned feature tensors are then fed into several Swin Transformer layers for deep feature extraction. One Swin Transformer layer consists of two Swin Transformer sub-blocks connected in series. The basic structures of the sub-blocks are the same: for the first sub-block within the group, first, an LN layer is used to normalize the input features, then the window self-attention mechanism is used to perform sliding window Transformer operations on the normalized features, then the output of the Transformer self-attention is added to the original features through a residual connection, and then the added features are normalized and non-linearly transformed through an LN layer and an MLP. Finally, the transformed features are input into the second sub-block for the same feature extraction operations; before performing the window self-attention operation, the second sub-block first slides the self-attention window a certain distance along the length and width dimensions to achieve feature interaction between windows; the features processed by the second sub-block are input into the next Swin Transformer layer for feature extraction, and the features processed by the last Swin Transformer layer are used as the output.
8. A multimodal gait recognition method according to claim 7, characterized in that: The calculation process of the window self-attention mechanism is as follows: The input feature tensor F w The i-th feature window of wi is denoted as F For the first sub-block of each Swin Transformer layer, it calculates self-attention for each feature window: Q wi = F wi W Qi , K wi = F wi W Ki , V wi = F wi W Vi A′ wi (j,k) = Softmax(A wi (j,k)) F out_i = A' wi V wi Among which Q wi , K wi , V wi are query, key, and value matrices obtained by linearly transforming F wi ; W Qi , W Ki , W Vi are transform matrices available for learning; A wi , A’ wi represent the attention scores before and after normalization; F out_i is the output of the sub-block. For the second sub-block of each Swin Transformer layer, it performs window sliding before calculating the multi-head self-attention for each feature window: wherein is the i-th feature window after sliding, M is the sliding step size, and the obtained is used as the input to perform window self-attention calculation as described in the first sub-block.
9. The multimodal gait recognition method according to claim 1, characterized in that: Adopt the following loss function as the guidance for network optimization: where, is the total loss, and γ1, γ2 are two hyperparameters for adjusting different losses, is the triplet loss finally output by the network; the triplet loss is a metric learning loss function used to optimize the feature space so that samples of the same class are closer in the feature space and samples of different classes are farther apart. Its expression is: Where a represents the anchor sample, p represents the positive sample of the same class as the anchor, n represents the negative sample of the same class as the anchor, d(·,·) represents the distance between features, and α represents the interval parameter that controls the distance between positive and negative samples; is the cross-entropy loss, a classification loss function used to measure the difference between the probability distribution predicted by the model and the true labels, and its expression is: where y i is the true label represented by one-hot code, indicating that the sample belongs to the i-th category; p i represents the probability that the model predicts the sample belongs to the i-th category; N represents the total number of categories.
10. A system based on multi-modal gait recognition, characterized in that, It includes the following units: A gait information acquisition unit, which is used to acquire gait information of the contour and skeleton modalities; A multi-modal gait feature extraction network construction unit, which is used to construct a multi-modal gait feature extraction network. The gait feature extraction network includes a two-branch shallow gait feature extraction module, a cross-modal attention fusion module, and a deep gait feature extraction module. Input the gait information of the two modalities into the two-branch shallow gait feature extraction module to initially extract shallow gait features, fuse the extracted features through the cross-modal attention fusion module, and then input the fused features into the deep gait feature extraction module to extract discriminative deep gait features; A training unit, which is used to train the multi-modal gait feature extraction network by combining a loss function and adopting a data augmentation strategy to obtain a gait feature extraction model with high accuracy in complex scenarios; An identification unit, which is used for a query gait sequence, input it into the trained gait feature extraction model to obtain a query gait feature vector, calculate the similarity between all feature vectors in the gait identity database and the query gait feature vector, and output the similarity ranking to achieve gait recognition.
Citation Information
Cited By
Animal gait recognition method and system based on cross-modal attention mechanism
CN120564269A
Gait recognition method based on contour and skeleton mixed attention feature fusion
CN121214547A
Gait recognition method, device and equipment and storage medium
CN121214552A
Gait recognition method and system based on multi-modal complementary learning
CN121482862A
A gait recognition method and system based on multi-modal complementary learning
CN121482862B