Multimodal gait recognition method based on SMPL modal decomposition and embedding fusion

Through the method of fusion of SMPL modal decomposition and embedding, the problem of difficulty in feature extraction and alignment in multimodal gait recognition is solved, and the accuracy and robustness of gait recognition are improved, especially the recognition performance in real scenarios.

CN120032427BActive Publication Date: 2025-08-19TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510177537.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-08-19
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing multimodal gait recognition method lacks targeted optimization of feature extraction during modal fusion, resulting in unsatisfactory recognition performance and difficult feature alignment between modals, which affects the recognition accuracy.

Method used

The SMPL modal decomposition and embedding fusion method is adopted to decompose the SMPL mannequin into SMPL poses and shapes, and an adaptive frame joint attention module is introduced. The modal embedding fusion module and joint loss function are aligned and fused in a unified semantic space to improve feature expression capabilities.

Benefits of technology

Through the application of adaptive frame joint attention module and modal embedding fusion module, the accuracy and robustness of gait recognition are improved, and more refined spatio-temporal feature extraction is achieved, and Rank-1 accuracy, Rank-5 accuracy, mAP and mINP indicators are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032427B_ABST
    Figure CN120032427B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of deep learning technology, and specifically relates to a multimodal gait recognition method based on SMPL modal decomposition and embedding fusion, comprising the following steps: constructing a multimodal gait recognition model DFGait that fuses the SMPL model and the contour; proposing an adaptive frame joint attention module to adaptively extract important joint information of key frames of the gait sequence in the SMPL posture branch; proposing a modal embedding fusion module to achieve efficient fusion of the two modal information by aligning and fusing SMPL model features and contour features in a unified semantic space; and using a joint loss function that combines triplet loss, cross entropy loss, and modal consistency loss to jointly supervise the training of the DFGait model. The present invention decomposes the SMPL human body model and fully extracts the gait information in terms of both dynamic posture and static shape contained in the SMPL model. Branch optimization is performed through the adaptive frame joint attention module to achieve more refined spatiotemporal feature extraction of gait information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning technology, and specifically relates to a multimodal gait recognition method based on SMPL modal decomposition and embedding fusion. Background Art

[0002] In recent years, with the rapid development of deep learning technology, biometric recognition technologies, such as facial recognition, fingerprint recognition, iris recognition, and gait recognition, have gradually become ubiquitous in daily life. Gait recognition offers significant advantages over other recognition technologies due to its ability to identify individuals from a distance, its lack of active cooperation, and its susceptibility to environmental factors. Therefore, it holds great potential in specific application scenarios such as security and intelligent transportation. However, due to complex environmental factors and high computing resource requirements, effective gait recognition technology remains difficult to implement on a large scale in real-world settings.

[0003] Previous gait recognition research can be roughly divided into two categories: mainstream appearance-based methods generally use silhouettes as input, providing rich appearance information. However, silhouettes are sensitive to appearance variations such as clothing and viewpoint, and fail to capture the body's internal structure. Model-based methods generally use skeletons as input. Skeletons provide precise information about the body's structure and motion, making them less susceptible to external interference, but they lack discriminative appearance information. This single modality has limitations when dealing with the complex and changing real-world scenarios.

[0004] Compared to single-modality methods, multimodal gait recognition technology can leverage the strengths of different modalities to complement each other. Even if one modality is interfered with, the other modality can still provide sufficient information to support the recognition task. Many works have explored multimodal methods, such as TriGait, MSAFF, and SkeletonGait. These methods have achieved some success in integrating visual and structural information by fusing contour and skeletal features. However, traditional skeletal models have limitations due to their lack of appearance information. In contrast, the SMPL human body model, as an emerging modality, demonstrates unique advantages. SMPL is a parametric three-dimensional human body model that, unlike other models, provides information on both shape and posture, encoding the human body model using these two parameters. Shape parameters control a person's height, weight, head-to-body ratio, and other factors, while posture parameters describe the relative angles of various joints. The SMPL human body model is a relatively new and promising model in gait recognition. Recent studies, such as SMPLGait and HybridGait, have attempted to combine the SMPL model with the contour modalities, promoting the further development of gait recognition technology.

[0005] However, multimodal gait recognition methods still face several challenges, which, to a certain extent, affect recognition accuracy. First, multimodal modeling often focuses on the fusion of modalities, but lacks targeted optimization of feature extraction within each modal branch, resulting in suboptimal recognition performance. Furthermore, due to the differences in data formats between modalities during modal fusion, feature alignment becomes a critical issue that needs to be addressed. Summary of the Invention

[0006] To address the technical problem that the above-mentioned multimodal modeling lacks targeted optimization of feature extraction of each modal branch, the present invention provides a multimodal gait recognition method based on SMPL modal decomposition and embedding fusion. The SMPL human body model is decomposed into SMPL posture and SMPL shape, and the two are used as input together with the silhouette modality. To improve the expressive power of the SMPL model features, an adaptive frame joint attention module is introduced in the SMPL posture branch to compensate for the shortcomings of the existing model in SMPL feature extraction. In addition, by using a modal embedding fusion module and a modal consistency loss function, efficient alignment and fusion of the SMPL model and the silhouette modality in feature expression are achieved, fully leveraging the complementary advantages of the two modalities in a unified feature expression.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0008] The multimodal gait recognition method based on SMPL modal decomposition and embedding fusion includes the following steps:

[0009] S1. Construct a multimodal gait recognition model DFGait that integrates the SMPL model and the contour;

[0010] S2. We propose an adaptive frame joint attention module to adaptively extract important joint information of gait sequence keyframes in the SMPL posture branch.

[0011] S3. We propose a modality embedding fusion module to achieve efficient fusion of two modal information by aligning and fusing SMPL model features and contour features in a unified semantic space.

[0012] S4. A joint loss function combining triplet loss, cross entropy loss, and modality consistency loss is used to jointly supervise the training of the DFGait model.

[0013] The method for constructing the multimodal gait recognition model DFGait that integrates the SMPL model and the contour in S1 is as follows: decomposing the SMPL model into two independent branches, SMPL shape and SMPL posture, focusing on the extraction of SMPL shape features and SMPL posture features respectively; thus, the model constructs three branches: the contour branch is responsible for capturing the appearance and motion features in the contour sequence, the SMPL shape branch is responsible for extracting the shape features of the SMPL human body model, and the SMPL posture branch is focused on extracting dynamic human posture information; through this multi-branch structure, the model can fully utilize the complementary information provided by each branch, effectively cope with illumination changes, perspective changes, and occlusion challenges, thereby improving the accuracy and robustness of gait recognition.

[0014] The method of using the adaptive frame joint attention module in S2 to adaptively capture the important joint information of the gait sequence keyframes in the SMPL posture branch is as follows:

[0015] First, the gait feature map is pooled in the time frame dimension and joint dimension respectively to obtain frame-level gait features and joint-level gait features, which are then fed into the adaptive feature aggregation module to obtain adaptive frame attention and adaptive joint attention.

[0016] Then, the adaptive frame attention and the adaptive joint attention are multiplied together to obtain the adaptive frame-joint attention. Finally, the frame-level and joint-level attention are allocated to the original gait feature map, and a residual connection is performed, so that the model can effectively enhance the representation of key frames and important joint information in the gait posture sequence.

[0017] The adaptive feature aggregation module consists of two parts: adaptive convolution kernel generation and adaptive convolution. The adaptive convolution kernel is generated based on the compressed frame-level features and joint-level features. The local features are learned through the convolution layer, and the global features are learned in combination with the linear layer. Finally, the weight distribution is generated through the SoftMax layer to obtain the adaptive convolution kernel. The frame adaptive convolution kernel and joint adaptive convolution kernel obtained in the adaptive convolution process are applied to the frame dimension and joint dimension of the original gait feature map respectively, and the adaptive frame attention and adaptive joint attention are obtained through the convolution operation.

[0018] In S3, the method for achieving efficient fusion of two modal information by aligning and fusing SMPL model features and contour features in a unified semantic space is as follows:

[0019] The contour feature map f generated by the contour branch S Adaptive joint attention Att generated by SMPL pose branch in S2 J , respectively, through the mapping matrix T composed of two layers of linear transformation including Dropout layer and ReLU layer f and Tp Mapping is performed to complete the embedding alignment of two different modal features in the same semantic space:

[0020] f e =T f f S and j e =T p Att J

[0021] where f e is the contour feature after embedding alignment; j e is the SMPL pose feature after embedding alignment;

[0022] Finally, the two embedded features are sent to the cross-modal fusion layer to obtain the cross-modal fusion features of SMPL posture and contour:

[0023] F f =Linear(Linear(CBP(f e ,j e )))

[0024] Among them, CBP is compressed bilinear pooling, which performs high-order information interaction fusion on the features of the two embedded modalities to capture the complex relationship between the modalities; then, through two linear transformations, the high-dimensional fusion features are mapped to the low-dimensional space to further extract effective fusion features.

[0025] The joint loss function in S4 is:

[0026] L=αL tri +βL ce +γL MC

[0027] Among them, α, β and γ are L tri 、L ce and L MC The weighting parameter of tri is the triplet loss, L ce is the cross entropy loss, L MC is the modal consistency loss;

[0028] The modality consistency loss function is defined as the distance between the features of the two modalities after they are embedded in the same semantic space. It explicitly measures and optimizes the distribution consistency of the SMPL model features and the contour features in the embedding alignment, thereby effectively constraining the feature expression of the two modalities in the semantic space and reducing the information difference between the modalities:

[0029]

[0030] in, Indicates that the embedded features of the contour mode and SMPL model are normalized separately to eliminate the interference of feature scale on distance calculation; the constraint condition ‖T f ‖2=‖T p ‖2=1T f and T p Normalization is performed to avoid the occurrence of trivial solutions that are mapped to zero vectors or mapped to the same point.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The present invention decomposes the SMPL human body model and fully extracts the gait information contained in the SMPL model in terms of both dynamic posture and static shape. Branch optimization is performed through the adaptive frame joint attention module to achieve more refined spatiotemporal feature extraction of gait information. According to ablation experiments on the SMPL posture branch of the DFGait model and the GaitGraph model, the results show that the adaptive frame joint attention module alone can achieve improvements of 1.9% and 2.0% in the Rank-1 indicator, respectively. Through the modal fusion embedding module, efficient alignment and fusion of the contour modality and the SMPL model are achieved, improving the accuracy and robustness of gait recognition. According to ablation experiments on the DFGait model, the modal embedding fusion module improves the model's recognition accuracy by 1.5% in the Rank-1 indicator. The present invention performs well in various evaluation indicators in the real-scene gait dataset Gait3D, achieving 70.40% Rank-1, 85.00% Rank-5, 61.04% mAP and 41.27% mINP, achieving certain improvements in gait recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0034] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.

[0035] Figure 1 This is the overall architecture diagram of the DFGait model of the present invention;

[0036] Figure 2 It is a structural diagram of the contour branch in the DFGait model of the present invention;

[0037] Figure 3 A structural diagram of the decomposition of the SMPL model into the SMPL posture branch and the SMPL shape branch in the DFGait model of the present invention;

[0038] Figure 4 : This is a structural diagram of the adaptive frame joint attention module in the DFGait model of the present invention;

[0039] Figure 5 Schematic diagram and structural diagram of the modal embedding fusion module in the DFGait model of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of this application, not all the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than to limit the claims of the present invention. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0041] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0042] The present invention is implemented under the pytorch deep learning framework. The present invention provides a gait recognition method based on SMPL modal decomposition and embedding fusion, which specifically includes the following steps:

[0043] 1. Data preparation

[0044] The data sample of the present invention is the Gait3D gait dataset:

[0045] The Gait3D dataset is the first large-scale dataset for gait recognition in real-world scenarios based on three-dimensional representations. Captured with 39 cameras in a large supermarket, people often walk in irregular routes and speeds, with changes in viewpoint and speed, which aligns with gait characteristics found in real-world scenarios. The Gait3D dataset contains 4,000 subjects and 25,309 video sequences. The training set consists of 3,000 subjects, and the test set consists of 1,000 subjects. For the test set, one sequence from each subject is randomly selected to form the validation set, and the remaining sequences form the registration set. Details of the three subsets are as follows: the training set consists of 3,000 subjects, 18,940 sequences, and 2,459,545 frames; the validation set consists of 1,000 subjects, 1,000 sequences, and 123,425 frames; and the registration set consists of 1,000 subjects, 5,369 sequences, and 696,269 frames.

[0046] During training, the silhouette data is cropped to a size of 64×44.

[0047] 2. Model construction

[0048] The main body of the constructed DFGait model is a three-branch network structure. The overall network model structure is as follows Figure 1 As shown, the gait silhouette and SMPL model are fed into the network. Specifically, the SMPL human body model is decomposed into two independent branches: SMPL posture and SMPL shape, respectively, to extract gait features. A modality embedding fusion module is then used to align and fuse features from the silhouette and SMPL model. Finally, the features from the silhouette, SMPL, and the fused modality are concatenated for feature mapping.

[0049] The detailed structure of the outline branch is as follows Figure 2 As shown in the figure, a shallow feature extraction is performed by an initial 3×3 convolution layer, and then deep feature learning is performed through four layers of spatial convolution and temporal convolution. The spatial convolution is a residual structure composed of two 3×3 convolution layers and a batch normalization layer, and the temporal convolution is a residual structure composed of a 3×1×1 convolution layer and a batch normalization layer. Finally, the gait profile feature F is obtained through temporal pooling TP and horizontal pyramid mapping HPP operations. sils .

[0050] The dual-branch network of SMPL is as follows Figure 3 The SMPLShape branch uses an MLP network for shape feature extraction. The MLP network consists of three linear layers, each of which is followed by batch normalization and RuLu activation functions to enhance the network's nonlinear expression capabilities. To prevent overfitting, dropout layers are added to the last two layers to further improve the model's generalization capabilities.

[0051] The SMPLPose branch uses ResGCN for posture feature extraction. It first normalizes the input through a batch normalization layer. Then, it performs deep spatiotemporal feature extraction of gait posture through seven spatiotemporal graph convolution layers, including one spatiotemporal graph convolution base layer and six bottleneck layers. Each spatiotemporal graph convolution base layer consists of a spatial graph convolution layer and an L×1 temporal convolution layer, and includes a batch normalization layer and a ReLU activation function. To reduce the number of parameters and computational complexity in the network, 1×1 convolution layers are inserted before and after the convolution layers to form bottleneck layers.

[0052] In particular, Figure 4 As shown in the figure, an adaptive frame-joint attention module is deployed after each spatiotemporal graph convolution layer. Specifically, the gait feature map output by the spatiotemporal graph convolution layer is pooled in the temporal frame dimension and the joint dimension to obtain frame-level gait features and joint-level gait features. These features are then fed into the adaptive feature aggregation module: a convolutional layer performs local feature learning, two linear layers perform global feature learning, and a softmax layer produces a frame-adaptive convolution kernel and a joint-adaptive convolution kernel. These two convolution kernels are applied to the frame dimension and the joint dimension of the gait feature map, respectively, to obtain adaptive frame attention and adaptive joint attention through convolution operations. Finally, the adaptive frame-joint attention is multiplied together and applied to the output feature map of the spatiotemporal graph convolution layer, achieving precise attention allocation to frame-level and joint-level features while performing residual connections.

[0053] Then the SMPL shape features and posture features are fused, and the dynamic posture features F P and the shape feature F obtained by the SMPL shape branch s Apply the maximum pooling operation separately to retain the most significant information in each feature, and then add the two pooled feature vectors to achieve the fusion of posture and shape features:

[0054] F SMPL =MaxPool2d(F P )+MaxPool1d(F S )

[0055] Among them, F SMPL That is, the gait feature of the SMPL model after the fusion of the SMPL posture branch and the SMPL shape branch.

[0056] At the same time, if Figure 5 As shown in Figure 2, the contour features and SMPL posture features are fused through the modality embedding fusion module:

[0057] F fusion =EmbeddingFusion(f S , Att J)

[0058] Among them, F fusion represents the fusion feature of contour and SMPL posture; EmbeddingFusion represents the modality embedding fusion module; f S It is the gait contour feature obtained after four layers of spatiotemporal convolution feature extraction in the contour branch; Att J It is the adaptive joint attention generated in the last layer of the adaptive frame joint attention module in the pose branch of SMPL.

[0059] Finally, the gait features F from the silhouette branch are sils , gait features F from the SMPL model SMPL And the feature F after modal fusion fusion Splicing is performed in the feature dimension to obtain the final output feature F out The feature vector is then mapped to the metric space through a separate fully connected layer, and the feature space is further adjusted using BNNeck.

[0060] 3. Model training

[0061] The network of the present invention is jointly trained by three loss functions, and the weight parameters of the model are iteratively updated through the back propagation algorithm and optimizer. The three loss functions are defined as follows:

[0062] Triplet loss L tri :

[0063]

[0064] Among them, N tri Indicates the total number of triplets contained in a batch; and They represent the anchor sample gait features of the i-th triplet in the batch, the positive sample gait features with the same identity as the anchor sample, and the negative sample gait features with different identities from the anchor sample; margin represents the margin;

[0065] Cross entropy loss L ce :

[0066]

[0067] Among them, N represents the number of samples in a batch; class represents the number of categories; y k is the true label vector of pedestrian sample k; is the prediction vector of pedestrian sample k; is the label vector y k The elements in represent the true label of pedestrian k belonging to the i-th category; is the prediction vector The value in represents the probability that the model predicts that pedestrian k belongs to the i-th category.

[0068] Modal consistency loss L MC :

[0069]

[0070] in, Indicates that the embedded features of the contour mode and SMPL model are normalized respectively to eliminate the interference of feature scale on distance calculation; the constraint condition ‖T f ‖2=‖T p ‖2=1T f and T p Normalization is performed to avoid the occurrence of trivial solutions that are mapped to zero vectors or mapped to the same point.

[0071] The joint loss function L is:

[0072] L=αL tri +βL ce +γL MC

[0073] where α, β and β are L tri 、L ce and L MC The weighting parameters of .

[0074] 4. Test results

[0075] In this example, the test set is divided into a registration set and a validation set. The gait sequences of the registration set and the validation set are input into a multimodal gait recognition network model based on SMPL modal decomposition and embedding fusion to obtain gait features. The pedestrian identity is determined by calculating the cosine similarity between the features of the validation set and the registration set.

[0076] 5. Model Evaluation

[0077] In order to evaluate the multimodal gait recognition method based on SMPL modal decomposition and embedding fusion proposed in the present invention, the present invention is systematically compared with 5 contour-based methods, 2 skeleton-based methods, and 5 multimodal methods based on contour and skeleton or SMPL models on the Gait3D dataset. The Rank-1 accuracy, Rank-5 accuracy, mean average precision (mAP), and mean inverse negative penalty (mINP) evaluation indicators are used to evaluate the performance of the model. Table 1 lists the comparison results of the present invention and other 12 advanced gait recognition methods on the Gait3D dataset in real scenes.

[0078] Table 1 Comparison results of different methods on Gait3D dataset

[0079]

[0080] As can be seen from Table 1, the proposed method achieved the best results in all evaluation indicators, and significantly outperformed other comparison methods in the key indicators of Rank-1 accuracy, Rank-5 accuracy, mAP, and mINP. This shows that the proposed method has demonstrated excellent performance in gait recognition tasks, can effectively capture the fine-grained features of gait, and further improves the robustness and accuracy of the model in complex situations facing real-world scenarios through modal fusion.

[0081] The above only describes in detail the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the purpose of the present invention, and various changes should be included in the scope of protection of the present invention.

Claims

1. A multimodal gait recognition method based on SMPL modal decomposition and embedding fusion, characterized by: The following steps are involved: S1. Construct a multimodal gait recognition model DFGait that integrates the SMPL model and the contour; S2. We propose an adaptive frame joint attention module to adaptively extract important joint information of gait sequence keyframes in the SMPL posture branch. S3. We propose a modality embedding fusion module to achieve efficient fusion of two modal information by aligning and fusing SMPL model features and contour features in a unified semantic space. The contour feature map f generated by the contour branch S Adaptive joint attention Att generated by SMPL pose branch in S2 J , respectively, through the mapping matrix T composed of two layers of linear transformation including Dropout layer and ReLU layer f and T p Mapping is performed to complete the embedding alignment of two different modal features in the same semantic space: f e =T f f S and j e =T p Att J where f e is the contour feature after embedding alignment; j e is the SMPL pose feature after embedding alignment; Finally, the two embedded features are sent to the cross-modal fusion layer to obtain the cross-modal fusion features of SMPL posture and contour: F f =Linear(Linear(CBP(f e ,j e ))) Among them, CBP is compressed bilinear pooling, which fuses the features of the two embedded modalities with high-order information interaction to capture the complex relationship between the modalities; then, through two linear transformations, the high-dimensional fusion features are mapped to a low-dimensional space to further extract effective fusion features; S4. A joint loss function combining triplet loss, cross entropy loss, and modality consistency loss is used to jointly supervise the training of the DFGait model.

2. The multimodal gait recognition method based on SMPL modal decomposition and embedding fusion according to claim 1 is characterized in that: The method for constructing the multimodal gait recognition model DFGait that integrates the SMPL model and the contour in S1 is as follows: decomposing the SMPL model into two independent branches, SMPL shape and SMPL posture, focusing on the extraction of SMPL shape features and SMPL posture features respectively; Therefore, the model constructs three branches: the contour branch is responsible for capturing the appearance and motion features in the contour sequence, the SMPL shape branch is responsible for extracting the shape features of the SMPL human body model, and the SMPL posture branch focuses on extracting dynamic human posture information; through this multi-branch structure, the model can make full use of the complementary information provided by each branch, effectively cope with illumination changes, perspective changes, and occlusion challenges, thereby improving the accuracy and robustness of gait recognition.

3. The multimodal gait recognition method based on SMPL modal decomposition and embedding fusion according to claim 1 is characterized in that: The method of using the adaptive frame joint attention module in S2 to adaptively capture the important joint information of the gait sequence keyframes in the SMPL posture branch is as follows: First, the gait feature map is pooled in the time frame dimension and joint dimension respectively to obtain frame-level gait features and joint-level gait features, which are then fed into the adaptive feature aggregation module to obtain adaptive frame attention and adaptive joint attention. Then multiply the adaptive frame attention and the adaptive joint attention to get the adaptive frame joint attention; Finally, attention is allocated at the frame level and joint level for the original gait feature map, and residual connections are performed, so that the model can effectively enhance the representation of key frames and important joint information in the gait posture sequence.

4. The multimodal gait recognition method based on SMPL modal decomposition and embedding fusion according to claim 3 is characterized in that: The adaptive feature aggregation module consists of two parts: adaptive convolution kernel generation and adaptive convolution. The adaptive convolution kernel is generated based on the compressed frame-level features and joint-level features. The local features are learned through the convolution layer, and the global features are learned in combination with the linear layer. Finally, the weight distribution is generated through the SoftMax layer to obtain the adaptive convolution kernel. The frame adaptive convolution kernel and joint adaptive convolution kernel obtained in the adaptive convolution process are applied to the frame dimension and joint dimension of the original gait feature map respectively, and the adaptive frame attention and adaptive joint attention are obtained through the convolution operation.

5. The multimodal gait recognition method based on SMPL modal decomposition and embedding fusion according to claim 1 is characterized in that: The joint loss function in S4 is: L=αL tri +βL ce +γL MC Among them, α, β and γ are L tri 、L ce and L MC The weighting parameter of tri is the triplet loss, L ce is the cross entropy loss, L MC is the modal consistency loss; The modality consistency loss function is defined as the distance between the features of the two modalities after they are embedded in the same semantic space. It explicitly measures and optimizes the distribution consistency of the SMPL model features and the contour features in the embedding alignment, thereby effectively constraining the feature expression of the two modalities in the semantic space and reducing the information difference between the modalities: in, Indicates that the embedded features of the contour mode and SMPL model are normalized respectively to eliminate the interference of feature scale on distance calculation; the constraint condition ‖T f ‖2=‖T p ‖2=1T f and T p Normalize, T f and T p Represents the mapping matrix consisting of two layers of linear transformation including the Dropout layer and the ReLU layer, avoiding the appearance of trivial solutions that are mapped to the zero vector or mapped to the same point.

Citation Information

Patent Citations

  • Gait recognition method based on three-dimensional human body modeling point cloud feature coding

    CN114973422A