A Gait Recognition Method Integrating Depth Maps Derived from RGB Images and Contour Sequences

By integrating RGB image-derived depth maps and silhouette sequences with a cross-level, multi-scale feature fusion network, the method enhances step recognition accuracy by capturing fine-grained spatial information and improving robustness against viewpoint and clothing changes.

CN119888867BActive Publication Date: 2025-07-15GUANGDONG ZHIYUN URBAN CONSTR TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369893.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

Existing step recognition methods face challenges in real-world scenarios due to insufficient utilization of three-dimensional information, particularly in fusing and integrating different modalities like SMPL parameters, bone structure graphs, and depth information, leading to suboptimal performance under varying viewpoints and clothing changes.

Method used

A method that combines RGB image-derived depth maps and silhouette sequences for step recognition, employing a cross-level, multi-scale feature fusion network with attention mechanisms to enhance the interaction between modalities, capturing fine-grained spatial information and improving robustness.

Benefits of technology

The method effectively leverages 3D geometric information from depth maps to enhance step recognition accuracy, addressing the limitations of 2D modalities and improving performance under varying viewpoints and clothing changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888867B_ABST
    Figure CN119888867B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of gait recognition, and provides a gait recognition method that fuses depth maps and contour sequences derived from RGB images. The present invention combines depth map sequences based on RGB images and traditional contour map sequences; uses existing RGB video datasets and the latest depth estimation model to estimate depth maps shown in a given RGB image sequence, and uses them as a new modality to capture distinctive features inherent in human motion; compared with traditional input modalities, depth maps provide more explicit 3D geometric information about the human body and its motion, enriching gait representation; the present invention provides a more comprehensive and accurate gait recognition solution; and achieves good recognition results on widely used gait recognition benchmarks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gait recognition, and particularly relates to a gait recognition method that fuses depth maps and contour sequences derived from RGB images. Background Art

[0002] Gait recognition is an authentication technology that confirms an individual's identity by analyzing their unique movement patterns while walking. It is commonly used in the field of public safety, such as criminal investigations and suspect tracking. Compared with traditional biometric technologies such as face, fingerprint, and iris recognition, gait recognition has the advantages of non-contact, high privacy, difficulty in forgery, and the ability to achieve long-distance recognition, and thus is widely used in various scenarios.

[0003] Currently, there are two main input modal forms for gait recognition methods, namely contour sequences and skeleton sequences. Contour sequences distinguish individuals by explicitly retaining appearance information, while skeleton sequences retain the internal structure information of the human body. When the appearance changes significantly, the skeleton sequence remains robust. However, both of these modalities have certain limitations. Contours are easily affected by significant changes in the external body shape caused by clothing changes, while although the skeleton is effective in solving clothing occlusion, it completely ignores the highly discriminative body shape information, resulting in poor performance. At the same time, to address the limitations of single modalities, recent research has also explored the possibility of multimodal fusion, thereby improving performance. However, existing solutions still cannot effectively solve the complex problems of real-world scenarios.

[0004] When the research scenario shifts from traditional laboratory settings to real-world scenarios, traditional gait recognition methods are no longer applicable. In recent years, researchers have gradually attempted to use new gait input modalities to achieve robust gait recognition results. The SMPLGait model proposes using features extracted from SMPL parameters as the 3D representation of the human body. It uses a dual-branch structure based on a deep neural network. One branch learns appearance features from the human silhouette, and the other branch learns 3D viewpoint and shape knowledge from 3D SMPL parameters, thus leveraging the 3D geometric information of the SMPL model to enhance gait appearance feature learning. However, SMPLGait simply concatenates the final global features of the two modalities and cannot effectively capture fine-grained spatial information. How to effectively fuse their features and capture and integrate the complex relationships between different gait modalities remains a problem. At the same time, although the SMPL model is usually regarded as a dense mesh, its feature vectors have only dozens of dimensions, showing a relatively sparse representation of body shape and posture, and the description of gait patterns is not fine-grained enough. The SkeletonGait++ model innovatively introduces the concept of a skeleton graph. Different from traditional skeleton modalities, the skeleton graph is not just a simple representation of joints but generates a heatmap for each joint point through Gaussian approximation, presenting a more intuitive structure. At the same time, the skeleton graph is closer to traditional image modalities in terms of data format, enabling the skeleton graph to better combine the advantages of the contour graph modality and having stronger compatibility in multi-modal fusion for gait recognition. Despite good performance, the skeleton graph modality still only contains 2D gait information and performs poorly when facing challenges such as viewpoint changes. The LidarGait model first proposed using 3D point clouds captured by LIDAR sensors for gait recognition. By projecting sparse point cloud data into a depth map and then combining a deep neural network to extract the fine-grained features required for gait recognition from 3D geometric information. However, the mainstream sensors for scenarios where gait recognition is currently required are still mainly low-cost cameras, and expensive lidar has not yet been popularized. Moreover, the current gait datasets are mainly RGB video datasets, and point cloud datasets are relatively scarce.

[0005] In addition, with the development of depth estimation technology, the depth information estimated from RGB images is becoming increasingly accurate. Compared with traditional input modalities, the depth map provides more explicit 3D geometric information about the human body and its movement, which cannot be obtained in contour and skeleton input modalities. Compared with 2D planar information, this additional dimension enriches the expression of gait features. Secondly, the depth map can more accurately analyze subtle gait movement changes, which is crucial for capturing the unique features of an individual's gait. In addition, depth information can help address the challenge of viewpoint changes because it provides a more consistent expression of the body structure information and movement information at different angles.

[0006] Although the existing gait recognition methods based on the fused contour sequence and SMPL parameter sequence can complete the personnel recognition in complex outdoor scenarios, they do not fully exploit the potential of 3D information in gait recognition. After replacing with a more powerful backbone network, the 3D branch will not significantly enhance the overall gait recognition performance. This is because in terms of feature fusion, SMPLGait only uses simple element-wise multiplication and addition operations, which cannot effectively bridge the gap between the two modal features. Moreover, the feature vector of the SMPL model has only dozens of dimensions, the representation of body shape and posture is relatively sparse, and the description of gait features is not fine-grained enough.

[0007] Although the existing gait recognition methods based on the skeleton graph sequence and contour graph sequence innovatively introduce the concept of the skeleton graph, generating a heat map for each joint point through Gaussian approximation, making the skeleton structure more intuitive, and retaining the data format of the traditional image modality. However, the skeleton graph modality only contains 2D planar structure information, which is not sufficient to express more subtle gait movement changes and still performs poorly in the face of challenges such as viewpoint changes. At the same time, SkeletonGait++ only uses a simple attention fusion operation when performing feature fusion, which is not sufficient to effectively capture and integrate the complex relationships between modalities.

[0008] The existing gait recognition methods based on 3D radar point clouds project the sparse point cloud data into a depth map and then extract the fine-grained features required for gait recognition from the 3D geometric information in combination with a deep neural network. Although good results have been achieved, currently, the main sensors used in the main scenarios where gait recognition is required are still low-cost cameras. Moreover, currently, the datasets are mainly RGB video datasets, and the point cloud datasets are relatively few.

[0009] In summary, although the gait recognition methods based on new input modalities such as SMPL parameters and skeleton graphs show certain potential, when solving practical scenario problems, various factors still need to be comprehensively considered, and relevant technologies and methods need to be continuously improved. Especially in terms of how to construct an input modality that can better contain rich gait information, so as to extract more discriminative gait features, deeply explore the potential of depth information in the gait recognition task, and how to effectively fuse the information of two different modalities of depth maps and contour graphs to obtain a more powerful gait representation, thereby improving the accuracy of gait recognition. Summary of the Invention

[0010] To solve the above technical problems, the present invention provides a gait recognition method that fuses the depth map derived from RGB images and the contour sequence to solve the problems in the prior art. The technical solution adopted by the present invention is:

[0011] A gait recognition method that fuses depth maps and contour sequences derived from RGB images, comprising the following steps:

[0012] S1: Export a contour map sequence and a depth map sequence from the original video clips in the public dataset;

[0013] S2: Perform cropping and normalization operations on the contour map sequence and the depth map sequence;

[0014] S3: Feed the obtained contour map sequence and depth map sequence into the contour feature extractor and the depth map feature extractor in the feature extraction module respectively; extract features from the contour map sequence and the depth map sequence respectively to obtain a contour feature map and a depth feature map , representing the encoding stage;

[0015] S4: Perform feature fusion through a cross-level, multi-scale fusion module;

[0016] S5: Feed the fused features into the feature aggregation module, and perform feature aggregation through temporal pooling and horizontal pyramid pooling to generate gait recognition features;

[0017] S6: Use the features obtained in S5 for training and inference;

[0018] S7: Finally, during inference, by comparing the cosine similarity of the features of the probe set and the gallery set, the most similar one is regarded as the prediction object.

[0019] Furthermore, step S1 includes:

[0020] S11: Through the pedestrian segmentation algorithm, set the background to black and the human contour to white. The contour sequence is denoted as S, with the size of , representing the number of channels, representing the length of the contour sequence, representing the height of each frame of the image, representing the width of each frame of the image;

[0021] S12: Adopt the Depth Anything basic model to estimate the depth map from the RGB image.

[0022] Furthermore, step S2 includes:

[0023] S21: Determine the positions of the top and bottom of the non-zero elements in the contour map sequence and the depth map sequence, and crop the images to remove the background;

[0024] S22: Set the contour input height to 64 pixels, and adjust the widths of the contour map sequence and the depth map sequence accordingly according to the human body aspect ratio, keeping the human body aspect ratio unchanged;

[0025] S23: Calculate the total number of pixels in the contour image, then calculate the cumulative number of pixels in each column, determine the position where the cumulative number of pixels exceeds half of the total number of pixels, and designate it as the vertical center of the image; The center of the contour image is set as the center of the depth image; The depth map sequence is represented as D, and the size is , where , , , represent the number of channels, sequence length, image frame height, and image frame width of the depth image respectively.

[0026] Further, step S4 includes:

[0027] S41: The multi-scale spatial extraction module connects the contour feature map and the channel feature map through the channel dimension to obtain a unified feature tensor , and the formula is as follows:

[0028] ;

[0029] S42: Perform multi-scale spatial extraction on the concatenated feature vector, and the formula is as follows:

[0030] ;

[0031] where , represents a 1×1 convolutional kernel; is the local score, is the global score; is the relu activation function, is the BatchNorm2d batch normalization layer;

[0032] S43: Calculate the attention weights by adding and through attention fusion, and the formula is as follows:

[0033] ;

[0034] where represents the activation function;

[0035] S44: Use cross-level fusion to obtain the final gait representation , while obtaining higher-level semantic information and retaining the spatial information contained in the shallow feature map, and the formula is as follows:

[0036] ;

[0037] wherein represents the features extracted at each encoding stage before the cross - level and multi - scale fusion module; the feature fusion process of steps S41 - S44 is applied at each stage of encoding.

[0038] Furthermore, step S5 includes:

[0039] S51: The fused features undergo a temporal pooling operation to aggregate the sequence of feature maps by maximizing along the temporal dimension, and the global understanding is output.

[0040] S52: The features after temporal pooling are further subjected to horizontal pyramid pooling to obtain multi - scale information of the input features. The number of blocks in the horizontal pyramid is specified as 16. The input features are horizontally segmented in the width dimension, and then average pooling and max - pooling operations are performed on each block. The results of the two pooling operations are concatenated along the last dimension to form the final output features.

[0041] Furthermore, step S6 includes:

[0042] S61: The output features obtained in S5 are linearly mapped through 16 independent fully - connected layers. Each fully - connected layer maps the input features to a new feature space with an output channel number of 256. The output features after this linear transformation are used to train the triplet loss.

[0043] S62: The features obtained in S61 are batch - normalized. The obtained features are divided into multiple parts, and each part is batch - normalized. The finally obtained features are used to train the cross - entropy loss and for inference.

[0044] S63: During the training process, the weighted sum of the triplet loss and the cross - entropy loss is used as the loss function. The triplet loss function has the following formula:

[0045] ;

[0046] wherein, represents the set of positive sample pairs, represents the set of negative sample pairs, represents the distance between the anchor and the positive sample, represents the distance between the anchor and the negative sample pair, represents the margin value set to 0.2;

[0047] The cross - entropy loss function has the following formula:

[0048] ;

[0049] Among them, is the number of categories, is the one-hot encoding of the true category label, is the category probability predicted by the model;

[0050] The formula for the total loss function is as follows:

[0051] ;

[0052] Among them is the triplet loss, is the cross-entropy loss, and are the weighting parameters;

[0053] S64: Measure the similarity by calculating the cosine similarity between the query feature vector and the gallery feature vector.

[0054] The present invention has the following beneficial effects:

[0055] The present invention integrates the depth map sequence based on RGB images and the traditional contour map sequence. The estimated depth map displayed from the given RGB image sequence is obtained by using the existing RGB video dataset and the latest depth estimation model, and is used as a new modality to capture the distinctive features inherent in human motion. Compared with the traditional input modality, the depth map provides more explicit 3D geometric information about the human body and its motion, enriching the gait representation. Secondly, the depth map can analyze the subtle gait motion changes more accurately, which is crucial for capturing the unique features of an individual's gait. In addition, the depth information can help solve the challenge of viewpoint change, because it provides a more consistent expression of the body structure information and motion information at different angles.

[0056] At the same time, in order to facilitate the feature fusion of the two modalities, the present invention also proposes a new cross-level, multi-stage, multi-scale feature fusion network. Through cross-level, multi-scale, multi-stage attention fusion, the interaction between modalities is enhanced, and the gait information between different modalities is captured in a finer granularity, so as to better bridge the gap between the two modalities and achieve more robust gait recognition, thereby providing a more comprehensive and accurate gait recognition solution. And good recognition results are obtained on the widely used gait recognition benchmarks. Description of the Drawings

[0057] Figure 1 is the overall flowchart of the present invention;

[0058] Figure 2 is the flowchart of the depth map and contour map cropping and normalization;

[0059] Figure 3 It is a flowchart of feature fusion. Specific implementation manners

[0060] Next, in combination with the Figures 1 - 3 in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. If not specifically specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0061] The present invention fuses a depth map sequence based on RGB images and a traditional contour map sequence. An estimated depth map is displayed from a given RGB image sequence by using an existing RGB video dataset and the latest depth estimation model, and it is used as a new modality to capture the distinctive features inherent in human motion. Compared with traditional input modalities, the depth map provides more explicit 3D geometric information about the human body and its motion, enriching gait representation. Secondly, the depth map can more accurately analyze subtle gait motion changes, which is crucial for capturing the unique features of an individual's gait. In addition, depth information can help address the challenge of viewpoint changes because it provides a more consistent expression of body structure information and motion information from different angles. At the same time, in order to facilitate the feature fusion of the two modalities, the present invention also proposes a new cross-level, multi-stage, multi-scale feature fusion network. Through cross-level, multi-scale, multi-stage attention fusion, the interaction between modalities is enhanced, and gait information between different modalities is captured in a more fine-grained manner, thereby better bridging the gap between the two modalities and achieving more robust gait recognition, thus providing a more comprehensive and accurate gait recognition solution. And good recognition results are obtained on widely used gait recognition benchmarks.

[0062] The present invention proposes a gait recognition method that fuses a depth map derived from RGB images and a contour sequence, and the flowchart is as Figure 1 shown. First, in the data preprocessing stage, a depth map sequence and a contour map sequence are derived from the sampled frames of an RGB motion video clip. Then, these sequences are fed into the corresponding feature extractors in the encoding module to generate corresponding feature maps. In the cross-level multi-scale fusion module, a multi-scale spatial feature extraction module is designed to extract multi-scale features from the feature maps. The obtained multi-scale features are fused with the original feature maps in the previous module through attention, and a cross-level fusion module is used to generate the final gait representation. Finally, feature aggregation is performed through a pooling operation, and the model is trained using cross-entropy loss and triplet loss, with cosine similarity as the metric for evaluating gait similarity. Specifically, it includes:

[0063] S1: In the preprocessing stage, a sequence of contour maps and a sequence of depth maps are derived from the original video clips in the publicly available dataset. The sequence of contour maps includes multiple contour maps, and the sequence of depth maps includes multiple depth maps. The publicly available dataset can be obtained from the Internet.

[0064] S11: Generally, in existing gait recognition datasets, there is usually a directly processed sequence of contour maps. Through a classical pedestrian segmentation algorithm, the background is set to black and the human contour is set to white. The contour sequence is denoted as S, with dimensions . denotes the number of channels, denotes the length of the contour sequence, denotes the height of each frame image, denotes the width of each frame image.

[0065] S12: The Depth Anything base model of the existing technology is used to estimate the depth map from the RGB image. The RGB image is the video frame in the original dataset. Specifically, to exclude the interference caused by the background factors, we use the contour as a mask to generate a depth map with only the region of interest. Other depth estimation models can also be used. Then, using the contour as a mask, the background is completely set to black to exclude the interference caused by the background factors.

[0066] S2: At the same time, in order to make the obtained sequence of contour maps and sequence of depth maps contain richer and more accurate gait information, cropping and normalization operations are performed on the sequence of contour maps and the sequence of depth maps. The process is as Figure 2 shown.

[0067] S21: First, we determine the positions of the top and bottom of the non-zero elements in the sequence of contour maps and the sequence of depth maps, and crop the images to delete the irrelevant background.

[0068] S22: Then, following the standard practice of gait recognition, the input height of the contour is set to 64 pixels, and the widths of the sequence of contour maps and the sequence of depth maps are adjusted accordingly according to the aspect ratio of the human body, keeping the aspect ratio of the human body unchanged.

[0069] The above process can ensure that the horizontal center is located at half of the image height, thus providing a consistent reference frame for all samples. When looking for the vertical center, since the vertical center of the contour image is better determined, the vertical center of the contour image is used as the vertical center of the depth image. First, calculate the total number of pixels in the contour image, and then calculate the cumulative number of pixels in each column from left to right. By traversing this cumulative pixel list, determine the position where the cumulative number of pixels exceeds half of the total pixels, and designate it as the vertical center of the image. The center of the found contour image is also set as the center of the depth image. The sequence of depth maps is denoted as D, with dimensions . Among them, , , , respectively represent the number of channels, sequence length, height of the image frame, and width of the image frame of the depth map.

[0070] S3: Then, the obtained contour map sequence S and depth map sequence D are respectively fed into the contour feature extractor and depth map feature extractor in the feature extraction module. We use the feature extractor introduced in DeepGaitV2 as the encoder because it has achieved optimal performance on common gait recognition datasets. Feature maps are obtained by performing feature extraction from S and D respectively and , represent the encoding stage.

[0071] S4: After obtaining the feature maps of the two modalities, feature fusion is performed through a cross-level and multi-scale fusion module. The process is as Figure 3 shown.

[0072] S41: First, the multi-scale spatial extraction module obtains a unified feature tensor by concatenating the contour feature map and the channel feature map along the channel dimension. The formula is as follows:

[0073] ;

[0074] This concatenation enables subsequent convolution operations to process these features simultaneously, thereby leveraging complementary information from the contour and depth modalities.

[0075] S42: Then, multi-scale spatial extraction is performed on the concatenated feature vector. The formula is as follows:

[0076] ;

[0077] Among them, , represents a 1×1 convolutional kernel. is the local score, is the global score. is the relu activation function, is the BatchNorm2d batch normalization layer; different sizes of convolutional kernels are used at this stage to capture both fine-grained features and global information simultaneously. By using this multi-scale method, the comprehensiveness of feature expression is strengthened to ensure that local details and global context information are retained. Thus, the local score and the global score are obtained.

[0078] S43: Finally, through attention fusion, and The attention weights are calculated by addition. Through the attention weights, the feature maps and are fused to obtain the fused output at this stage , enabling the model to selectively focus on the important parts of the input, thereby improving the effectiveness of the gait fusion features. The formula is as follows:

[0079] ;

[0080] Among them, represents the activation function.

[0081] S44: To enhance the feature expression ability, cross-level fusion is used to obtain the final gait representation . While obtaining higher-level semantic information, the rich spatial information contained in the shallow feature maps is retained. The formula is as follows:

[0082] ;

[0083] Among them represents the features extracted at each encoding stage before the cross-level and multi-scale fusion module. At the same time, to solve the problem of insufficient information interaction between modalities, the feature fusion process of steps S41 - S44 is applied at each stage of encoding. By fusing the feature information of different depths layer by layer, the feature representation is gradually optimized, thereby improving the overall performance of the model.

[0084] S5: The fused final features are fed into the feature aggregation module, and feature aggregation is performed through temporal pooling and horizontal pyramid pooling to generate the final gait recognition features.

[0085] S51: The fused features first undergo a temporal pooling operation, and the feature map sequence is aggregated by maximizing along the temporal dimension to output a global understanding.

[0086] S52: The features after temporal pooling are then subjected to horizontal pyramid pooling to obtain multi-scale information of the input features. The number of blocks in the horizontal pyramid is specified as 16, and horizontal segmentation is performed on the width dimension of the input features. Then, average pooling and max pooling operations are performed on each block. By combining average pooling and max pooling, both the details of the local features (max pooling) and the statistical information of the global features (average pooling) are retained. Finally, the results of the two pooling operations are concatenated along the last dimension to form the final output features.

[0087] S6: The features obtained in S5 are used for training and inference.

[0088] S61: Linearly map the output features obtained in S5 through 16 independent fully-connected layers. Each fully-connected layer maps the input features to a new feature space with an output channel number of 256. Use the output features after this linear transformation to train the triplet loss.

[0089] S62: Batch-normalize the features obtained in S61. Divide the obtained features into multiple parts, perform batch normalization on each part, and finally use the obtained features to train the cross-entropy loss and for inference.

[0090] S63: During the training process, use the weighted sum of the triplet loss and the cross-entropy loss as the loss function. The triplet loss function has the following formula:

[0091] ;

[0092] where represents the set of positive sample pairs, represents the set of negative sample pairs. represents the distance between the anchor and the positive sample, represents the distance between the anchor and the negative sample pair, represents the margin value set to 0.2.

[0093] The cross-entropy loss function has the following formula:

[0094] ;

[0095] where is the number of classes, that is, the number of identity IDs in the dataset, is the one-hot encoding of the true class label, is the class probability predicted by the model, obtained through softmax calculation.

[0096] The formula for the total loss function is as follows:

[0097] ;

[0098] where is the triplet loss, is the cross-entropy loss, and are the weighting parameters.

[0099] S64: Measure the similarity by calculating the cosine similarity between the query feature vector and the gallery feature vector. Cosine similarity: used to measure how similar the directions of these two vectors are. The closer the value is to 1, the more similar they are; the closer it is to 0, the more different they are.

[0100] S7: When finally making inferences, by comparing the cosine similarities of the features between the probe set and the gallery set, the most similar one is regarded as the prediction object; for example, when obtaining a photo or a video of a person walking, the image features of the person are extracted, compared with the image features in the gallery set, and the cosine similarity is calculated. The similar one is regarded as the prediction object.

[0101] The embodiments described above are only descriptions of the preferred modes of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations, variations, modifications, and replacements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A gait recognition method that fuses depth maps and contour sequences derived from RGB images, characterized in that It includes the following steps: S1: Export a sequence of contour maps and a sequence of depth maps from the original video clips in the public dataset; S2: Perform cropping and normalization operations on the sequence of contour maps and the sequence of depth maps; S3: Feed the obtained sequence of contour maps and sequence of depth maps into the contour feature extractor and the depth map feature extractor in the feature extraction module respectively; Feature extraction is respectively performed on the contour map sequence and the depth map sequence to obtain a contour feature map and a depth feature map , representing the encoding stage; S4: Perform feature fusion through a cross-level and multi-scale fusion module; S5: Feed the fused features into the feature aggregation module, and perform feature aggregation through temporal pooling and horizontal pyramid pooling to generate gait recognition features; S6: Use the features obtained in S5 for training and inference; S7: Finally, during inference, by comparing the cosine similarities of the features of the probe set and the gallery set, the most similar one is regarded as the predicted object; Step S1 includes: S11: Through the pedestrian segmentation algorithm, set the background to black, set the human body contour to white, represent the contour sequence as S, and the size is , represents the number of channels, represents the length of the contour sequence, represents the height of each frame of image, represents the width of each frame of image; S12: Adopt the Depth Anything basic model to estimate the depth map from the RGB image; Step S2 includes: S21: Determine the positions of the top and bottom of the non-zero elements in the sequence of contour maps and the sequence of depth maps, crop the images, and remove the background; S22: Set the contour input height to 64 pixels, and adjust the widths of the sequence of contour maps and the sequence of depth maps accordingly according to the aspect ratio of the human body, keeping the aspect ratio of the human body unchanged; S23: Calculate the total number of pixels in the contour image, then calculate the cumulative number of pixels in each column, determine the position where the cumulative number of pixels exceeds half of the total number of pixels, and designate it as the vertical center of the image; set the center of the contour image as the center of the depth image; the depth map sequence is represented as D, with dimensions of , where , , , respectively represent the number of channels, sequence length, image frame height, and image frame width of the depth image; Step S4 includes: S41: The multi-scale spatial extraction module connects the contour feature map and the channel feature map through the channel dimension to obtain a unified feature vector, as shown in the following formula: , as shown in the following formula: ; S42: Perform multi-scale spatial extraction on the concatenated feature vectors, and the formula is as follows: ; Among them, , represents a 1×1 convolutional kernel; is the local score, is the global score; is the relu activation function, is the BatchNorm2d batch normalization layer; S43: Calculate the attention weights by adding and as follows: ; Among them, denotes activation function; S44: Obtain the final gait representation using cross-level fusion , while obtaining higher-level semantic information, retain the spatial information contained in the shallow feature maps. The formula is as follows: ; Among them represents the features extracted at each encoding stage before the cross-level and multi-scale fusion module; the feature fusion process of steps S41 - S44 is applied at each stage of encoding; Step S5 includes: S51: The fused features undergo a temporal pooling operation, and aggregate the sequence of feature maps by maximizing along the temporal dimension to output a global understanding; S52: Perform horizontal pyramid pooling on the features after temporal pooling to obtain multi-scale information of the input features. Specify the number of blocks in the horizontal pyramid as 16, perform horizontal segmentation on the width dimension of the input features, and then perform average pooling and max pooling operations on each block, and concatenate the results of the two pooling operations along the last dimension to form the final output features.

2. The gait recognition method for fusing a depth map and a contour sequence derived from an RGB image according to claim 1, wherein, Step S6 includes: S61: Linearly map the output features obtained in S5 through 16 independent fully connected layers. Each fully connected layer maps the input features to a new feature space, and the number of output channels is 256. Use this linearly transformed output feature to train the triplet loss; S62: Perform batch normalization on the features obtained in S61, divide the obtained features into multiple parts, perform batch normalization on each part, and finally use the obtained features to train the cross-entropy loss and perform inference; S63: During the training process, the weighted sum of the triplet loss and the cross-entropy loss is used as the loss function, and the formula of the triplet loss function is as follows: ; Among them, represents the set of positive sample pairs, represents the set of negative sample pairs, represents the distance between the anchor point and the positive sample, represents the distance between the anchor point and the negative sample pair, represents that the boundary value is set to 0.2; Cross-entropy loss function The formula is as follows: ; Among them, is the number of categories, is the one-hot encoding of the true category label, is the category probability predicted by the model; The formula of the total loss function is as follows: ; Among them is the triplet loss, is the cross-entropy loss, and is the weighting parameter; S64: Measure the similarity by calculating the cosine similarity between the query feature vector and the gallery feature vector.

Citation Information

Patent Citations

  • Scene recognition method based on multi-modal features and graph attention mechanism

    CN118486026A

  • Behavior recognition method, system and equipment for passenger station group and medium

    CN119251773A