A gait feature extraction method, device, and computer-readable storage medium

By segmenting the gait image into image segments of the head, hands and feet and processing it in combination with attention mechanism, the problem of existing gait feature extraction methods ignore partial feature changes is solved, and the accuracy and efficiency of extraction are improved.

CN114724171BActive Publication Date: 2025-06-03SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011510090.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-18
Publication Date
2025-06-03
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

The existing gait feature extraction method performs overall extraction without ignoring the changes in part of the feature during the walking process, resulting in incomplete extraction and reducing accuracy.

Method used

By segmenting each frame of the M-frame image into three image segments, extracting the features of the head, hands and feet respectively, and processing these feature maps through attention mechanisms, a more comprehensive gait feature is generated.

Benefits of technology

It improves the accuracy of gait feature extraction, can more comprehensively capture the changing characteristics of the human body during walking, simplifies the network structure, and improves the extraction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724171B_ABST
    Figure CN114724171B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a gait feature extraction method, apparatus, and computer-readable storage medium, including: dividing each frame of M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set, where the M frames of images are consecutive M frames of images in a first video, each frame of image in the first video includes a first person, M is the number of images included in one gait cycle, and M is an integer greater than 1; performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set; splicing the feature maps corresponding to the same frame of image in the first feature map set, the second feature map set, and the third feature map set into a feature map to obtain a fourth feature map set; processing the fourth feature map set through an attention mechanism to obtain a first gait feature. The embodiment of the present application can improve the accuracy of gait feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image recognition technology, and in particular, to a gait feature extraction method, device, and computer-readable storage medium. Background Art

[0002] With the increasing attention to public safety issues, the application scope of video surveillance is getting larger and larger. Traditional identity feature extraction methods usually require high-definition information such as faces, irises, and fingerprints. In cases where the shooting distance is far, the light is dim, or there is no human cooperation, high-definition information cannot be obtained, so that these identity feature extraction methods are restricted.

[0003] Gait is a physiological and behavioral biometric that describes a person's walking pattern. Gait feature extraction can extract features from a person's walking posture to complete long-distance recognition. Currently, gait features can be extracted based on convolutional neural networks. However, the above-mentioned extracted gait features are extracted for the whole person, so that the changes in some features during the walking process of a person will be ignored, making the extraction of gait features not comprehensive enough, thus reducing the accuracy of gait feature extraction. Summary of the Invention

[0004] The embodiments of the present application disclose a gait feature extraction method, device, and computer-readable storage medium for improving the accuracy of gait feature extraction.

[0005] In a first aspect, a gait feature extraction method is provided, including:

[0006] Dividing each frame of M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set, where the M frames of images are consecutive M frames of images in a first video, each frame of image in the first video includes a first person, M is the number of images included in a gait cycle, and M is an integer greater than 1;

[0007] Performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set;

[0008] Concatenating the feature maps corresponding to the same frame of image in the first feature map set, the second feature map set, and the third feature map set into a single feature map to obtain a fourth feature map set;

[0009] Processing the fourth feature map set through an attention mechanism to obtain a first gait feature.

[0010] In the embodiments of the present application, since during the process of a person walking, the changes in human body movements are mainly concentrated on the limbs, only extracting the overall gait characteristics of a single person will ignore the characteristics of the changing parts of the human body during walking. At this time, the video picture set of a person in a gait cycle can be segmented into three image segment sets, and features can be extracted respectively. The feature sets extracted from each image segment set can reflect the gait characteristics of the human body in this part, which can improve the accuracy of gait feature extraction. This gait feature extraction network introduces an attention mechanism, which can adaptively learn the weights of features, distinguish the proportions of different gait features, and thus can further improve the accuracy of gait feature extraction. In addition, splicing the segmented feature maps together can not only consider the features extracted from each part, but also obtain the weights of all features by only one attention mechanism, which can simplify the structure of the gait feature extraction network and thus improve the efficiency of gait feature extraction.

[0011] As a possible implementation manner, the method further includes:

[0012] Obtain a second video, where the second video includes multiple frames of images, and the multiple frames of images include the first person;

[0013] Segment the images of the same person from the images included in the second video to obtain a first video.

[0014] In the embodiments of the present application, before M frames of images are input into the gait feature extraction network, the pictures of consecutive frames can be segmented from the video first. The segmented pictures can include a person, which can reduce the screening process for this person during the gait feature extraction process and improve the efficiency of gait feature extraction. In addition, since the human body is constantly changing during the walking process, it is necessary to obtain multiple frames of images in the video, and these multiple frames of images can reflect the changing conditions of the human body, so that the characteristics of the human body changes can be extracted, thereby improving the accuracy of gait feature extraction.

[0015] As a possible implementation manner, the step of segmenting each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set includes:

[0016] Perform a second convolution calculation on the M frames of images to obtain M frames of feature maps;

[0017] Segment each frame of the M frames of feature maps into three feature picture segments to obtain a first feature picture segment set, a second feature picture segment set, and a third feature picture segment set;

[0018] The step of performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set includes:

[0019] Perform a first convolution calculation on the first set of feature image segments, the second set of feature image segments, and the third set of feature image segments to obtain the first set of feature images, the second set of feature images, and the third set of feature images.

[0020] In the embodiments of the present application, when directly segmenting and extracting features from M frames of images, the overall features of a complete human picture will be damaged. Therefore, the complete features of the M frames of images can be extracted first, and then segmented and extracted. Features can be further extracted on the basis of retaining the complete features, which can ensure that the complete features and local features of the picture are extracted, thereby improving the accuracy of gait feature extraction.

[0021] As a possible implementation manner, the step of splitting each frame of the M frames of images into three image segments to obtain a first set of image segments, a second set of image segments, and a third set of image segments includes:

[0022] Split each frame of the M frames of images into three image segments according to a preset ratio to obtain the first set of image segments, the second set of image segments, and the third set of image segments;

[0023] The first set of image segments includes the head of the first person;

[0024] The second set of image segments includes the hands of the first person;

[0025] The third set of image segments includes the feet of the first person.

[0026] In the embodiments of the present application, the swinging amplitude of a person's hands and feet is the largest during walking. Therefore, the pictures or feature maps of the head, hands, and feet of the human body can be separated according to a certain ratio. Extract the features of the head, hands, and feet respectively, and the primary and secondary of the features of each part can be determined, thereby improving the accuracy of gait feature extraction. In addition, since the M frames of pictures include a picture of a person, the segmentation can be directly performed according to the ratio in the human height direction without further determining the positions of the head, hands, and feet, which can reduce the steps of picture segmentation and thus improve the efficiency of gait feature extraction.

[0027] As a possible implementation manner, the step of performing a first convolution calculation on the first set of image segments, the second set of image segments, and the third set of image segments to obtain a first set of feature images, a second set of feature images, and a third set of feature images includes:

[0028] Perform a first sub-convolution calculation on the first image segment to obtain a first set of feature images;

[0029] Perform a second sub-convolution calculation on the second image segment to obtain a second set of feature images;

[0030] Perform third sub-convolution calculation on the third image segment to obtain a third feature map set.

[0031] In the embodiments of the present application, when extracting the features of each image segment set, each image segment set can be separately subjected to sub-convolution calculation, and the convolution calculation of each part can extract the features corresponding to the head, hands, and feet in different ways to ensure that the important features of each part can be obtained, so as to ensure that both local features and overall features can be extracted, thereby improving the accuracy of gait extraction.

[0032] As a possible implementation manner, the process of obtaining the first gait feature by processing the fourth feature map set through an attention mechanism includes:

[0033] Transpose the fourth feature map set to obtain a transposed feature map set;

[0034] Perform convolution processing and normalization processing on the transposed fourth feature map set to obtain a weight set of the transposed feature map set;

[0035] Multiply and map the weight set of the transposed feature map set and the transposed feature map set to obtain the first gait feature.

[0036] In the embodiments of the present application, since local gait features are extracted by segmenting in the direction of human height, when calculating the weights of the feature maps through the attention mechanism, the feature weights in the direction of human height can be calculated. Transposing the fourth feature map set can realize the calculation of the feature weights in the direction of human height in the fourth feature map set, so as to ensure that features corresponding to different heights have different weights. Thus, the proportion of the feature weights of the head, hands, and feet can be distinguished, thereby improving the accuracy of gait feature extraction.

[0037] As a possible implementation manner, the method further includes:

[0038] Calculate the similarity between the first gait feature and the stored second gait feature;

[0039] When the similarity is greater than the threshold, determine that the first person is the person corresponding to the second gait feature.

[0040] In the embodiments of the present application, after extracting gait features, gait recognition can be performed to ensure the integrity of gait feature extraction and recognition. In addition, by calculating the similarity between gait features, the accuracy of gait feature extraction can be further verified, which is beneficial to the subsequent improvement and perfection of the gait feature extraction network.

[0041] A second aspect of the embodiments of the present application provides a gait feature extraction device, including a unit for performing the gait feature extraction provided in the first aspect or any embodiment of the first aspect. The gait feature extraction device may include:

[0042] A segmentation unit, configured to segment each frame of the M frames of images into three image segments, obtaining a first image segment set, a second image segment set, and a third image segment set. The M frames of images are consecutive M frames of images in a first video, each frame of the first video includes a first person, and M is the number of images included in one gait cycle, and M is an integer greater than 1;

[0043] A first convolution unit, configured to perform a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set, obtaining a first feature map set, a second feature map set, and a third feature map set;

[0044] A splicing unit, configured to splice the feature maps corresponding to the same frame of image in the first feature map set, the second feature map set, and the third feature map set into one feature map, obtaining a fourth feature map set;

[0045] A processing unit, configured to process the fourth feature map set through an attention mechanism, obtaining a first gait feature.

[0046] In one embodiment, the gait feature extraction device may further include:

[0047] An acquisition unit, configured to acquire a second video, where the second video includes multiple frames of images, and the multiple frames of images include the first person;

[0048] The segmentation unit is further configured to segment the images of the same person from the images included in the second video, obtaining the first video.

[0049] In one embodiment, the segmentation unit segmenting each frame of the M frames of images into three image segments, obtaining a first image segment set, a second image segment set, and a third image segment set includes:

[0050] Performing a second convolution calculation on the M frames of images, obtaining M frames of feature maps;

[0051] Segmenting each frame of the M frames of feature maps into three feature map segments, obtaining a first feature map segment set, a second feature map segment set, and a third feature map segment set;

[0052] The first convolution unit is specifically configured to:

[0053] Performing a first convolution calculation on the first feature map segment set, the second feature map segment set, and the third feature map segment set, obtaining the first feature map set, the second feature map set, and the third feature map set.

[0054] In one embodiment, the splitting unit splits each frame of the M-frame images into three image segments, and obtaining a first image segment set, a second image segment set, and a third image segment set includes:

[0055] Splitting each frame of the M-frame images into three image segments according to a preset ratio to obtain the first image segment set, the second image segment set, and the third image segment set;

[0056] The first image segment set includes the head of the first person;

[0057] The second image segment set includes the hands of the first person;

[0058] The third image segment set includes the feet of the first person.

[0059] In one embodiment, the first convolutional unit is specifically configured to:

[0060] Perform a first sub-convolution calculation on the first image segment to obtain a first feature map set;

[0061] Perform a second sub-convolution calculation on the second image segment to obtain a second feature map set;

[0062] Perform a third sub-convolution calculation on the third image segment to obtain a third feature map set.

[0063] In one embodiment, the processing unit is specifically configured to:

[0064] Transpose the fourth feature map set to obtain a transposed feature map set;

[0065] Perform convolution processing and normalization processing on the transposed fourth feature map set to obtain a weight set of the transposed feature map set;

[0066] Multiply and map the weight set of the transposed feature map set and the transposed feature map set correspondingly to obtain the first gait feature.

[0067] In one embodiment, the gait feature extraction device may further include:

[0068] A calculation unit, configured to calculate the similarity between the first gait feature and a stored second gait feature;

[0069] A determination unit, configured to determine that the first person is the person corresponding to the second gait feature when the similarity is greater than a threshold.

[0070] The third aspect of the embodiments of the present application provides a gait feature extraction device, including a processor and a memory, the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to call the program instructions to execute the gait feature extraction method provided by the first aspect or any one of the embodiments of the first aspect.

[0071] The fourth aspect provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the gait feature extraction method provided by the first aspect or any one of the embodiments of the first aspect.

[0072] The fifth aspect provides an application program, which is used to execute the gait feature extraction method provided by the first aspect or any one of the embodiments of the first aspect during operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0074] Figure 1 is a schematic flowchart of a gait feature extraction method disclosed in the embodiments of the present application;

[0075] Figure 2 is a schematic diagram of an image segmentation method disclosed in the embodiments of the present application;

[0076] Figure 3 is a schematic diagram of a training disclosed in the embodiments of the present application;

[0077] Figure 4 is a schematic structural diagram of a convolution module disclosed in the embodiments of the present application;

[0078] Figure 5 is a schematic structural diagram of another convolution module disclosed in the embodiments of the present application;

[0079] Figure 6 is a schematic structural diagram of an attention module disclosed in the embodiments of the present application;

[0080] Figure 7 is a schematic structural diagram of a gait feature extraction device provided by the embodiments of the present invention;

[0081] Figure 8 is a schematic structural diagram of another gait feature extraction device provided by the embodiments of the present invention. Specific Embodiments

[0082] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present invention.

[0083] The embodiments of the present application disclose a gait feature extraction method, device, and computer-readable storage medium for improving the accuracy of gait feature extraction. The following will be described in detail respectively.

[0084] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a gait feature extraction method disclosed in the embodiments of the present application. According to different requirements, Figure 1 certain steps in the flowchart shown in Figure 1 can be split into several steps. As shown in

[0085] 101. Divide each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set.

[0086] Before inputting the M frames of images of the first video into the gait feature extraction network, a second video can be obtained first. The second video can be multiple frames of images from a camera device or a storage device, and these multiple frames of images can include a first person. The first person is segmented from the second video to obtain the first video. The first video can be consecutive multiple images of the same person. The first video can be a binary video or a non-binary video. It should be understood that the first video can be images of one gait cycle or images of multiple gait cycles.

[0087] When the first video can be a binarized video, the second video of the first person segmented can also be binarized to obtain the first video, which can be consecutive multiple binarized images of the same person. Specifically, the images of the first video are binarized according to the contour of the first person, whereby other interference factors in the picture can be removed, and other non-gait features can be avoided from being extracted as well, thus improving the accuracy of gait feature extraction. When segmenting the first video from the second video, it can be segmented according to the position of a certain person. The first video can be an image segment of the second video. The sizes of the first videos should be the same. It should be understood that the first video can be the same as or different from the second video. For example, one of the people in the second video, i.e., the first person, can be recognized first. Then, the picture size of the segmented first video can be determined according to the ratio of the body height and body width of the person. Among them, this ratio can be set in advance, and the first video segmented according to this ratio can include a complete person. It should be understood that the above is an example of how to segment the first video and does not constitute a limitation.

[0088] After obtaining the first video, each frame of the M frames of images can be segmented into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set. The images are segmented according to a ratio. Figure 2 It is a schematic diagram of an image segmentation method disclosed in an embodiment of the present application. As Figure 2 shown, the M frames of pictures can be divided into three parts of image segments according to the upper, middle, and lower ratios of the human body height to obtain a first image segment set, a second image segment set, and a third image segment set. M can be the number of images included in the first video, and M can be an integer greater than 1. Since the sizes of the M frames of images segmented can be determined, the segmentation ratio can be set in advance. The first image segment set can include the head of the first person, the second image segment set can include the hands of the first person, and the third image segment set can include the feet of the first person. For example, for M frames of images of 100*60, a segmentation of 3:4:3 is performed to obtain a first image segment set with a size of 30*60, a second image segment set with a size of 40*60, and a third image segment set with a size of 30*60. It should be understood that the segmentation ratio can be set in advance, and the size of the ratio is not limited here. In addition, the segmented images can also be the feature maps of the first video after convolution.

[0089] After obtaining the first video, the M frames of images in the first video can be first subjected to convolution processing. Figure 3 It is a schematic diagram of a training disclosed in an embodiment of the present application. As Figure 3As shown, M-frame images can be input into a network for gait feature extraction for training. The gait feature extraction network may include a segmentation module to complete the above-mentioned segmentation process of the pictures. The M-frame images of the first video can be subjected to a second convolution calculation to obtain M-frame feature maps. The gait feature extraction network may further include a second convolution module, Figure 4 which is a schematic structural diagram of a convolution module disclosed in an embodiment of the present application, as Figure 4 shown. The second convolution module may include two convolutional layers and a pooling layer. Among them, the two convolutional layers may be convolutional layer 1 and convolutional layer 2 respectively. The size of the convolution kernel of convolutional layer 1 may be 5*5, and the stride may be 1; the size of the convolution kernel of convolutional layer 2 may be 3*3, and the stride may be 1; the size of the pooling layer may be 2*2, and the stride may be 2. For example, in the M-frame gait cycle images, the size of each image is 64*44. After passing through convolutional layer 1 and convolutional layer 2, the size of the image data may remain unchanged. Then, after passing through the pooling layer, the size of the image data may become 32*22, and M-frame feature maps of 32*22 are output. Then, each frame of the feature map in the M-frame feature map can be segmented into three feature picture segments by the segmentation module to obtain a first feature picture segment set, a second feature picture segment set, and a third feature picture segment set. The first feature picture segment set, the second feature picture segment set, and the third feature picture segment set are respectively subjected to a first convolution calculation to obtain a first feature map set, a second feature map set, and a third feature map set. For example, the above-mentioned M-frame feature maps of 32*22 are segmented in the direction of the human body height according to a ratio of 5:6:5, and the size of the obtained first feature picture segment set is 10*22, the size of the second feature picture segment set is 12*22, and the size of the third feature picture segment set is 10*22. It should be understood that the feature picture segment set at this time is the above-mentioned image segment set.

[0090] 102. The first image segment set, the second image segment set, and the third image segment set are subjected to a first convolution calculation to obtain a first feature map set, a second feature map set, and a third feature map set.

[0091] After obtaining the first image segment set, the second image segment set, and the third image segment set, these three image segment sets can be subjected to a first convolution calculation to obtain corresponding feature map sets. It can be understood that the first image segment is subjected to a first sub-convolution calculation to obtain a first feature map set, the second image segment is subjected to a second sub-convolution calculation to obtain a second feature map set, and the third image segment is subjected to a third sub-convolution calculation to obtain a third feature map set. The gait feature extraction network may further include a first convolution module. As Figure 3As shown, the first convolutional module may include a first sub-convolutional module, a second sub-convolutional module, and a third sub-convolutional module. The first sub-convolutional module, the second sub-convolutional module, and the third sub-convolutional module may respectively perform the first sub-convolutional calculation, the second sub-convolutional calculation, and the third sub-convolutional calculation. That is, it can be understood that the first set of image segments can be input into the first sub-convolutional module for the first sub-convolutional calculation to obtain the first set of feature maps; the second set of image segments can be input into the second sub-convolutional module for the second sub-convolutional calculation to obtain the second set of feature maps; the third set of image segments can be input into the third sub-convolutional module for the third sub-convolutional calculation to obtain the third set of feature maps. Figure 5 It is a schematic structural diagram of another convolutional module disclosed in an embodiment of the present application. As Figure 5 shown, the sub-convolutional module in the first convolutional module may include three convolutional layers and one pooling layer. For example, these three convolutional layers may include convolutional layer 3, convolutional layer 4, and convolutional layer 5. The sizes of the convolutional kernels of convolutional layer 3, convolutional layer 4, and convolutional layer 5 may be 3*3, 1*1, and 3*3 in sequence, and the stride of each convolutional kernel may be 1. The size of the pooling layer may be 2*2, and the stride may be 2. It should be understood that the structures of the first sub-convolutional module, the second sub-convolutional module, and the third sub-convolutional module may be the same, and the convolutional kernels of convolutional layer 3, convolutional layer 4, and convolutional layer 5 of the three sub-convolutional modules may be different.

[0092] 103. Concatenate the feature maps corresponding to the same frame of image in the first set of feature maps, the second set of feature maps, and the third set of feature maps into a single feature map to obtain the fourth set of feature maps.

[0093] After obtaining the first feature map set, the second feature map set, and the third feature map set, the first feature map set, the second feature map set, and the third feature map set can be spliced, and correspondingly, a complete fourth feature map set can be spliced. The gait feature extraction network may further include a splicing module, and the first feature map set, the second feature map set, and the third feature map set can be input into the splicing module to complete the splicing of the feature maps corresponding to the same frame of the M-frame images after segmentation. It can be understood that each of the M-frame images can be segmented, and then the feature maps of the segmented image segments can be determined. Then, the feature maps after segmentation in the same frame can be spliced to obtain a complete feature map, that is, the fourth feature map. For example, when M is 24, after the first frame of the image is segmented, image segments a, b, and c can be obtained. Then, feature extraction is respectively performed on the image segments a, b, and c to obtain feature maps a1, a2, and a3. The feature maps a1, a2, and a3 can be spliced in the segmentation direction to form a complete feature map. Therefore, the extraction of local features of the gait map can be ensured. The feature maps passing through the splicing module can be merged into one feature map, and this one feature map can be used for subsequent processing, instead of processing the three parts separately. This can not only ensure that each part of the features is processed, but also simplify the gait feature extraction network and improve the processing efficiency of gait features.

[0094] 104. Process the fourth feature map set through an attention mechanism to obtain the first gait feature.

[0095] After obtaining the fourth feature map set, the fourth feature map set can be processed through an attention mechanism to obtain the first gait feature. The gait feature extraction network may further include an attention module, and the fourth feature map set can be processed through the attention mechanism by the attention module. Figure 6 It is a schematic structural diagram of an attention module disclosed in an embodiment of the present application. As Figure 6As shown in the figure, the attention module may include a transpose module, a convolutional layer, a normalization module, a multiplication module, and a fully connected layer (FC). The fourth feature map set can be input into the transpose module for transposition to obtain a transposed feature map set. After transposition, the attention module can calculate weights for the features in the human height direction in the fourth feature map set. Then, the transposed feature map set can be input into the convolutional layer and the normalization module for convolutional processing and normalization processing of the transposed feature map to obtain a weight set of the transposed feature map set. Among them, the convolutional kernel size of the convolutional layer can be 1*1, the stride can be 1, and the normalization module can be a sigmiod layer or a softmax layer, etc. Then, the transposed feature map set and the transposed feature weight set can be multiplied correspondingly in the multiplication module to obtain a feature map. Then, this feature map can be input into the fully connected layer for mapping to obtain the corresponding first gait feature. Thus, the attention module can dynamically obtain the feature weights of the feature map set, so as to effectively distinguish the importance of different features, and thus improve the accuracy of gait feature extraction.

[0096] The gait feature extraction network may include several modules. According to the gait feature extraction network, an image can be input into a second convolutional module, a segmentation module, a first convolutional module, a splicing module, and an attention module for processing to obtain a first gait feature.

[0097] After obtaining the first gait feature map, the first gait feature can be further recognized to verify the accuracy of the above-mentioned gait feature extraction model. Specifically, the similarity between the first gait feature and the second gait feature that has been stored can be calculated. For example, the similarity between the first gait feature map and the second gait feature map can be determined by the Euclidean distance and the cosine similarity, etc. When the similarity is greater than the threshold, it can be determined that the first person is the person corresponding to the second gait feature.

[0098] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a gait feature extraction device provided by an embodiment of the present invention. As Figure 7 shown, the gait feature extraction device may include:

[0099] A segmentation unit 701, configured to segment each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set. The M frames of images are consecutive M frames of images in a first video. Each frame of image in the first video includes a first person. M is the number of images included in a gait cycle, and M is an integer greater than 1;

[0100] The first convolution unit 702 is configured to perform a first convolution calculation on the first set of image segments, the second set of image segments, and the third set of image segments to obtain a first set of feature maps, a second set of feature maps, and a third set of feature maps;

[0101] The splicing unit 703 is configured to splice the feature maps corresponding to the same frame of image in the first set of feature maps, the second set of feature maps, and the third set of feature maps into a single feature map to obtain a fourth set of feature maps;

[0102] The processing unit 704 processes the fourth set of feature maps through an attention mechanism to obtain the first gait feature.

[0103] In one embodiment, the gait feature extraction device may further include:

[0104] The acquisition unit 705 is configured to acquire a second video, where the second video includes multiple frames of images, and the multiple frames of images include the first person;

[0105] The segmentation unit 701 is further configured to segment the images of the same person from the images included in the second video to obtain a first video.

[0106] In one embodiment, the segmentation unit 701 segmenting each frame of the M frames of images into three image segments to obtain the first set of image segments, the second set of image segments, and the third set of image segments includes:

[0107] Performing a second convolution calculation on the M frames of images to obtain M frames of feature maps;

[0108] Segmenting each frame of the M frames of feature maps into three feature map segments to obtain a first set of feature map segments, a second set of feature map segments, and a third set of feature map segments;

[0109] The first convolution unit 702 is specifically configured to:

[0110] Performing a first convolution calculation on the first set of feature map segments, the second set of feature map segments, and the third set of feature map segments to obtain the first set of feature maps, the second set of feature maps, and the third set of feature maps.

[0111] In one embodiment, the segmentation unit 701 segmenting each frame of the M frames of images into three image segments to obtain the first set of image segments, the second set of image segments, and the third set of image segments includes:

[0112] Segmenting each frame of the M frames of images into three image segments according to a preset ratio to obtain the first set of image segments, the second set of image segments, and the third set of image segments;

[0113] The first set of image segments includes the head of the first person;

[0114] The second set of image segments includes the hand of the first person;

[0115] The third set of image segments includes the feet of the first person.

[0116] In one embodiment, the first convolutional unit 702 is specifically configured to:

[0117] Perform a first sub-convolution calculation on the first image segment to obtain a first atlas of features;

[0118] Perform a second sub-convolution calculation on the second image segment to obtain a second atlas of features;

[0119] Perform a third sub-convolution calculation on the third image segment to obtain a third atlas of features.

[0120] In one embodiment, the processing unit 704 is specifically configured to:

[0121] Transpose the fourth atlas of features to obtain a transposed atlas of features;

[0122] Perform convolutional processing and normalization processing on the transposed fourth atlas of features to obtain a weight set of the transposed atlas of features;

[0123] Multiply and map the weight set of the transposed atlas of features and the transposed atlas of features to obtain the first gait feature.

[0124] In one embodiment, the gait feature extraction device may further include:

[0125] A calculation unit 706, configured to calculate the similarity between the first gait feature and a stored second gait feature;

[0126] A determination unit 707, configured to determine that the first person is the person corresponding to the second gait feature when the similarity is greater than a threshold.

[0127] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of another gait feature extraction device provided by an embodiment of the present invention. As Figure 8As shown, the gait feature extraction device may include a processor 801, a memory 802, and a bus 803. The processor 801 may be a general-purpose central processing unit (CPU) or multiple CPUs, a single or multiple graphics processing units (GPUs), microprocessors, application-specific integrated circuits (ASICs), or one or more integrated circuits for controlling the execution of the program of the present invention. The memory 802 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions, or it may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 802 may exist independently or be integrated with the processor 801. The bus 803 is connected to the processor 801. The bus 803 transmits information between the above components.

[0128] A set of program codes is stored in the memory 802, and the processor 801 is configured to call the program codes stored in the memory 802 to execute the operations performed by the above-mentioned splitting unit 701, first convolution unit 702, splicing unit 703, processing unit 704, obtaining unit 705, calculation unit 706, and determination unit 707.

[0129] In one embodiment, a computer-readable storage medium is provided, which is used to store an application program, and the application program is used to execute Figure 1 the gait feature extraction method.

[0130] In one embodiment, an application program is provided, and the application program is used to execute Figure 1 the gait feature extraction method.

[0131] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, etc.

[0132] The above has introduced the embodiments of the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A gait feature extraction method, characterized in that, it includes: Segmenting the image of the first person from a second video including multiple frames of images to obtain a first video, where the multiple frames of images include the first person, and the first person is a complete person; Dividing each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set. The M frames of images are consecutive M frames of images in the first video, each frame of image in the first video is an image of the first person, M is the number of images included in a gait cycle, and M is an integer greater than 1; Performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set; Stitching the feature maps corresponding to the same frame of image in the first feature map set, the second feature map set, and the third feature map set into a single feature map to obtain a fourth feature map set; Processing the fourth feature map set through an attention mechanism to obtain a first gait feature.

2. The method according to claim 1, characterized in that, the method further includes: Obtaining a second video, where the second video includes multiple frames of images, and the multiple frames of images include the first person; Segmenting the images of the same person from the images included in the second video to obtain a first video.

3. The method according to claim 1, characterized in that, the step of dividing each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set includes: Performing a second convolution calculation on the M frames of images to obtain M frames of feature maps; Dividing each frame of the M frames of feature maps into three feature map segments to obtain a first feature map segment set, a second feature map segment set, and a third feature map segment set; The step of performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set includes: Performing a first convolution calculation on the first feature map segment set, the second feature map segment set, and the third feature map segment set to obtain the first feature map set, the second feature map set, and the third feature map set.

4. The method according to claim 1, characterized in that, the step of dividing each frame of the M frames of images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set includes: Dividing each frame of the M frames of images into three image segments according to a preset ratio to obtain the first image segment set, the second image segment set, and the third image segment set; The first image segment set includes the head of the first person; The second image segment set includes the hands of the first person; The third image segment set includes the feet of the first person.

5. The method according to claim 1, characterized in that, the step of performing a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set includes: Perform a first sub-convolution calculation on the first image segment to obtain a first feature map set; Perform a second sub-convolution calculation on the second image segment to obtain a second feature map set; Perform a third sub-convolution calculation on the third image segment to obtain a third feature map set.

6. The method according to claim 1, wherein, the process of obtaining the first gait feature by processing the fourth feature map set through an attention mechanism includes: transpose the fourth feature map set to obtain a transposed feature map set; perform convolution processing and normalization processing on the transposed feature map set to obtain a weight set of the transposed feature map set; multiply and map the weight set of the transposed feature map set and the transposed feature map set to obtain the first gait feature.

7. The method according to any one of claims 1-6, wherein, the method further includes: calculate the similarity between the first gait feature and the stored second gait feature; when the similarity is greater than a threshold, determine that the first person is the person corresponding to the second gait feature.

8. A gait feature extraction device, wherein, it includes: a segmentation unit, configured to segment each frame of the M-frame images into three image segments to obtain a first image segment set, a second image segment set, and a third image segment set. The M-frame images are consecutive M-frame images in a first video, each frame of the first video includes a first person, and M is the number of images included in a gait cycle, and M is an integer greater than 1; a first convolution unit, configured to perform a first convolution calculation on the first image segment set, the second image segment set, and the third image segment set to obtain a first feature map set, a second feature map set, and a third feature map set; a splicing unit, configured to splice the feature maps corresponding to the same frame of image in the first feature map set, the second feature map set, and the third feature map set into a single feature map to obtain a fourth feature map set; a processing unit, configured to process the fourth feature map set through an attention mechanism to obtain a first gait feature.

9. A gait feature extraction device, wherein, it includes a processor and a memory, the processor and the memory are connected to each other. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to call the program instructions to execute the gait feature extraction method according to any one of claims 1-7.

10. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the gait feature extraction method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Set-based cross-view gait recognition method

    CN109583298A

  • Gait feature extraction method and pedestrian identity recognition method based on gait features

    CN110084156A