High-rank convolution guided gait recognition method and system

By adopting a high-rank convolution-oriented method in gait recognition technology, an efficient feature extraction module is built, which solves the recognition problems in occlusion and uncertain gait cycle scenarios, achieves higher accuracy and efficiency, and reduces the amount of model parameters.

CN120220231APending Publication Date: 2025-06-27XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297737.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing gait recognition technologies perform poorly in occlusion and uncertain gait cycle scenarios, and the linear correlation of convolution kernel coefficients leads to feature extraction redundancy.

Method used

The high-rank convolution-oriented gait recognition method is adopted to construct efficient feature extraction modules GLDFA and GLDFB, combining space-time-channel decoupling convolution, grouping convolution, point-by-point convolution and normalization layers to extract global and local gait features, and enhance feature representation through superposition or splicing of local global features.

Benefits of technology

It improves the accuracy and efficiency of gait recognition, reduces the amount of model parameters, and avoids redundancy problems in feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220231A_ABST
    Figure CN120220231A_ABST
Patent Text Reader

Abstract

The invention discloses a high-rank convolution-guided gait recognition method and system. The method comprises the following steps: acquiring a gait video of a pedestrian; extracting a pedestrian contour map sequence from the gait video; and inputting the pedestrian contour map sequence into a trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain final gait discrimination features and complete gait recognition, and the method and system can accurately perform gait recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision pedestrian identity recognition, and relates to a gait recognition method and system guided by high-rank convolution. Background Art

[0002] Gait recognition is a technology for identifying identities through the walking patterns of pedestrians in videos. With advantages such as long distance and no cooperation required, it has good application prospects in public safety fields such as finding lost elderly people, children, and tracking suspects. However, the performance of gait recognition technology is affected by external factors such as clothing, carrying objects, and shooting angles, and a stable and efficient gait recognition method is still expected by all parties.

[0003] In the gait recognition technology based on video pedestrian appearance features, according to the obtained gait features, it can be roughly divided into four categories: average feature (abbreviated as GEI), global feature (abbreviated as GFR), local feature (abbreviated as LFR), and global and local fusion feature (abbreviated as GLFR). The gait recognition methods based on the average feature GEI first normalize multiple frames of pedestrian contour maps, and then generate a gait energy image by averaging the contour sequence images for gait recognition. The average gait feature can use the complete gait feature within the gait cycle to make up for the lack of local features caused by occlusion and walking habits. However, this type of method requires a complete gait cycle data support and is not suitable for scenarios with frame extraction, missing frames, or uncertain cycles caused by occlusion. The GFR class of methods based on global features directly extracts gait features from the original contour sequence or contour template, and tends to capture the global context relationship of gait. Whether there is a lack of contour integrity or not, it will look for the common features of the missing part and the non-missing part, and it is difficult to highlight the discriminative gait features of the occluded and unoccluded parts of the pedestrian. To solve this defect, some scholars have proposed the LFR class of methods based on local features. The gait contour map of the human body is divided into several local regions, and CNN is used to extract local features from each region. By restricting the receptive field of the network, the motion features of different body parts can be obtained. In the scenario of missing pedestrian contours caused by occlusion, it can highlight the gait information of the occluded and unoccluded parts. By adjusting the weights of different parts, it has a good effect on gait recognition with missing contours. However, the method based on LFR only focuses on the body part segmentation and ignores the cooperative motion relationship of each limb part during gait movement. Therefore, it cannot obtain a more complete gait feature representation in the case of no occlusion. In order to take into account both local details and the acquisition of global features, the GLFR class of gait recognition methods based on the fusion of global and local features has emerged. Among them, GaitGL (Lin B, Zhang S, Yu X. Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation[C]. IEEE / CVF International Conference on Computer Vision(ICCV), 2021: 14628–14636) uses 3D convolution to extract global gait features, and uses multiple shared 3D convolutions to act on the horizontal blocks of the pedestrian contour map to extract local features. Finally, the global and local fusion gait feature expression is realized through the superposition and splicing of features.Gaitedeg (Liang J, Fan C, Hou S, et al. Gaitedge: Beyond plain end-to-end gait recognition for better practicality[C]. European Conference on Computer Vision (ECCV), 2022: 375–390) proposed a micro-action capturer in the global and local feature expression framework to capture the motion patterns of different limbs. In occlusion scenarios, these methods achieved good results by extracting local discriminative features to resist the feature loss caused by limb occlusion. GaitMask (Lin B, Liu Y, Zhang S. GaitMask: Mask-based model for gait recognition[C]. British Machine Vision Conference (BMVC), 2021: 1–8) and GaitGL2 (Lin B, Zhang S, Wang M, Li L, Yu X. GaitGL: Learning Discriminative Global-Local Feature Representations for Gait Recognition. arXiv preprint arXiv:2208.01380, 2022 Aug 2. Available from: https: / / arxiv.org / abs / 2208.01380) introduced masks in the local feature extraction module to simulate occlusion, enabling the network to learn more robust local features; however, when occlusion occurs, they did not consider the feature learning problem of the global feature extraction module regarding occlusion.

[0004] The above methods all focus on how to obtain richer gait expressions, aiming to limit the limb shapes entering the gait algorithm and acquire local and global features of different limbs. They did not consider the problem of the feature extraction ability of each CNN convolution and the effectiveness of model inference. If there is a linear correlation in the convolution kernel coefficients, it may lead to a linear correlation between the gait features extracted by the convolution, that is, for a certain category of gait, the learned features are a combination of the gait features of other categories, indicating that there is learning redundancy in the feature extraction module of the gait recognition system.

[0005] To solve the balance problem between the efficiency and accuracy of gait recognition, a high-rank convolution-guided gait recognition method and system are proposed. The inherent rank of the convolution is used to evaluate the learning performance of each gait feature extraction module, and improved efficient feature extraction modules GLDFA and GLDFB are designed based on GaitGL. Summary of the Invention

[0006] The object of the present invention is to overcome the above-mentioned disadvantages of the prior art, and provide a gait recognition method and system guided by high-rank convolution, which can accurately perform gait recognition.

[0007] To achieve the above object, the present invention discloses a gait recognition method guided by high-rank convolution, including:

[0008] Obtain the gait video of a pedestrian;

[0009] Extract a sequence of pedestrian silhouette maps from the gait video;

[0010] Input the sequence of pedestrian silhouette maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature, and complete gait recognition.

[0011] The further improvement of the gait recognition method guided by high-rank convolution of the present invention lies in:

[0012] Furthermore, before inputting the sequence of pedestrian silhouette maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF, it further includes:

[0013] Use a semantic segmentation module to extract the sequence of pedestrian silhouettes in the gait video dataset, and construct a pedestrian gait dataset therefrom;

[0014] Construct efficient feature extraction modules GLDFA and GLDFB;

[0015] Utilize the efficient feature extraction modules GLDFA and GLDFB, and combine with the local-global feature extraction module GLConvA, the local temporal aggregation module LTA, the spatial aggregation module SP, the generalized average pooling layer, the fully connected layer and the hybrid loss function to construct the high-rank convolution-guided gait recognition network GaitGL-DF;

[0016] Utilize the gait dataset to train and test the high-rank convolution-guided gait recognition network GaitGL-DF to obtain the trained high-rank convolution-guided gait recognition network GaitGL-DF.

[0017] Furthermore, use the semantic segmentation module DeeplabV2 to extract the sequence of pedestrian silhouettes in the gait video dataset.

[0018] Furthermore, both the efficient feature extraction modules GLDFA and GLDFB include a spatio-temporal-channel decoupled convolution, a grouped convolution, a pointwise convolution and a normalization layer.

[0019] Furthermore, the efficient feature extraction modules GLDFA and GLDFB include a spatio-temporal-channel decoupled convolution SC(·), a grouped convolution GC(·), a pointwise convolution PC(·), and a batch normalization layer BN(·). Among them, the spatio-temporal-channel decoupled convolution SC is used to separate the spatio-temporal sampling and feature channel fusion operations, forming the spatio-temporal-channel decoupled convolution SC(·); the grouped convolution GC acts on each horizontal block of the pedestrian silhouette map to extract the motion features of different horizontal blocks, forming the local motion feature extraction convolution GC(·) of the gait; the pointwise convolution PC is used to extract the linkage relationship between each horizontal block, constituting the global feature extraction convolution PC(·) of the gait.

[0020] The GLDFA module is composed of the superposition of local features and global features to enhance the representation of gait features; the GLDFB module is composed of the concatenation of local features and global features to enrich the representation of gait features.

[0021] Furthermore, the spatio-temporal-channel decoupling module SCG is composed of the spatio-temporal-channel decoupled convolution SC and the grouped convolution GC.

[0022] Furthermore, the global and local gait features of the pedestrian obtained by the efficient feature extraction module GLDFA are X Aout = BN[PC(GC(SC(X Ain )))+GC(SC(X Ain ))], where X Aout ∈R C×t×h×w represents the global and local features output by the module. Among them, C is the number of channels of the feature map, t is the length of the feature map mapping, and (h, w) are the height and width of the feature map; SC(·) represents the spatio-temporal-channel decoupled convolution; GC(·) represents the grouped convolution; PC(·) represents the pointwise convolution; BN[·] represents batch normalization.

[0023] The global and local gait features of the pedestrian obtained by the efficient feature extraction module GLDFB are X Bout = BN{concat[PC(GC(SC(X Bin ))),GC(SC(X Bin ))]}, where X Bout ∈R C×t×2h×w represents the global and local features output by the module.

[0024] Furthermore, the process of inputting the pedestrian silhouette sequence into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature is as follows:

[0025] Input the pedestrian silhouette sequence into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the local feature Xlocal and the global feature X global , concatenate the local feature X local and the global feature X global to obtain the final gait feature.

[0026] The present invention discloses a gait recognition system guided by high-rank convolution, including:

[0027] An acquisition module, configured to acquire a gait video of a pedestrian;

[0028] A semantic segmentation module, configured to extract a sequence of pedestrian contour maps from the gait video;

[0029] An identification module, configured to input the sequence of pedestrian contour maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature and complete gait recognition.

[0030] A further improvement of the high-rank convolution-guided gait recognition system according to the present invention lies in:

[0031] Furthermore, it further includes:

[0032] A first construction module, configured to use the semantic segmentation module to extract a sequence of pedestrian contours from the gait video dataset and construct a pedestrian gait dataset therefrom;

[0033] A second construction module, configured to construct efficient feature extraction modules GLDFA and GLDFB;

[0034] A third construction module, configured to use the efficient feature extraction modules GLDFA and GLDFB, and in combination with a local-global feature extraction module GLConvA, a local temporal aggregation module LTA, a spatial aggregation module SP, a generalized average pooling layer, a fully connected layer, and a hybrid loss function, to construct a high-rank convolution-guided gait recognition network GaitGL-DF;

[0035] A training module, configured to use the gait dataset to train and test the high-rank convolution-guided gait recognition network GaitGL-DF to obtain the trained high-rank convolution-guided gait recognition network GaitGL-DF.

[0036] The present invention has the following beneficial effects:

[0037] When the high-rank convolution-guided gait recognition method and system according to the present invention are specifically operated, the pedestrian contour sequence is input into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait feature representation, and then the feature similarity is calculated to complete pedestrian gait discrimination. Among them, the efficient feature extraction modules GLDFA and GLDFB can obtain more discriminative global and local gait features from the pedestrian contour sequence with fewer parameters than conventional convolutions, realizing gait discrimination with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0039] Figure 1 is a structural diagram of the high-rank convolution-guided gait recognition network GaitGL-DF;

[0040] Figure 2 is a structural diagram of the efficient feature extraction module GLDFA;

[0041] Figure 3 is a structural diagram of the efficient feature extraction module GLDFB;

[0042] Figure 4 is a structural diagram of the spatio-temporal-channel decoupling module SCG;

[0043] Figure 5 is a schematic diagram of the global and local feature extraction of the efficient feature extraction modules GLDFA and GLDFB. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] In the description of the present invention, it should be understood that the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0046] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0047] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the present invention, the character " / " generally represents an "or" relationship between the contextually related objects.

[0048] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0049] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0050] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0051] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0052] The high-rank convolution-guided gait recognition method and system of the present invention include the following steps:

[0053] 1) Use the semantic segmentation module DeeplabV2 to extract the pedestrian contour sequence in the gait video dataset and construct a gait dataset therefrom.

[0054] 2) Construct an efficient feature extraction module;

[0055] Both the efficient feature extraction modules GLDFA and GLDFB include a spatio-temporal-channel decoupled convolution SC(·), a grouped convolution GC(·), a pointwise convolution PC(·), and a normalization layer BN(·). Among them, the spatio-temporal-channel decoupled convolution SC is used to separate the spatio-temporal sampling and feature channel fusion operations to form the spatio-temporal-channel decoupled convolution SC(·); the grouped convolution GC acts on each horizontal block of the pedestrian contour map to extract the motion features of each horizontal block and form the local motion feature extraction convolution GC(·) of the gait; the pointwise convolution PC is used to extract the linkage relationship between each horizontal block and constitute the global feature extraction convolution PC(·) of the gait; to enhance the representation of gait features, the GLDFA module is formed by superimposing local features and global features; to enrich the representation of gait features, the GLDFB module is formed by splicing local features and global features and is used to extract the spatio-temporal features of the pedestrian gait.

[0056] 3) Construct a gait recognition network;

[0057] Utilize the efficient feature extraction modules GLDFA and GLDFB constructed in step 2), in combination with the local-global feature extraction module GLConvA, the local temporal aggregation module LTA, the spatial aggregation module SP, the generalized mean pooling layer, the fully connected layer, and the hybrid loss function, to construct the high-rank convolution-guided gait recognition network GaitGL-DF.

[0058] 4) Gait recognition;

[0059] The gait dataset is divided into a training set and a test set. The test set is further divided into a query set and a registration set. Then, the gait sample data in the training set is used to train the gait recognition network GaitGL-DF guided by high-rank convolutional gait. During testing, the gait sample data of the query set and the registration set are fed into the trained high-rank convolutional gait recognition network GaitGL-DF for feature extraction. By measuring the feature similarity with the gait features of the registration set, the pedestrian gait discrimination in the query samples is completed.

[0060] The specific operation of step 1) is as follows:

[0061] The semantic segmentation module DeeplabV2 is used to extract the pedestrian contour sequence from the video.

[0062] The specific operation of step 2) is as follows:

[0063] The structures of the efficient feature extraction modules GLDFA and GLDFB are as Figure 2 and Figure 3 shown, both consisting of a spatio-temporal-channel decoupling convolution SC(·), a grouped convolution GC(·), a pointwise convolution PC(·), and a normalization layer BN(·), which are used to extract the spatio-temporal features of pedestrian gait.

[0064] The structure of the spatio-temporal-channel decoupling module SCG is as Figure 4 shown. First, the feature channels are adjusted through a 1×1×1 convolution, and then 3D convolution is used for sampling, separating the channel increase and sampling.

[0065] The structures of the global and local feature extractors in the efficient feature extraction modules GLDFA and GLDFB are as Figure 5 shown, including a grouped convolution GC(·) for extracting the local features of gait, and a pointwise convolution PC(·) for extracting the mutual relationship between local features to construct global features. Let its input be where C1 is the number of channels of the feature map, t is the length of the feature map mapping, and (h, w) are the height and width of the feature map. First, the local features are extracted through grouped convolution, and the corresponding local feature is where g is the number of groups of grouped convolution, corresponding to the i-th local feature. Then, the global features are extracted through pointwise convolution. C is the number of channels of the output feature map, that is, each channel has a 1×1×1 convolution kernel, and the output result is X global ∈R C×t× h ×w .

[0066] The enhanced gait features are constructed by superimposing the local features and the global features, and combined with the normalization layer BN to construct the GLDFA module as Figure 2 shown.

[0067] Construct gait features with richer scales by splicing local features and global features, and combine with the Batch Normalization (BN) layer to construct the GLDFB module as shown in Figure 3 .

[0068] The operation process of step 3) is as follows:

[0069] Construct a high-rank convolutional guided gait recognition network, the structure of which is as shown in Figure 1 . Specifically, construct a high-rank convolutional guided gait recognition network GaitGL-DF composed of a local-global feature extraction module GLConvA, a local temporal aggregation module LTA, an efficient feature extraction module GLDFA, a spatial aggregation module SP, an efficient feature extraction module GLDFB, a generalized mean pooling layer, a fully connected layer, and a hybrid loss function.

[0070] The specific operation of step 4) is as follows:

[0071] The gait dataset is derived from the publicly available CASIA-B dataset of the Institute of Automation, Chinese Academy of Sciences. This dataset contains 124 subjects, and each subject has 10 groups of gait videos. Among them, 6 groups are normal walking, 2 groups are walking with a backpack, and 2 groups are walking while wearing a coat. For each video, there are data of 11 shooting angles evenly distributed between 0° and 180°. Use these videos to construct a contour sequence dataset of pedestrian gait.

[0072] Divide the pedestrian contour dataset constructed in step 1) into a training set and a test set, and divide the test set into a query set and a registration set. That is, all the walking state data of subjects #001 - #074 are designated as the training set The normal walking state data #01 - 04 of subjects #075 - #124 are designated as the registration set The normal walking state data #05 - 06, backpack walking state data #01 - 02, and coat-wearing walking state data #01 - 02 of subjects #075 - #124 are designated as the query set

[0073] Then use the training set data to train the GaitGL-DF gait recognition network, and use a combined hybrid loss function of the triplet loss function Tripleloss and the cross-entropy loss function Cross Entropy for network training. During testing, send the query set and registration set data into the trained gait recognition network for gait feature extraction, and finally perform feature similarity measurement with the gait features of the registration set to complete the gait identification of the query samples.

[0074] To verify the advancement of the present invention, the performance of the present invention was compared with six existing gait recognition methods based on appearance features on the CASIA-B dataset. These six methods include OpenGait, GaitSet, GaitPart, GaitGL, GaitBase, and DeepGaitV2-3D. The implementation effect of the present invention is shown in Table 2.

[0075] OpenGait is by Fan C, Liang J, Shen C, et al. OpenGait: Revisiting Gait Recognition Toward Better Practicality[C]. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023: 9707–9716.

[0076] GaitSet is by Chao H, He Y, Zhang J, Feng J. Gaitset: Regarding gait as a set for cross-view gait recognition[C]. AAAI Conference on Artificial Intelligence, 2019, 33: 8126–8133.

[0077] GaitPart is by Fan C, Peng Y, Cao C, Liu X. Gaitpart: Temporal part-based model for gait recognition[C]. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020: 14213–14221.

[0078] GaitGL is by Lin B, Zhang S, Yu X. Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation[C]. IEEE / CVF International Conference on Computer Vision (ICCV), 2021: 14628–14636.

[0079] GaitBase is by Fan C, Liang J, Shen C, et al. OpenGait: Revisiting Gait Recognition Toward Better Practicality[C]. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023: 9707–9716.

[0080] DeepGaitV2-3D is by Fan C, Hou S, Huang Y, Yu S. Exploring Deep Models for Practical Gait Recognition. arXiv preprint arXiv:2303.03301, 2023. Retrieved from https: / / arxiv.org / pdf / 2303.03301v1.

[0081] As can be seen from Table 2, the average recognition rate of the present invention in CASIA-B is 91.99%, and the number of parameters of the gait recognition network is 2.43 Million (million). Compared with the original GaitGL

[15] network, the average recognition rate has increased by 0.19%, and the number of parameters has decreased by 21.36% ((3.09 - 2.43) / 3.09). In addition, the recognition rate of the present invention also exceeds other methods, and the number of parameters is small, verifying the high efficiency and robustness of the present invention under cross-walking conditions.

[0082] Table 2

[0083]

[0084] Example Two

[0085] The gait recognition system guided by high-rank convolution according to the present invention includes:

[0086] An acquisition module for acquiring the gait video of a pedestrian;

[0087] A semantic segmentation module for extracting a sequence of pedestrian contour maps from the gait video;

[0088] An identification module for inputting the sequence of pedestrian contour maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature and complete gait recognition.

[0089] In this embodiment, it further includes:

[0090] The first construction module is used to extract the sequence of pedestrian contours in the gait video dataset using a semantic segmentation module and construct a pedestrian gait dataset based on this;

[0091] The second construction module is used to construct the efficient feature extraction modules GLDFA and GLDFB;

[0092] The third construction module is used to utilize the efficient feature extraction modules GLDFA and GLDFB, and in combination with the local-global feature extraction module GLConvA, the local temporal aggregation module LTA, the spatial aggregation module SP, the generalized average pooling layer, the fully connected layer, and the hybrid loss function, construct the high-rank convolution-guided gait recognition network GaitGL-DF;

[0093] The training module is used to train and test the high-rank convolution-guided gait recognition network GaitGL-DF using the gait dataset to obtain the trained high-rank convolution-guided gait recognition network GaitGL-DF.

[0094] In the embodiments of the present application, the division of the modules is illustrative, merely a logical function division. In actual implementation, there may be other division methods. Additionally, in each embodiment of the present application, the various functional modules may be integrated in one processor, may also exist physically separately, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0095] Embodiment III

[0096] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the high-rank convolution-guided gait recognition method. For example, it includes: obtaining the gait video of a pedestrian; extracting the sequence of pedestrian contour maps from the gait video; inputting the sequence of pedestrian contour maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature and complete gait recognition. Among them, the memory may include internal memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, which may be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory and provide instructions and data to the processor.

[0097] Embodiment IV

[0098] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the high-rank convolution-guided gait recognition method. For example, it includes: obtaining a gait video of a pedestrian; extracting a sequence of pedestrian silhouette maps from the gait video; inputting the sequence of pedestrian silhouette maps into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature, thereby completing gait recognition. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disk, etc.

[0099] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0100] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.

[0103] Those skilled in the art will readily conceive of other embodiments of the present invention upon consideration of the specification and the disclosure of the invention. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0104] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

[0105] The above are only the preferred embodiments of the present invention, and do not limit the present invention in any way. Any simple modifications, changes, and equivalent structural changes made to the above embodiments according to the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A high-rank convolution-guided gait recognition method, characterized in that: include: Get the pedestrian's gait video; Extracting a pedestrian profile sequence from the gait video; The pedestrian profile sequence is input into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination features and complete gait recognition.

2. The high-rank convolution-guided gait recognition method according to claim 1, characterized in that: Before inputting the pedestrian profile sequence into the trained high-rank convolution-guided gait recognition network GaitGL-DF, the method further includes: Use the semantic segmentation module to extract pedestrian silhouette sequences from the gait video dataset and use it to construct a pedestrian gait dataset; Construct efficient feature extraction modules GLDFA and GLDFB; Using the efficient feature extraction modules GLDFA and GLDFB, combined with the local-global feature extraction module GLConvA, the local-time aggregation module LTA, the spatial aggregation module SP, the generalized average pooling layer, the fully connected layer and the hybrid loss function, a high-rank convolution-guided gait recognition network GaitGL-DF is constructed; The high-rank convolution-guided gait recognition network GaitGL-DF is trained and tested using the gait data set to obtain a trained high-rank convolution-guided gait recognition network GaitGL-DF.

3. The high-rank convolution-guided gait recognition method according to claim 2, characterized in that: The semantic segmentation module DeeplabV2 is used to extract pedestrian silhouette sequences in the gait video dataset.

4. The high-rank convolution-guided gait recognition method according to claim 2, characterized in that: The efficient feature extraction modules GLDFA and GLDFB each include a spatiotemporal-channel decoupled convolution, a grouped convolution, a point-by-point convolution and a normalization layer.

5. The high-rank convolution-guided gait recognition method according to claim 4, characterized in that: The efficient feature extraction modules GLDFA and GLDFB include a spatiotemporal-channel decoupled convolution SC(·), a grouped convolution GC(·), a point-by-point convolution PC(·) and a normalization layer BN(·), wherein the spatiotemporal-channel decoupled convolution SC is used to separate the spatiotemporal sampling and feature channel fusion operations to form a spatiotemporal-channel decoupled convolution SC(·); the grouped convolution GC acts on each horizontal block of the pedestrian profile to extract the motion features of different horizontal blocks to form a local motion feature extraction convolution GC(·) of the gait; the point-by-point convolution PC is used to extract the linkage relationship between each horizontal block to form a global feature extraction convolution PC(·) of the gait; The GLDFA module is composed of the superposition of local features and global features to enhance the representation of gait features; the GLDFB module is composed of the concatenation of local features and global features to enrich the representation of gait features.

6. The high-rank convolution-guided gait recognition method according to claim 4, characterized in that: The space-time-channel decoupling module SCG is composed of space-time-channel decoupling convolution SC and group convolution GC.

7. The high-rank convolution-guided gait recognition method according to claim 4, characterized in that: The global and local gait features of pedestrians obtained by the efficient feature extraction module GLDFA are X Aout =BN[PC(GC(SC(X Ain )))+GC(SC(X Ain ))], where X Aout ∈R C×t×h×w Represents the global and local features of the module output, where C is the number of channels of the feature map, t is the length of the feature map mapping, and (h, w) are the height and width of the feature map; SC(·) represents spatiotemporal-channel decoupled convolution; GC(·) represents grouped convolution; PC(·) represents point-by-point convolution; BN[·] represents batch normalization; The global and local gait features of pedestrians obtained by the efficient feature extraction module GLDFB are X Bout =BN{concat[PC(GC(SC(X Bin ))),GC(SC(X Bin ))]}, where X Bout ∈R C×t×2h×w Represents the global and local features of the module output.

8. The high-rank convolution-guided gait recognition method according to claim 2, characterized in that: The pedestrian profile sequence is input into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination feature: The pedestrian profile sequence is input into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the local features X local and the global feature X global , the local feature X local and the global feature X global Splicing to get the final gait features.

9. A high-rank convolution-guided gait recognition system, characterized in that: include: An acquisition module, used to acquire the gait video of pedestrians; A semantic segmentation module, used for extracting a pedestrian profile sequence from the gait video; The recognition module is used to input the pedestrian profile sequence into the trained high-rank convolution-guided gait recognition network GaitGL-DF to obtain the final gait discrimination features and complete gait recognition.

10. The high-rank convolution-guided gait recognition system according to claim 9, characterized in that: Also includes: The first construction module is used to extract pedestrian contour sequences in the gait video dataset using a semantic segmentation module, and construct a pedestrian gait dataset based on this; The second building module is used to build efficient feature extraction modules GLDFA and GLDFB; The third construction module is used to use the efficient feature extraction modules GLDFA and GLDFB, and combine the local global feature extraction module GLConvA, the local time aggregation module LTA, the spatial aggregation module SP, the generalized average pooling layer, the fully connected layer and the hybrid loss function to construct a high-rank convolution-guided gait recognition network GaitGL-DF; A training module is used to train and test the high-rank convolution-guided gait recognition network GaitGL-DF using the gait data set to obtain the trained high-rank convolution-guided gait recognition network GaitGL-DF.