Training Method of Gait Recognition Model Augmented by Motion Mode, Gait Recognition Method, Medium, and Controller

Through the training method of augmenting gait recognition model, combining the first branch network and the second branch network to extract features, and fusion of multi-stage feature aggregation networks, the problem of poor adaptability of gait recognition method in complex scenarios is solved, achieving higher identity recognition accuracy.

CN119723245BActive Publication Date: 2025-07-22ANHUI GUANGCHENG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411801793.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-07-22
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

The existing representation-based gait recognition methods are poorly adaptable in complex scenarios, and it is difficult to make full use of dynamic human movement information in gait sequences, resulting in low accuracy of identity recognition.

Method used

The gait recognition model with augmented motion mode is adopted, and the features of the original gait silhouette sequence and the augmented gait silhouette sequence are extracted respectively through the first branch network and the second branch network, and the feature fusion is performed through the multi-stage feature aggregation network. The model parameters are optimized using the triple loss function and the cross entropy loss function to generate diverse motion apparent information and pedestrian identity clues.

Benefits of technology

The accuracy of identity recognition of the gait recognition model in complex scenarios has been improved, the ability to utilize dynamic human movement information has been enhanced, and the adaptability and discrimination ability of the model has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723245B_ABST
    Figure CN119723245B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, a gait recognition method, a medium, and a controller for a gait recognition model with augmented motion patterns. The training method includes: performing multiple periodic trainings on the gait recognition model using a training image set. For each training cycle, the following operations are performed: obtaining the training image set, where the training image set includes an original gait silhouette sequence and its corresponding augmented gait silhouette sequence; inputting the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain target gait features and their corresponding classification probabilities; obtaining a triplet loss function based on the target gait features, obtaining a cross-entropy loss function based on the classification probabilities corresponding to the target gait features, and obtaining a total loss function based on the triplet loss function and the cross-entropy loss function; adjusting the parameters of the gait recognition model according to the total loss function to obtain a trained gait recognition model. This training method improves the accuracy of identity recognition of the gait recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gait recognition, and in particular, to a training method for a gait recognition model with augmented motion patterns, a gait recognition method, a medium, and a controller. Background Art

[0002] Gait recognition is a technology for identity recognition based on human body shape and motion patterns, and has been widely applied in many fields such as public security and finance in recent years. Gait recognition technology can obtain gait data under the conditions of long distance, non-controlled, and non-perceived. By extracting the apparent features of the human body shape and the dynamic features of the motion pattern from the gait data, accurate recognition of pedestrian identities can be achieved.

[0003] According to the data types of the research objects, gait recognition can be roughly divided into model-based methods and representation-based methods. The model-based method first uses a human pose estimation method to extract the coordinates of human key points, and then performs explicit modeling on the human body structure and walking manner, and extracts gait features. Although the model-based method can robustly cope with perspective changes, it is greatly affected by the performance of the human pose estimation model and has a large computational cost. The representation-based method directly uses a convolutional neural network to extract discriminative identity features from gait images.

[0004] In recent years, benefiting from the powerful feature extraction ability of CNN (Convolutional Neural Network), the representation-based gait recognition method is often superior to the model-based method, and thus has received more attention from researchers. The representation-based gait recognition method has the problem of poor adaptability to complex scenarios. Summary of the Invention

[0005] The present invention aims to at least solve one of the technical problems in the related technologies to some extent. For this reason, one objective of the present invention is to propose a training method for a gait recognition model with augmented motion patterns, which improves the identity recognition accuracy of the gait recognition model.

[0006] The second objective of the present invention is to propose a gait recognition method with augmented motion patterns.

[0007] The third objective of the present invention is to propose a computer-readable storage medium.

[0008] The fourth objective of the present invention is to propose a controller.

[0009] To achieve the above object, an embodiment of the first aspect of the present invention proposes a training method for a gait recognition model with augmented motion patterns. The gait recognition model includes a first branch network, a second branch network, a multi-stage feature aggregation network, an element-wise feature summation layer, a fully connected layer, and a classifier. Among them, the first branch network is used to input the original gait silhouette sequence, the second branch network is used to input the augmented gait silhouette sequence, the output ends of the first branch network and the second branch network are connected to the input end of the multi-stage feature aggregation network, the output end of the first branch network and the output end of the multi-stage feature aggregation network are connected to the input end of the element-wise feature summation layer, the output end of the element-wise feature summation layer is connected to the input end of the classifier through the fully connected layer, and the classifier is used to output the target gait features and their corresponding classification probabilities. The training method includes: using the training image set to perform multiple periodic trainings on the gait recognition model. For each training cycle, the following operations are performed: obtaining the training image set, where the training image set includes the original gait silhouette sequence and its corresponding augmented gait silhouette sequence; inputting the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain the target gait features and their corresponding classification probabilities; obtaining a triplet loss function according to the target gait features, obtaining a cross-entropy loss function according to the classification probabilities corresponding to the target gait features, and obtaining a total loss function according to the triplet loss function and the cross-entropy loss function; adjusting the parameters of the gait recognition model according to the total loss function to obtain a trained gait recognition model.

[0010] According to the training method of the gait recognition model with augmented motion patterns according to the embodiment of the present invention, the original gait silhouette sequence is amplified to increase the augmented gait silhouette sequence, providing diverse motion appearance information and pedestrian identity clues for the gait recognition model. By extracting and fusing features from the original gait silhouette sequence and the augmented gait features, the identity recognition accuracy of the gait recognition model is improved.

[0011] In addition, the training method of the gait recognition model with augmented motion patterns proposed according to the above embodiment of the present invention may further have the following additional technical features:

[0012] According to an embodiment of the present invention, inputting the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain the target gait feature and its corresponding classification probability includes: inputting the original gait silhouette sequence into the first branch network to obtain the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature; inputting the augmented gait silhouette sequence into the second branch network to obtain the augmented gait feature; inputting the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the augmented gait feature into the multi-stage feature aggregation network to obtain the first fusion feature; inputting the original gait feature and the first fusion feature into the element-level feature summation layer to obtain the second fusion feature; and sequentially inputting the second fusion feature into the fully connected layer and the classifier to obtain the target gait feature and its corresponding classification probability.

[0013] According to an embodiment of the present invention, the first branch network includes a first feature extraction module, a second feature extraction module, a first time aggregation module, a third feature extraction module, a time pooling module, and a space pooling module connected in sequence. Inputting the original gait silhouette sequence into the first branch network to obtain the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature includes: using the first feature extraction module to respectively perform dynamic change feature extraction and multi-granularity feature extraction on the original gait silhouette sequence, and performing an element-level feature summation operation on the extracted first dynamic change feature and first multi-granularity feature to obtain the first-layer fusion feature; using the second feature extraction module to respectively perform dynamic change feature extraction and multi-granularity feature extraction on the first-layer fusion feature, and performing an element-level feature summation operation on the extracted second dynamic change feature and second multi-granularity feature to obtain the second-layer fusion feature; using the first time aggregation module to perform a time-dimensional compression and aggregation operation on the second-layer fusion feature to obtain the first aggregation feature; using the third feature extraction module to respectively perform dynamic change feature extraction and multi-granularity feature extraction on the first aggregation feature, and performing an element-level feature summation operation on the extracted third dynamic change feature and third multi-granularity feature to obtain the third-layer fusion feature; using the time pooling module to perform a time-dimensional feature pooling operation on the third-layer fusion feature to obtain a pooled feature; and using the space pooling module to perform a space-dimensional feature compression on the pooled feature to obtain the original gait feature.

[0014] According to an embodiment of the present invention, the augmented gait silhouette sequence includes at least two of a forward walking sequence, a backward walking sequence, an upstairs walking sequence, and a downstairs walking sequence. The second branch network includes a shallow feature extraction module and a channel-level feature aggregation module. Inputting the augmented gait silhouette sequence into the second branch network to obtain augmented gait features includes: using the shallow feature extraction module to respectively perform feature extraction on at least two of the forward walking sequence, the backward walking sequence, the upstairs walking sequence, and the downstairs walking sequence to obtain at least two of a first gait feature, a second gait feature, a third gait feature, and a fourth gait feature; using the channel-level feature aggregation module to perform feature aggregation on at least two of the first gait feature, the second gait feature, the third gait feature, and the fourth gait feature to obtain the augmented gait features.

[0015] According to an embodiment of the present invention, the multi-stage feature aggregation network includes a first feature aggregation module, a second temporal aggregation module, a second feature aggregation module, and a third feature aggregation module. Inputting the first-layer fusion feature, the first aggregated feature, the third-layer fusion feature, and the augmented gait features into the multi-stage feature aggregation network to obtain a first fusion feature includes: using the first feature aggregation module to perform a channel-level feature splicing operation on the augmented gait features and the first-layer fusion feature, performing feature channel adjustment on the first spliced feature obtained by the splicing operation to obtain a first adjusted feature, and performing feature channel adjustment and feature compression processing on the first spliced feature obtained by the splicing operation to obtain a first compressed feature; using the second temporal aggregation module to perform a compression aggregation operation on the first adjusted feature in the temporal dimension to obtain a second aggregated feature; using the second feature aggregation module to perform a channel-level feature splicing operation on the first aggregated feature and the second aggregated feature, performing feature channel adjustment on the second spliced feature obtained by the splicing operation to obtain a second adjusted feature, and performing feature channel adjustment and feature compression processing on the first spliced feature obtained by the splicing operation to obtain a second compressed feature; using the third feature aggregation module to perform a channel-level feature splicing operation on the second adjusted feature and the third-layer fusion feature, and performing feature channel adjustment and feature compression processing on the third spliced feature obtained by the splicing operation to obtain a third compressed feature; performing a channel-level feature splicing operation on the first compressed feature, the second compressed feature, and the third compressed feature to obtain a first fusion feature.

[0016] According to an embodiment of the present invention, obtaining an augmented gait silhouette sequence includes: obtaining an original gait silhouette sequence; for each gait silhouette image in the original gait silhouette sequence, evenly dividing each image in the gait silhouette image in the horizontal direction to obtain r slice images, fixing the first slice image at the top, and shifting the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels to the right in sequence along the vertical direction to generate a forward walking sequence; or, for each gait silhouette image in the original gait silhouette sequence, evenly dividing each image in the gait silhouette image in the horizontal direction to obtain r slice images, fixing the first slice image at the top, and shifting the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels to the left in sequence along the vertical direction to generate a backward walking sequence; or, for each gait silhouette image in the original gait silhouette sequence, evenly dividing each image in the gait silhouette image in the vertical direction to obtain r slice images, fixing the first slice image at the leftmost side, and shifting the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels downwards in sequence along the horizontal direction to generate an upstairs walking sequence; or, for each gait silhouette image in the original gait silhouette sequence, evenly dividing each image in the gait silhouette image in the vertical direction to obtain r slice images, fixing the first slice image at the top, and shifting the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels upwards in sequence along the horizontal direction to generate a downstairs walking sequence.

[0017] According to an embodiment of the present invention, the expression of the total loss function is:

[0018] L all = L triplet + L ce

[0019] wherein, L all represents the total loss function, L triplet represents the triplet loss function, and L ce represents the cross - entropy loss function;

[0020]

[0021] wherein, K represents the number of triplets, and the triplet is (F pos , F neg , F anc ), F pos represents the positive sample feature, F neg represents the negative sample feature, and F anc represents the anchor feature, is the Euclidean distance between the feature F pos and the feature F neg , is the feature Fpos the Euclidean distance to feature F anc where \(1\leq k\leq K\), \(\max(\cdot)\) represents the maximum function, and margin represents the boundary threshold;

[0022]

[0023] where \(M\) represents the number of samples, \(1\leq i\leq M\), \(y\) i represents the one - hot encoded type label of the \(i\) - th sample, and \(p\) i represents the classification result of the \(i\) - th sample.

[0024] To achieve the above object, an embodiment of the second aspect of the present invention proposes a gait recognition method with augmented motion patterns. The recognition method includes: obtaining an original gait silhouette sequence to be measured, augmenting the gait silhouette sequence to be measured to obtain an augmented gait silhouette sequence to be measured; inputting the original gait silhouette sequence to be measured and the augmented gait silhouette sequence to be measured into a pre - trained gait recognition model to obtain target gait features and their corresponding classification probabilities, where the pre - trained gait recognition model is obtained by using the training method of the gait recognition model with augmented motion patterns proposed in the embodiment of the first aspect of the present invention.

[0025] To achieve the above object, an embodiment of the third aspect of the present invention proposes a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the gait recognition model with augmented motion patterns proposed in the embodiment of the first aspect of the present invention, or the gait recognition method with augmented motion patterns proposed in the embodiment of the second aspect of the present invention.

[0026] To achieve the above object, an embodiment of the fourth aspect of the present invention proposes a controller, including a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the training method of the gait recognition model with augmented motion patterns proposed in the embodiment of the first aspect of the present invention, or the gait recognition method with augmented motion patterns proposed in the embodiment of the second aspect of the present invention.

[0027] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a structural diagram of a gait recognition model according to an embodiment of the present invention;

[0029] Figure 2 is a flowchart of a training method of a gait recognition model according to an embodiment of the present invention;

[0030] Figure 3 Schematic diagram of generating an augmented gait silhouette sequence according to an embodiment of the present invention;

[0031] Figure 4 Structural diagram of a gait recognition model according to a specific embodiment of the present invention;

[0032] Figure 5 Flowchart of generating target gait features and classification probabilities according to a specific embodiment of the present invention;

[0033] Figure 6 Flowchart of obtaining original gait features according to a specific embodiment of the present invention;

[0034] Figure 7 Flowchart of the processing of the dynamic change perception unit according to a specific embodiment of the present invention;

[0035] Figure 8 Flowchart of the processing of the multi - granularity feature extraction unit according to a specific embodiment of the present invention;

[0036] Figure 9 Flowchart of obtaining augmented gait features according to a specific embodiment of the present invention;

[0037] Figure 10 Flowchart of obtaining the first fusion feature according to a specific embodiment of the present invention;

[0038] Figure 11 Flowchart of the processing of the feature aggregation module according to an embodiment of the present invention;

[0039] Figure 12 Flowchart of a gait recognition method according to an embodiment of the present invention;

[0040] Figure 13 Structural block diagram of the controller according to an embodiment of the present invention. Detailed implementation manners

[0041] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0042] It should be noted that gait recognition methods based on representation can generally be divided into three categories. The first category of methods regards the human body contour sequence as an unordered set of image frames, extracts frame-level features and set features respectively, and uses a horizontal pyramid mapping module to aggregate global and local features to enhance the representation ability of gait features; the second category of methods focuses on the division method of gait features on the spatial scale, and uses 3D convolution to obtain the spatio-temporal representation of the whole and each local part of the human gait; the third category of methods focuses on modeling the temporal relationship of the gait sequence and capturing the motion characteristics, and tries to design multi-scale temporal feature extraction methods and motion feature capture methods.

[0043] The above methods often focus on the extraction of spatial features or temporal features, lacking effective means for extracting highly discriminative motion appearance features, resulting in these methods being difficult to fully utilize the dynamic human motion information in the gait sequence, and thus having poor adaptability to complex scenarios.

[0044] The following will describe in detail the training method, gait recognition method, medium, and controller of the gait recognition model with augmented motion patterns according to the embodiments of the present invention in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0045] As Figure 1 shown, the gait recognition model with augmented motion patterns in the embodiments of the present invention may include a first branch network, a second branch network, a multi-stage feature aggregation network, an element-level feature summation layer, a fully connected layer, and a classifier. Among them, the first branch network is used to input the original gait silhouette sequence, the second branch network is used to input the augmented gait silhouette sequence, the output ends of the first branch network and the second branch network are connected to the input end of the multi-stage feature aggregation network, the output end of the first branch network and the output end of the multi-stage feature aggregation network are connected to the input end of the element-level feature summation layer, and the output end of the element-level feature summation layer is connected to the input end of the classifier through the fully connected layer. The classifier is used to output the target gait features and their corresponding classification probabilities.

[0046] To improve the adaptability of the gait recognition model to complex scenarios, in the embodiments of the present invention, the gait recognition model is provided with a first branch network for extracting the features of the original gait silhouette sequence, a second branch network for extracting the features of the augmented gait silhouette sequence, and a multi-stage feature aggregation network for fusing the features extracted from the original gait silhouette sequence by the first branch network and the features extracted from the augmented gait silhouette sequence by the second branch network, and performs identity recognition based on the features output by the first branch network and the multi-stage feature aggregation network.

[0047] It should be noted that the augmented gait silhouette sequence is obtained based on the original gait silhouette sequence, so that when the gait recognition model does not perform identity recognition, it can make full use of the dynamic human motion information in the original gait silhouette sequence.

[0048] Figure 2 is a flowchart of a method for training a gait recognition model according to an embodiment of the present invention. As Figure 2 shown, the method for training a gait recognition model with augmented motion patterns may include:

[0049] Performing multiple periodic trainings on the gait recognition model using a training image set, wherein for each training cycle, the following operations are performed:

[0050] S101, obtaining a training image set, wherein the training image set includes an original gait silhouette sequence and its corresponding augmented gait silhouette sequence;

[0051] S102, inputting the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain target gait features and their corresponding classification probabilities;

[0052] S103, obtaining a triplet loss function according to the target gait features, obtaining a cross-entropy loss function according to the classification probabilities corresponding to the target gait features, and obtaining a total loss function according to the triplet loss function and the cross-entropy loss function;

[0053] S104, adjusting the parameters of the gait recognition model according to the total loss function to obtain a trained gait recognition model.

[0054] Specifically, when obtaining the gait silhouette image sequence , electronic devices such as cameras can be used to capture gait silhouette data of different people, and then the captured gait silhouette data is processed to obtain a gait silhouette image sequence. It is also possible to directly obtain the gait silhouette image sequence from the CASIA-B dataset and the CCPG dataset. It should be noted that the CASIA-B dataset contains gait silhouette image sequences of 124 pedestrians, and the CCPG dataset contains gait silhouette image sequences of 200 pedestrians. The embodiments of the present invention do not limit the method for obtaining the gait silhouette image sequence.

[0055] In a specific embodiment of the present invention, the image size during the training of the gait recognition model is set to H×W = 64×44, that is, the image height is 64, the width is 44, the input sequence length T is 30 frames, and the input channel C = 1, indicating that the image is a grayscale image.

[0056] Performing data augmentation on the obtained original gait silhouette sequence , performing splitting and translation operations on the silhouette images in the original gait silhouette sequence to generate an augmented gait silhouette sequence. The augmented gait silhouette sequence in the embodiments of the present invention includes at least a forward walking sequence D lf , a backward walking sequence D bb , an upstairs walking sequence D gu and a downstairs walking sequence D gdboth of them in

[0057] Input the original gait silhouette sequence D in the training image set and its corresponding augmented gait silhouette sequence into the gait recognition model. The first branch network extracts features from the original gait silhouette sequence D, the second branch network extracts features from the augmented gait silhouette sequence, the multi-stage feature aggregation network fuses the features extracted by the first branch network and the second network, and the element-wise feature summation layer performs element-wise feature summation on the features extracted by the first branch network and the multi-stage feature aggregation network. The features obtained by element-wise feature summation are sequentially input into the fully connected layer and the classifier, and the classifier outputs the target gait feature F fc and its corresponding classification probability P.

[0058] For the target gait feature F fc Calculate the triplet loss function L triplet , and calculate the cross-entropy loss function L for the classification probability P ce . Therefore, the total loss function L of the gait recognition model all can be expressed as: L all = L triplet + L ce . The gradient descent method can be used to train the gait recognition model and calculate the total loss function L all to update the network parameters. When the number of training iterations reaches the set number or the total loss function L all converges, the training stops, thus obtaining the trained gait recognition model (optimal gait recognition model). The trained gait recognition model is used to identify the identity of the gait silhouette sequence, and the identity recognition result is obtained.

[0059] In an embodiment of the present invention, as Figure 3 shown, obtaining the augmented gait silhouette sequence includes:

[0060] Obtain the original gait silhouette sequence D;

[0061] For each gait silhouette image in the original gait silhouette sequence D, each image in the gait silhouette image is evenly divided horizontally to obtain r slice images. Fix the first slice image at the top, and move the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels to the right vertically in sequence to generate the forward walking sequence D lf ; or,

[0062] For each gait silhouette image in the original gait silhouette sequence, each image in the gait silhouette image is evenly divided horizontally to obtain r slice images. Fix the first slice image at the top, and move the remaining r - 1 slice images l, 2×l,..., (r - 1)×l pixels to the left vertically in sequence to generate the backward walking sequence D bb; or,

[0063] For each gait silhouette image in the original gait silhouette sequence, each image in the gait silhouette image is evenly sliced in the vertical direction to obtain r sliced images. Fix the first sliced image on the leftmost side, and move the remaining r - 1 sliced images downward by l, 2×l, ……, (r - 1)×l pixels in sequence along the horizontal direction to generate an upstairs walking sequence D gu ; or,

[0064] For each gait silhouette image in the original gait silhouette sequence, each image in the gait silhouette image is evenly sliced in the vertical direction to obtain r sliced images. Fix the first sliced image on the top, and move the remaining r - 1 sliced images upward by l, 2×l, ……, (r - 1)×l pixels in sequence along the horizontal direction to generate a downstairs walking sequence D gd .

[0065] The embodiments of the present invention perform data augmentation on the original gait silhouette sequence . First, perform slicing and translation operations on each gait silhouette image in the original gait silhouette sequence D

[0066] Specifically, for each gait silhouette image, the image is evenly divided into r parts in the horizontal dimension (the height dimension of the gait silhouette image) of the gait silhouette image. Fix the position of the top local image block, and move the remaining r - 1 parts to the right by l, 2×l, ……, (r - 1)×l pixels in sequence along the horizontal direction to generate a gait image similar to forward leaning walking, see Figure 3 . Move the remaining r - 1 parts to the left by l, 2×l, ……, (r - 1)×l pixels in sequence along the horizontal direction to generate a gait image similar to backward leaning walking, see Figure 3 .

[0067] For each gait silhouette image, the image is evenly divided into r parts in the vertical dimension (the width dimension of the gait silhouette image) of the gait silhouette image. Fix the position of the leftmost local image block, and move the remaining r - 1 parts downward by l, 2×l, ……, (r - 1)×l pixels in sequence along the height direction to generate a gait image similar to upstairs walking, see Figure 3 . Move the remaining r - 1 parts upward by l, 2×l, ……, (r - 1)×l pixels in sequence along the height direction to generate a gait image similar to downstairs walking, see Figure 3 .

[0068] Perform cropping and filling operations on the images obtained by performing slicing and translation operations on each gait silhouette image in the original gait silhouette sequence D, and four types of gait sequences can be generated: forward leaning walking sequence D lf , backward leaning walking sequence D bb , upstairs walking sequence Dgu and the downstairs walking sequence D gd 。

[0069] In an embodiment of the present invention, a motion pattern augmentation strategy is adopted to augment the original gait silhouette sequence D to generate four types of new walking data, providing diverse motion appearance information and pedestrian identity clues for the gait recognition network. And combining the gait features of the augmented data to enhance the discriminative ability of the gait features.

[0070] In an embodiment of the present invention, as Figure 4 and Figure 5 shown, the original gait silhouette sequence and the augmented gait silhouette sequence are input into the gait recognition model to obtain the target gait features and their corresponding classification probabilities, which may include:

[0071] S201, input the original gait silhouette sequence into the first branch network to obtain the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature;

[0072] S202, input the augmented gait silhouette sequence into the second branch network to obtain the augmented gait feature;

[0073] S203, input the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the augmented gait feature into the multi-stage feature aggregation network to obtain the first fusion feature;

[0074] S204, input the original gait feature and the first fusion feature into the element-wise feature summation layer to obtain the second fusion feature;

[0075] S205, input the second fusion feature into the fully connected layer and the classifier in sequence to obtain the target gait features and their corresponding classification probabilities.

[0076] Specifically, the first branch network extracts features from the original gait silhouette sequence D. The first branch network (spatiotemporal feature extraction network) in an embodiment of the present invention may include a first feature extraction module, a second feature extraction module, a first time aggregation module, a third feature extraction module, a time pooling module, and a space pooling module connected in sequence. The first feature extraction module outputs the first-layer fusion feature The first time aggregation module outputs the first aggregation feature The third feature extraction module outputs the third-layer fusion feature The space pooling module outputs the original gait feature F sp 。

[0077] The second branch network processes the augmented gait silhouette sequence corresponding to the original gait silhouette sequence D (the forward walking sequence D lf , the backward walking sequence D bb , the upstairs walking sequence D guand the downstairs walking sequence D gd ) perform feature extraction to obtain the augmented gait feature F da .

[0078] The multi-stage feature aggregation network processes the first-layer fusion feature The first aggregation feature The third-layer fusion feature and the augmented gait feature F da to perform fusion, obtaining the first fusion feature F ms .

[0079] The first fusion feature F ms and the original gait feature F sp perform an element-wise feature summation operation to obtain the final fusion feature (the second fusion feature) The final fusion feature F fusion successively passes through a fully connected layer and a classifier, respectively outputting the target gait feature F fc and the classification probability P.

[0080] In a specific embodiment of the present invention, the classifier may be composed of a fully connected layer and a batch normalization layer.

[0081] In an embodiment of the present invention, as Figure 4 and Figure 6 shown, the first branch network includes a first feature extraction module, a second feature extraction module, a first time aggregation module, a third feature extraction module, a time pooling module, and a space pooling module connected in sequence. Inputting the original gait silhouette sequence into the first branch network, obtaining the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature may include:

[0082] S301, using the first feature extraction module, respectively perform dynamic change feature extraction and multi-granularity feature extraction on the original gait silhouette sequence, and perform an element-wise feature summation operation on the extracted first dynamic change feature and the first multi-granularity feature to obtain the first-layer fusion feature;

[0083] S302, using the second feature extraction module, respectively perform dynamic change feature extraction and multi-granularity feature extraction on the first-layer fusion feature, and perform an element-wise feature summation operation on the extracted second dynamic change feature and the second multi-granularity feature to obtain the second-layer fusion feature;

[0084] S303, using the first time aggregation module, perform a compression aggregation operation on the second-layer fusion feature in the time dimension to obtain the first aggregation feature;

[0085] S304. Using the third feature extraction module, perform dynamic change feature extraction and multi-granularity feature extraction on the first aggregated feature respectively, and perform an element-wise feature summation operation on the extracted third dynamic change feature and third multi-granularity feature to obtain the third-layer fusion feature;

[0086] S305. Use the time pooling module to perform feature pooling operation on the third-layer fusion feature in the time dimension to obtain the pooled feature;

[0087] S306. Use the spatial pooling module to perform feature compression on the pooled feature in the spatial dimension to obtain the original gait feature.

[0088] In this embodiment, the first feature extraction module may include a first-layer dynamic change perception unit, a first-layer multi-granularity feature extraction unit, and a first-layer element-wise feature summation unit. The first-layer dynamic change perception unit and the first-layer multi-granularity feature extraction unit respectively extract dynamic change features and multi-granularity features from the original gait silhouette sequence D. The first dynamic change feature and the

[0089] first multi-granularity feature After an element-wise feature summation operation, generate the first-layer fusion feature and multi-granularity features The second dynamic change feature and the

[0090] second multi-granularity feature After an element-wise feature summation operation, generate the second-layer fusion feature

[0091] In a specific embodiment of the present invention, the time aggregation module is composed of a 3D convolution with a convolution kernel of (3, 3, 3) and a stride of (3, 1, 1), and the activation function LeakyReLU.

[0092] In this embodiment, the third feature extraction module may include a third-layer dynamic change perception unit, a third-layer multi-granularity feature extraction unit, and a third-layer element-wise feature summation unit. The third-layer dynamic change perception unit and the third-layer multi-granularity feature extraction unit perform operations on the first aggregated feature Extract dynamic change features separately and multi-granularity features The third dynamic change feature and the third multi-granularity feature Through the element-wise feature summation operation, generate the third-layer fusion feature

[0093] In this embodiment, the time pooling module performs a feature pooling operation on the third-layer fusion feature to generate a pooled feature The spatial pooling module performs a feature compression operation on the pooled feature F tp to generate the original gait feature

[0094] In a specific embodiment of the present invention, the time pooling operation may refer to a 3D max pooling layer. The spatial pooling operation can be expressed as: where represents a 3D average pooling layer with a pooling kernel of (1, 1, W), and p is a learnable network parameter

[0095] In a specific embodiment of the present invention, as Figure 7 shown, the dynamic change perception unit is composed of a global feature extraction branch and a dynamic difference feature extraction branch

[0096] In this embodiment, the global feature extraction branch is composed of a local feature extractor consisting of a 3D convolution with a convolution kernel of (3, 3, 3) and an activation function LeakyReLU. This branch extracts features from the input features of the dynamic change perception unit to obtain the feature

[0097] In this embodiment, the dynamic difference feature extraction branch is also composed of a 3D convolution with a convolution kernel of (3, 3, 3) and an activation function LeakyReLU. This branch first calculates the difference between adjacent frame features of the input features of T in frames, and a difference sequence of length T in -1 frames can be obtained. For the convenience of subsequent calculations, here, by copying the difference feature of the T in -1th frame, a difference feature sequence of length T in frames can be obtained Then, use 3D convolution and activation function to extract features from it to obtain the feature

[0098] The dynamic change perception unit passes the feature F global and the feature F difPerforming an element-wise feature summation operation can obtain the output feature F of this module dvp .

[0099] In a specific embodiment of the present invention, as Figure 8 shown, the multi-granularity feature extraction unit consists of multiple local feature extractors with non-shared parameters. The multi-granularity feature extraction unit first divides the input feature along the height dimension into multiple local features according to the prior knowledge of human body structure. For the multi-granularity feature extraction units of different layers, the number of divided local features is set as follows: the first-layer multi-granularity feature extraction unit contains 3 local features, the second-layer multi-granularity feature extraction unit contains 5 local features, and the third-layer multi-granularity feature extraction unit contains 9 local features.

[0100] In the first-layer multi-granularity feature extraction unit, there are a total of 3 groups of local feature extractors composed of 3D convolution with a convolution kernel of (3, 3, 3) and the activation function LeakyReLU. Each input feature After being sliced along the height dimension, 3 local features and can be obtained. Then, these local features pass through the local feature extractors to obtain the corresponding locally convolved features and Subsequently, and After passing through the feature concatenation operation along the height dimension again, the output feature of the first-layer multi-granularity feature extraction unit is obtained In the second-layer multi-granularity feature extraction unit and the third-layer multi-granularity feature extraction unit, there are 5 groups and 9 groups of local feature extractors respectively for extracting local features.

[0101] As a specific example, in the three-layer multi-granularity feature extraction unit, the first layer divides the input feature according to the height of and into 3 parts of local features; the second layer divides the input feature according to the height of and into 5 parts of local features; the third layer divides the input feature according to the height of and into 9 parts of local features.

[0102] For each layer of the multi-granularity feature extraction unit, first divide the input feature along the height dimension into multiple local features, then use the local feature extractors to perform convolution and non-linear activation operations on each local feature, and then perform a feature concatenation operation on the output features of each local feature extractor along the height dimension to obtain the output feature of the multi-granularity feature extraction unit.

[0103] In an embodiment of the present invention, in the first branch network, a dynamic change perception unit is used to calculate the dynamic changes between adjacent frames in the gait sequence, and the motion features of the dynamic region are extracted from the dynamic change sequence; a multi-granularity feature extraction unit is used to divide the input features into multiple local blocks along the height dimension according to human prior knowledge in different network layers, and extract the local spatio-temporal features respectively. As the network depth increases, the number of local feature divisions in the multi-granularity feature extraction unit gradually increases, so as to obtain local gait features including coarse-grained and fine-grained. The dynamic change perception unit and the multi-granularity feature extraction unit can provide key dynamic differential change features and multi-granularity local gait representations, which can effectively improve the representation ability of the model.

[0104] In an embodiment of the present invention, as Figure 4 and Figure 9 shown, the augmented gait silhouette sequence includes at least two of the forward walking sequence, the backward walking sequence, the upstairs walking sequence, and the downstairs walking sequence. The second branch network includes a shallow feature extraction module and a channel-level feature aggregation module. The augmented gait silhouette sequence is input into the second branch network to obtain augmented gait features, which may include:

[0105] S401, using the shallow feature extraction module to extract features from at least two of the forward walking sequence, the backward walking sequence, the upstairs walking sequence, and the downstairs walking sequence respectively, to obtain at least two of the first gait feature, the second gait feature, the third gait feature, and the fourth gait feature;

[0106] S402, using the channel-level feature aggregation module to perform feature aggregation on at least two of the first gait feature, the second gait feature, the third gait feature, and the fourth gait feature to obtain augmented gait features.

[0107] The second branch network of the embodiment of the present invention may include a shallow feature extraction module and a channel-level feature aggregation module.

[0108] Specifically, a shallow feature extraction module is constructed. For the four types of augmented gait sequences: the forward walking sequence D lf , the backward walking sequence D bb , the upstairs walking sequence D gu , and the downstairs walking sequence D gd , the gait features are extracted respectively to obtain the first gait feature the second gait feature the third gait feature and the fourth gait feature Design a channel-level feature aggregation operation for the four types of features, the first gait feature F lf , the second gait feature F bb , the third gait feature Fgu and the fourth gait feature F gd perform feature aggregation to obtain augmented gait features

[0109] In a specific embodiment of the present invention, the shallow feature extractor consists of a 3D convolution with a convolution kernel of (3, 3, 3) and the activation function LeakyReLU.

[0110] In an embodiment of the present invention, as Figure 4 and Figure 10 shown, the multi-stage feature aggregation network includes a first feature aggregation module, a second temporal aggregation module, a second feature aggregation module, and a third feature aggregation module. Input the first-layer fusion feature, the first aggregated feature, the third-layer fusion feature, and the augmented gait feature into the multi-stage feature aggregation network to obtain the first fusion feature, which may include:

[0111] S501: Use the first feature aggregation module to perform a channel-level feature splicing operation on the augmented gait feature and the first-layer fusion feature, perform feature channel adjustment on the first spliced feature obtained from the splicing operation to obtain the first adjusted feature, and perform feature channel adjustment and feature compression processing on the first spliced feature obtained from the splicing operation to obtain the first compressed feature;

[0112] S502: Use the second temporal aggregation module to perform a compression aggregation operation on the first adjusted feature in the time dimension to obtain the second aggregated feature;

[0113] S503: Use the second feature aggregation module to perform a channel-level feature splicing operation on the first aggregated feature and the second aggregated feature, perform feature channel adjustment on the second spliced feature obtained from the splicing operation to obtain the second adjusted feature, and perform feature channel adjustment and feature compression processing on the first spliced feature obtained from the splicing operation to obtain the second compressed feature;

[0114] S504: Use the third feature aggregation module to perform a channel-level feature splicing operation on the second adjusted feature and the third-layer fusion feature, and perform feature channel adjustment and feature compression processing on the third spliced feature obtained from the splicing operation to obtain the third compressed feature;

[0115] S505: Perform a channel-level feature splicing operation on the first compressed feature, the second compressed feature, and the third compressed feature to obtain the first fusion feature.

[0116] In the embodiment of the present invention, the multi-stage feature aggregation network is used for the augmented gait feature F da , the first-layer fusion feature the first aggregated feature and the third-layer fusion feature for fusion.

[0117] The multi-stage feature aggregation network in the embodiment of the present invention may include a first feature aggregation module, a second temporal aggregation module, a second feature aggregation module, and a third feature aggregation module. It should be noted that the structures and functions of the first temporal aggregation module and the second temporal aggregation module in the embodiment of the present invention are the same.

[0118] In this embodiment, the first feature aggregation module includes a channel-level feature splicing operation, a first-layer feature compression unit, and a first-layer channel mapping unit. As Figure 11 shown, the augmented gait feature F da (input feature 1) and the first-layer fusion feature (input feature 2) first pass through the channel-level feature splicing operation to obtain the first spliced feature Subsequently, the first spliced feature undergoes feature channel adjustment and feature compression through the first-layer feature compression unit. The first-layer feature compression unit includes a 3D convolution with a convolution kernel of (3, 3, 3), an activation function LeakyReLU, a temporal pooling layer, and a spatial pooling layer. The first spliced feature passes through the 3D convolution with a convolution kernel of (3, 3, 3) and the LeakyReLU activation function to obtain Then, the feature can be obtained through the temporal pooling layer, and the first compressed feature

[0119] In addition, the first spliced feature undergoes feature channel adjustment through the first-layer channel mapping unit. The channel mapping module includes a 3D convolution with a convolution kernel of (3, 3, 3) and LeakyReLU to obtain the first adjusted feature

[0120] The first adjusted feature obtains the second aggregated feature

[0121]

[0122] In this embodiment, the second feature aggregation module includes a channel-level feature splicing operation, a second-layer feature compression unit, and a second-layer channel mapping unit. As Figure 11 shown, the second aggregated feature (input feature 1) and the first aggregated feature (input feature 2) first pass through the channel-level feature splicing operation to obtain the second spliced feature The structure and function of the second-layer feature compression unit are the same as those of the above-mentioned first-layer feature compression unit. The second spliced feature The feature channels are adjusted and the features are compressed through the second-layer feature compression unit to obtain the second compressed feature The second concatenated feature The feature channels are adjusted through the second-layer channel mapping unit to obtain the second adjusted feature

[0123] In this embodiment, the third feature aggregation module includes a channel-level feature concatenation operation and a third-layer feature compression unit. The second adjusted feature (Input feature 1) and the third-layer fused feature (Input feature 2) Similarly, first through the channel-level feature concatenation operation, the third concatenated feature is obtained The third concatenated feature Only the feature channels need to be adjusted and the features compressed through the third-layer feature compression unit to obtain the third compressed feature

[0124] For the three compressed features generated in the multi-stage feature aggregation network: the first compressed feature The second concatenated feature and the third compressed feature Through the channel-level feature concatenation operation, the first fused feature is obtained Among them, by setting 3C fccm = C s3 , the first fused feature F with the same feature dimension can be obtained ms and the original gait feature F sp .

[0125] In a specific embodiment of the present invention, the feature compression unit in the feature aggregation module includes a 3D convolution with a convolution kernel of (3, 3, 3), an activation function LeakyReLU, a temporal pooling layer, and a spatial pooling layer. The channel mapping unit consists of a 3D convolution with a convolution kernel of (3, 3, 3) and LeakyReLU

[0126] The embodiment of the present invention realizes multi-stage feature fusion of the original gait feature and the augmented gait feature of the input original gait silhouette sequence by constructing a multi-stage feature aggregation network. The fused feature is sent to the subsequent multi-stage feature aggregation network through the channel mapping module, and the compressed features of each stage feature aggregation module are output through the feature compression module. The gait features of multiple stages are used to improve the model identity recognition accuracy

[0127] In an embodiment of the present invention, the expression of the total loss function is

[0128] L all = L triplet + L ce

[0129] Among them, L all represents the total loss function, and L triplet represents the triplet loss function, and L ce represents the cross-entropy loss function;

[0130]

[0131] Among them, K represents the number of triplets, and the triplet is (F pos , F neg , F anc ), where F pos represents the positive sample feature, F neg represents the negative sample feature, and F anc represents the anchor feature, is the Euclidean distance between the feature F pos and the feature F neg , is the Euclidean distance between the feature F pos and the feature F anc , 1 ≤ k ≤ K, max(·) represents the maximum value function, and margin represents the margin threshold;

[0132]

[0133] Among them, M represents the number of samples, 1 ≤ i ≤ M, and y i represents the label of the i-th sample in one-hot encoding type, and p i represents the classification result of the i-th sample.

[0134] In the embodiment of the present invention, in the triplet loss function L triplet , given a feature triplet (F pos , F neg , F anc ), where F pos represents the positive sample feature, F neg represents the negative sample feature, and F anc represents the anchor feature. Under the condition of given K groups of feature triplets, the triplet loss can be expressed as:

[0135]

[0136] Specifically, the gradient descent method is used to train the gait recognition model with motion pattern augmentation, and the total loss function L all is calculated to update the model parameters. When the number of training iterations reaches the set number or the total loss function L all converges, the training stops, so as to obtain a trained gait recognition model. This model is used to identify the identity of the gait silhouette sequence to obtain the identity recognition result.

[0137] It is feasible to adopt the Adam optimizer and perform a total of 100,000 iterations of training. The initial learning rate is set to 0.0001, and when the training iteration reaches 80,000 times, the learning rate is reduced to 0.00001.

[0138] The augmented gait silhouette sequence in the embodiments of the present invention provides diverse motion appearance information and pedestrian identity clues for the gait recognition model. The gait recognition model extracts features based on the original gait silhouette sequence and its corresponding augmented gait silhouette sequence. The original gait features are combined with the gait features of the augmented gait silhouette sequence (augmented gait features) to enhance the discriminative ability of the gait features of the gait recognition model.

[0139] The first branch network in the embodiments of the present invention uses a dynamic change perception unit to calculate the dynamic changes between adjacent frames in the gait sequence and extracts the motion features of the dynamic regions from the dynamic change sequence; uses a multi-granularity feature extraction unit. In different network layers, according to human prior knowledge, the input features are divided into multiple local blocks along the height dimension, and the local spatio-temporal features of each are extracted. As the network depth increases, the number of local feature partitions in the multi-granularity feature extraction unit gradually increases, thereby obtaining local gait features including coarse-grained and fine-grained features. The dynamic change perception unit and the multi-granularity feature extraction unit can provide key dynamic differential change features and multi-granularity local gait representations, which can effectively improve the representation ability of the model.

[0140] The multi-stage feature aggregation network in the embodiments of the present invention realizes the multi-stage feature fusion of the original gait features and the augmented gait features of the input original gait silhouette sequence. The fused features are sent to the subsequent multi-stage feature aggregation network through a channel mapping module, and the compressed features of each stage feature aggregation module are output through a feature compression module. By using the gait features of multiple stages, the accuracy of model identity recognition is improved.

[0141] The training method of the gait recognition model with augmented motion patterns in the embodiments of the present invention amplifies the original gait silhouette sequence, adds an augmented gait silhouette sequence, provides diverse motion appearance information and pedestrian identity clues for the gait recognition model, and improves the accuracy of model identity recognition by performing feature extraction and fusion on the original gait silhouette sequence and the augmented gait features.

[0142] The present invention provides a gait recognition method with augmented motion patterns.

[0143] Figure 12 It is a flowchart of the gait recognition method according to an embodiment of the present invention. As Figure 12 shown, the gait recognition method with augmented motion patterns may include:

[0144] S601. Obtain the original gait silhouette sequence to be measured, and amplify the gait silhouette sequence to be measured to obtain the augmented gait silhouette sequence to be measured;

[0145] S602. Input the original gait silhouette sequence to be measured and the augmented gait silhouette sequence to be measured into a pre-trained gait recognition model to obtain the target gait features and their corresponding classification probabilities, where the pre-trained gait recognition model is obtained by using the training method of the gait recognition model as described above.

[0146] The gait recognition model in the embodiment of the present invention may include a first branch network, a second branch network, a multi-stage feature aggregation network, an element-level feature summation layer, a fully connected layer, and a classifier. Among them, the first branch network is used to input the original gait silhouette sequence, the second branch network is used to input the augmented gait silhouette sequence, the output ends of the first branch network and the second branch network are connected to the input end of the multi-stage feature aggregation network, the output end of the first branch network and the output end of the multi-stage feature aggregation network are connected to the input end of the element-level feature summation layer, the output end of the element-level feature summation layer is connected to the input end of the classifier through the fully connected layer, and the classifier is used to output the target gait features and their corresponding classification probabilities.

[0147] The gait recognition method with motion pattern augmentation in the embodiment of the present invention performs data augmentation on the input original gait silhouette sequence to be measured, and generates gait sequences under four different walking patterns through a motion pattern augmentation strategy. The input original gait silhouette sequence to be measured is used to obtain the original gait features through the first branch network; the four gait sequences generated by data augmentation are used to obtain the augmented gait features through the second branch network. The original gait features and the augmented gait features are used to obtain the fused gait features through the multi-stage feature aggregation network. The original gait features and the fused gait are used to obtain the final gait features through a feature fusion operation.

[0148] The trained gait recognition model with motion pattern augmentation in the embodiment of the present invention mines effective motion appearance features and discriminative identity features from diverse motion performance forms, and combines the first branch network, the second branch network, and the multi-stage feature aggregation network to improve the accuracy of gait recognition.

[0149] The gait recognition method with motion pattern augmentation in the embodiment of the present invention mines effective motion appearance features and discriminative identity features from diverse motion performance forms, and uses the trained gait recognition model to recognize the original gait silhouette sequence to be measured and the augmented gait silhouette sequence to be measured, thereby improving the accuracy of gait recognition.

[0150] The present invention provides a computer-readable storage medium.

[0151] In one embodiment, a computer program is stored on a computer-readable storage medium. When the computer program is executed by a processor, the training method of the gait recognition model with augmented motion patterns as described above is implemented.

[0152] In another embodiment, a computer program is stored on a computer-readable storage medium. When the computer program is executed by a processor, the gait recognition method with augmented motion patterns as described above is implemented.

[0153] The present invention provides a controller.

[0154] In one embodiment, the controller includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the training method of the gait recognition model with augmented motion patterns as described above is implemented.

[0155] In one embodiment, the controller includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the gait recognition method with augmented motion patterns as described above is implemented.

[0156] Figure 13 It is a structural block diagram of the controller according to an embodiment of the present invention.

[0157] As Figure 13 shown, the controller 500 includes: a processor 501 and a memory 503. Among them, the processor 501 and the memory 503 are connected, such as connected through a bus 502. Optionally, the controller 500 may further include a transceiver 504. It should be noted that in practical applications, the transceiver 504 is not limited to one, and the structure of the controller 500 does not constitute a limitation to the embodiments of the present invention.

[0158] The processor 501 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present invention. The processor 501 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0159] The bus 502 may include a path for transmitting information between the above components. The bus 502 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 502 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 13 it is only represented by a thick line in Figure 13 , but it does not mean that there is only one bus or one type of bus.

[0160] The memory 503 is used to store the training method of the gait recognition model with motion mode augmentation in the above embodiments of the present invention, or a computer program corresponding to the gait recognition method with motion mode augmentation. This computer program is controlled and executed by the processor 501. The processor 501 is used to execute the computer program stored in the memory 503 to implement the content shown in the foregoing method embodiments.

[0161] Among them, the controller 500 includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 13 The shown controller 500 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0162] The computer-readable storage medium and the controller of the embodiments of the present invention use the above training method of the gait recognition model with motion mode augmentation to perform on the above gait recognition model, so as to improve the identity recognition accuracy of the trained gait recognition model.

[0163] The computer-readable storage medium and the controller of the embodiments of the present invention use the above gait recognition method with motion mode augmentation to mine effective motion appearance features and discriminative identity features from diverse motion performance forms, and use the trained gait recognition model to identify the to-be-tested original gait silhouette sequence and the to-be-tested augmented gait silhouette sequence, so as to improve the accuracy of gait recognition.

[0164] It should be noted that the logic and / or steps represented in the flowchart or described otherwise herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions), or in combination with these instruction execution systems, apparatuses or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in combination with an instruction execution system, apparatus or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting or otherwise processing it as appropriate, and then storing it in a computer memory.

[0165] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0166] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0167] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.

[0168] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0169] In the present invention, unless otherwise clearly defined and limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0170] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0171] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as a limitation on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A training method for a gait recognition model augmented with motion patterns, characterized in that, The gait recognition model includes a first branch network, a second branch network, a multi-stage feature aggregation network, an element-wise feature summation layer, a fully connected layer, and a classifier. Among them, the first branch network is used to input the original gait silhouette sequence, the second branch network is used to input the augmented gait silhouette sequence. The output ends of the first branch network and the second branch network are connected to the input end of the multi-stage feature aggregation network. The output end of the first branch network and the output end of the multi-stage feature aggregation network are connected to the input end of the element-wise feature summation layer. The output end of the element-wise feature summation layer is connected to the input end of the classifier through the fully connected layer. The classifier is used to output the target gait feature and its corresponding classification probability. The training method includes: Using the training image set to perform multiple periodic trainings on the gait recognition model. Among them, for each training cycle, the following operations are performed: Obtain the training image set, where the training image set includes the original gait silhouette sequence and its corresponding augmented gait silhouette sequence; Input the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain the target gait feature and its corresponding classification probability; Obtain the triplet loss function according to the target gait feature, obtain the cross-entropy loss function according to the classification probability corresponding to the target gait feature, and obtain the total loss function according to the triplet loss function and the cross-entropy loss function; Adjust the parameters of the gait recognition model according to the total loss function to obtain the trained gait recognition model; Among them, obtaining the augmented gait silhouette sequence includes: Obtain the original gait silhouette sequence; For each gait silhouette image in the original gait silhouette sequence, perform average slicing on each image in the height direction of the gait silhouette image to obtain r sliced images. Fix the first sliced image at the top, and shift the remaining r - 1 sliced images to the right by l, 2×l,..., (r - 1)×l pixels in sequence along the width direction of the gait silhouette image to generate a forward walking sequence; or, For each gait silhouette image in the original gait silhouette sequence, perform average slicing on each image in the height direction of the gait silhouette image to obtain r sliced images. Fix the first sliced image at the top, and shift the remaining r - 1 sliced images to the left by l, 2×l,..., (r - 1)×l pixels in sequence along the width direction of the gait silhouette image to generate a backward walking sequence; or, For each gait silhouette image in the original gait silhouette sequence, perform average slicing on each image in the width direction of the gait silhouette image to obtain r sliced images. Fix the first sliced image at the leftmost side, and shift the remaining r - 1 sliced images downward by l, 2×l,..., (r - 1)×l pixels in sequence along the height direction of the gait silhouette image to generate an up-stair walking sequence; or, For each gait silhouette image in the original gait silhouette sequence, each image in the gait silhouette image is evenly divided in the width direction of the gait silhouette image to obtain r sliced images. The first sliced image at the top is fixed, and the remaining r - 1 sliced images are successively shifted upward by l, 2×l, ……, (r - 1)×l pixels along the height direction of the gait silhouette image to generate a downstairs walking sequence.

2. The training method of the gait recognition model according to claim 1, wherein Inputting the original gait silhouette sequence and the augmented gait silhouette sequence into the gait recognition model to obtain the target gait feature and its corresponding classification probability includes: Inputting the original gait silhouette sequence into the first branch network to obtain the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature; Inputting the augmented gait silhouette sequence into the second branch network to obtain the augmented gait feature; Inputting the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the augmented gait feature into the multi-stage feature aggregation network to obtain the first fusion feature; Inputting the original gait feature and the first fusion feature into the element-level feature summation layer to obtain the second fusion feature; Successively inputting the second fusion feature into the fully connected layer and the classifier to obtain the target gait feature and its corresponding classification probability.

3. The training method of the gait recognition model according to claim 2, characterized in that The first branch network includes a first feature extraction module, a second feature extraction module, a first time aggregation module, a third feature extraction module, a time pooling module, and a space pooling module connected in sequence. Inputting the original gait silhouette sequence into the first branch network to obtain the first-layer fusion feature, the first aggregation feature, the third-layer fusion feature, and the original gait feature includes: Using the first feature extraction module to respectively extract the dynamic change feature and the multi-granularity feature from the original gait silhouette sequence, and performing an element-level feature summation operation on the extracted first dynamic change feature and the first multi-granularity feature to obtain the first-layer fusion feature; Using the second feature extraction module to respectively extract the dynamic change feature and the multi-granularity feature from the first-layer fusion feature, and performing an element-level feature summation operation on the extracted second dynamic change feature and the second multi-granularity feature to obtain the second-layer fusion feature; Using the first time aggregation module to perform a compression aggregation operation on the second-layer fusion feature in the time dimension to obtain the first aggregation feature; Using the third feature extraction module to respectively extract the dynamic change feature and the multi-granularity feature from the first aggregation feature, and performing an element-level feature summation operation on the extracted third dynamic change feature and the third multi-granularity feature to obtain the third-layer fusion feature; Using the time pooling module to perform a feature pooling operation on the third-layer fusion feature in the time dimension to obtain the pooled feature; Using the space pooling module to perform feature compression on the pooled feature in the space dimension to obtain the original gait feature.

4. The training method of the gait recognition model according to claim 2, wherein The augmented gait silhouette sequence includes at least two of a forward walking sequence, a backward walking sequence, an upstairs walking sequence, and a downstairs walking sequence. The second branch network includes a shallow feature extraction module and a channel-level feature aggregation module. Inputting the augmented gait silhouette sequence into the second branch network to obtain augmented gait features includes: Using the shallow feature extraction module to respectively extract features from at least two of the forward walking sequence, the backward walking sequence, the upstairs walking sequence, and the downstairs walking sequence, to obtain at least two of a first gait feature, a second gait feature, a third gait feature, and a fourth gait feature; Using the channel-level feature aggregation module to perform feature aggregation on at least two of the first gait feature, the second gait feature, the third gait feature, and the fourth gait feature, to obtain the augmented gait features.

5. The training method of the gait recognition model according to claim 2, wherein The multi-stage feature aggregation network includes a first feature aggregation module, a second temporal aggregation module, a second feature aggregation module, and a third feature aggregation module. Inputting the first-layer fusion feature, the first aggregated feature, the third-layer fusion feature, and the augmented gait features into the multi-stage feature aggregation network to obtain a first fusion feature includes: Using the first feature aggregation module to perform a channel-level feature splicing operation on the augmented gait features and the first-layer fusion feature, performing feature channel adjustment on the first spliced feature obtained by the splicing operation to obtain a first adjusted feature, and performing feature channel adjustment and feature compression processing on the first spliced feature obtained by the splicing operation to obtain a first compressed feature; Using the second temporal aggregation module to perform a compression aggregation operation on the first adjusted feature in the temporal dimension to obtain a second aggregated feature; Using the second feature aggregation module to perform a channel-level feature splicing operation on the first aggregated feature and the second aggregated feature, performing feature channel adjustment on the second spliced feature obtained by the splicing operation to obtain a second adjusted feature, and performing feature channel adjustment and feature compression processing on the second spliced feature obtained by the splicing operation to obtain a second compressed feature; Using the third feature aggregation module to perform a channel-level feature splicing operation on the second adjusted feature and the third-layer fusion feature, and performing feature channel adjustment and feature compression processing on the third spliced feature obtained by the splicing operation to obtain a third compressed feature; Performing a channel-level feature splicing operation on the first compressed feature, the second compressed feature, and the third compressed feature to obtain a first fusion feature.

6. The training method of the gait recognition model according to claim 1, characterized in that The expression of the total loss function is: Among them, represents the total loss function, represents the triplet loss function, represents the cross-entropy loss function; Among them, represents the number of triples, and the triple is , represents the positive sample feature, represents the negative sample feature, represents the anchor feature, is the feature and the feature the Euclidean distance between them, is the feature and the feature the Euclidean distance between them, , represents the maximum value function, represents the boundary threshold; Among them, represents the number of samples, , represents the label of the one-hot encoded type of the th sample, represents the classification result of the th sample.

7. A gait recognition method with augmented motion patterns, characterized in that, The recognition method includes: Obtaining a to-be-tested original gait silhouette sequence, and performing augmentation on the to-be-tested original gait silhouette sequence to obtain a to-be-tested augmented gait silhouette sequence; Inputting the to-be-tested original gait silhouette sequence and the to-be-tested augmented gait silhouette sequence into a pre-trained gait recognition model to obtain target gait features and their corresponding classification probabilities, where the pre-trained gait recognition model is obtained by using the training method of the gait recognition model with motion pattern augmentation as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the training method of the gait recognition model with augmented motion patterns according to any one of claims 1-6, or the gait recognition method with augmented motion patterns according to claim 7.

9. A controller, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the computer program is executed by the processor, it implements the training method of the gait recognition model with augmented motion patterns according to any one of claims 1-6, or the gait recognition method with augmented motion patterns according to claim 7.

Citation Information

Patent Citations

  • Abnormal behavior identification method and device

    CN118570872A