Gait recognition method and system based on mask slice pyramid spatio-temporal feature fusion
Through the method of spatial and temporal feature fusion of mask slice pyramids, the problem of low gait recognition accuracy in complex environments in the existing technology is solved, and more efficient gait feature extraction and feature information fusion are achieved, which improves the recognition accuracy.
Patent Information
- Application Number
- CN202510103984.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-17
AI Technical Summary
When faced with complex environments such as viewing angle changes, occlusions and wear changes, existing gait recognition technologies are difficult to effectively extract fine-grained gait features and fuse feature information at different levels, resulting in low recognition accuracy and efficiency.
The method of spatial and temporal feature fusion of mask slice pyramids is adopted, and through dynamic mask slice feature extraction, multi-scale temporal feature extraction and pyramid adaptive feature fusion, fine-grained information of gait, model time features, and fuse multi-level features.
It significantly improves the accuracy of gait recognition in complex environments, enhances the ability to capture dynamic changes of gait and the efficiency of utilization of feature information.
Smart Images

Figure CN120164252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gait recognition, and in particular, to a gait recognition method and system for spatio-temporal feature fusion of mask slice pyramids. Background Art
[0002] Gait recognition is a technology for identity recognition by analyzing an individual's walking pattern in video images. Compared with traditional identity recognition methods, such as face recognition and fingerprint recognition, gait recognition can complete identity recognition at a long distance, without contact, and even at low resolution. Based on evaluating an individual's unique walking pattern, gait recognition is widely used in criminal investigations, mental health and other fields. However, this technology still faces many challenges in practical applications, mainly including factors such as perspective changes, occlusions, and clothing changes, which will affect the accuracy and efficiency of actual recognition.
[0003] Currently, many deep learning-based gait recognition methods have been proposed. In terms of spatial feature extraction, some methods use gait silhouettes as inputs and extract features through global feature extractors (such as global convolution and pooling operations). However, due to the small differences between the gait silhouettes of different individuals, this method has defects in extracting fine-grained gait features. In addition, local feature extraction methods extract local features of gaits through predefined region segmentation strategies, but such methods usually rely on fixed regions for feature extraction and are difficult to flexibly handle complex environments such as perspective changes, occlusions, and clothing changes. In terms of temporal feature extraction, most advanced gait recognition methods model the dynamic changes of gaits by extracting temporal features from gait sequences, and the local dynamic features in gaits contain important identity information. Single-time-scale methods often ignore the dynamic changes across time scales, while multi-time-scale strategies may lead to blurred resolutions of short-time features. In addition, most existing methods do not make full use of features at different levels in neural networks and fail to effectively fuse low-level and high-level feature information. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides a gait recognition method and system for spatio-temporal feature fusion of mask slice pyramids, taking into account the feature changes at different time scales, better extracting fine-grained gait features and fusing feature information at different levels, and effectively improving the accuracy of gait recognition in complex environments.
[0005] The present invention provides a gait recognition method for spatio-temporal feature fusion of mask slice pyramids, and the method includes: Performing data preprocessing operations on the input video image data; Performing spatio-temporal feature extraction processing on the video image data after the data preprocessing operation to obtain first-stage features; Perform dynamic mask slicing feature extraction processing on the first-stage features to obtain second-stage features; Perform multi-scale temporal feature extraction processing on the second-stage features to obtain third-stage features; Perform secondary dynamic mask slicing feature extraction processing on the third-stage features to obtain fourth-stage features; Perform pyramid adaptive feature fusion processing on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain gait features after fusion processing; Perform training optimization processing on the gait features after fusion processing.
[0006] Furthermore, the data preprocessing operation on the input video image data includes: Divide the input video image data into several gait contour sequences according to categories.
[0007] Furthermore, the spatio-temporal feature extraction processing on the video image data after data preprocessing operation to obtain first-stage features includes: Perform spatio-temporal feature extraction processing on the video image data after data preprocessing operation based on an initial 3D convolutional layer to obtain the first-stage features of the video image data after data preprocessing operation.
[0008] Furthermore, the dynamic mask slicing feature extraction processing on the first-stage features to obtain second-stage features includes: Perform max pooling processing and average pooling processing on the first-stage features to obtain the maximum value and mean value of the first-stage features in a preset time dimension; Calculate the difference between the maximum value and the mean value, and generate a first-stage dynamic mask based on the difference; Convert the first-stage features into important features and unimportant features based on the first-stage dynamic mask; Horizontally divide the important features and unimportant features of the first-stage features into several first-stage feature slices, and apply independent 3D convolutional layers to process the feature components of each first-stage feature slice; Aggregate and connect the outputs of the feature components of all first-stage feature slices, and perform batch normalization and activation processing on the aggregated and connected first-stage feature slices to obtain second-stage features.
[0009] Furthermore, the multi-scale temporal feature extraction processing on the second-stage features to obtain third-stage features includes: Horizontally divide the second-stage features into several second-stage feature slices, and perform max pooling on the feature components of the two time scales in each second-stage feature slice to respectively obtain the feature components of each second-stage feature slice at the first time scale and the second time scale; Apply independent 1D convolutional layers to the feature components of each second-stage feature slice at the first time scale and the second time scale respectively to generate corresponding second-stage attention weights; Normalize the second-stage attention weights between the first time scale and the second time scale to obtain normalized second-stage attention weights; Based on element-wise multiplication, apply the normalized second-stage attention weights to perform weighted summation on the feature components of each second-stage feature slice at the first time scale and the second time scale, and add them to the feature components at the first time scale to obtain the output features of each second-stage feature slice; Concatenate the output features of all second-stage feature slices in the height dimension to obtain third-stage features.
[0010] Further, the process of performing secondary dynamic mask slice feature extraction on the third-stage features to obtain fourth-stage features includes: Perform max pooling and average pooling on the third-stage features to obtain the maximum value and the mean value of the third-stage features in a preset time dimension; Calculate the difference between the maximum value and the mean value, and generate a third-stage dynamic mask based on the difference; Based on the third-stage dynamic mask, convert the third-stage features into important features and unimportant features; Horizontally divide the important features and unimportant features of the third-stage features into several third-stage feature slices, and apply independent 3D convolutional layers to the feature components of each third-stage feature slice respectively; Aggregate and connect the outputs of the feature components of all third-stage feature slices, and perform batch normalization and activation processing on the aggregated and connected third-stage feature slices to obtain fourth-stage features.
[0011] Further, the process of performing pyramid adaptive feature fusion on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain the fused gait features includes: Perform slice-level pooling on the first-stage features, second-stage features, third-stage features, and fourth-stage features respectively to obtain the first-stage features, second-stage features, third-stage features, and fourth-stage features after slice-level pooling; Perform adaptive fusion processing on the one-stage features, two-stage features, three-stage features, and four-stage features after slice-level pooling of the slices to obtain the gait features after fusion processing.
[0012] Further, the step of respectively performing slice-level pooling processing on the one-stage features, two-stage features, three-stage features, and four-stage features to obtain the one-stage features, two-stage features, three-stage features, and four-stage features after slice-level pooling processing includes: Horizontally divide the one-stage features, two-stage features, three-stage features, and four-stage features into several final-stage feature slices; Perform a max-pooling operation on each final-stage feature slice in a preset time dimension to extract the most important feature information of each final-stage feature slice; Perform feature compression processing on the most important feature information of each final-stage feature slice based on a fully connected layer; Based on a generalized mean pooling layer, perform global fusion on the most important feature information after feature compression processing of all final-stage feature slices, and splice them in the height dimension to generate the one-stage features, two-stage features, three-stage features, and four-stage features after slice-level pooling processing.
[0013] Further, the step of performing adaptive fusion processing on the one-stage features, two-stage features, three-stage features, and four-stage features after slice-level pooling processing to obtain the gait features after fusion processing includes: Splice the one-stage features, two-stage features, three-stage features, and four-stage features after slice-level pooling processing pairwise in the feature length dimension; Perform batch normalization on the pairwise-spliced features and apply a 1D convolutional layer for processing to generate an attention map; Perform activation processing on the attention map to obtain the normalized final-stage attention weights; Based on element-wise multiplication, apply the normalized final-stage attention weights to weight the pairwise-spliced features respectively to obtain the weighted features; Perform a residual connection on the weighted features and the normalized final-stage attention weights, and output the gait features after fusion processing.
[0014] The present invention also provides a gait recognition system for mask slice pyramid spatio-temporal feature fusion. The gait recognition system for mask slice pyramid spatio-temporal feature fusion is used to implement the above-mentioned gait recognition method for mask slice pyramid spatio-temporal feature fusion. The system includes: A data preprocessing module, which is used to perform data preprocessing operations on the input video image data; The first-stage feature generation module is used to perform spatio-temporal feature extraction on the video image data after data preprocessing operations to obtain first-stage features; The second-stage feature generation module is used to perform dynamic mask slice feature extraction on the first-stage features to obtain second-stage features; The third-stage feature generation module is used to perform multi-scale temporal feature extraction on the second-stage features to obtain third-stage features; The fourth-stage feature generation module is used to perform secondary dynamic mask slice feature extraction on the third-stage features to obtain fourth-stage features; The feature fusion processing module is used to perform pyramid adaptive feature fusion on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain the gait features after fusion processing; The training and optimization processing module is used to perform training and optimization processing on the gait features after fusion processing.
[0015] The present invention provides a gait recognition method and system based on masked slice pyramid spatio-temporal feature fusion. By adopting dynamic mask slice feature extraction, multi-scale temporal feature extraction, and pyramid adaptive feature fusion, improvements are achieved in capturing fine-grained gait information, modeling temporal features, and fusing multi-level features respectively. Among them, by adopting dynamic mask slice feature extraction, after dynamically generating important features and unimportant features, they are segmented into several feature slices, and independent 3D convolution operations are applied to each slice, effectively improving the performance of extracting dynamic features in complex environments; by adopting multi-scale temporal feature extraction, by aggregating long-term and short-term temporal features at different time scales, highly discriminative temporal features are effectively extracted, while paying attention to global temporal features and key features in short time series; by adopting pyramid adaptive feature fusion, including slice-level scaling processing and adaptive fusion processing, through multi-level feature fusion, combining the advantages of low-level features and high-level features, the spatial and temporal features of gait recognition are fully modeled from a global perspective, overall improving the accuracy of gait recognition and having high application value. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is the flowchart of the gait recognition method for spatio-temporal feature fusion of masked slice pyramid in the first embodiment of the present invention; Figure 2 It is the schematic structural diagram of the gait recognition method for spatio-temporal feature fusion of masked slice pyramid in the first embodiment of the present invention; Figure 3 It is the flowchart of the process for extracting dynamic masked slice features in the first embodiment of the present invention; Figure 4 It is the structural diagram of the dynamic masked slice feature module in the first embodiment of the present invention; Figure 5 It is the flowchart of the process for extracting multi-scale temporal features in the first embodiment of the present invention; Figure 6 It is the structural diagram of the multi-scale temporal attention module in the first embodiment of the present invention; Figure 7 It is the flowchart of the process for extracting secondary dynamic masked slice features in the first embodiment of the present invention; Figure 8 It is the flowchart of the process for pyramid adaptive feature fusion in the first embodiment of the present invention; Figure 9 It is the structural diagram of the pyramid adaptive feature fusion module in the first embodiment of the present invention; Figure 10 It is the flowchart of the process for slice-level pooling in the first embodiment of the present invention; Figure 11 It is the flowchart of the process for adaptive fusion in the first embodiment of the present invention; Figure 12 It is the structural diagram of the adaptive fusion module in the first embodiment of the present invention; Figure 13 It is the architecture diagram of the gait recognition system for spatio-temporal feature fusion of masked slice pyramid in the second embodiment of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0019] In the present invention, it should be understood that terms such as "including" or "having" are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and do not intend to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0020] In addition, it should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0021] Embodiment 1 Embodiment 1 of the present invention provides a gait recognition method for spatio-temporal feature fusion of masked slice pyramids, and the method includes: performing data preprocessing operations on the input video image data; performing spatio-temporal feature extraction processing on the video image data after the data preprocessing operations to obtain first-stage features; performing dynamic masked slice feature extraction processing on the first-stage features to obtain second-stage features; performing multi-scale temporal feature extraction processing on the second-stage features to obtain third-stage features; performing secondary dynamic masked slice feature extraction processing on the third-stage features to obtain fourth-stage features; performing pyramid adaptive feature fusion processing on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain gait features after the fusion processing; and performing training and optimization processing on the gait features after the fusion processing.
[0022] In an alternative implementation manner of this embodiment, as Figure 1 and Figure 2 shown, Figure 1 FIG. shows the flowchart of the gait recognition method for spatio-temporal feature fusion of masked slice pyramids in Embodiment 1 of the present invention, Figure 2 FIG. shows the structural schematic diagram of the gait recognition method for spatio-temporal feature fusion of masked slice pyramids in Embodiment 1 of the present invention, and includes the following steps: S101. Perform data preprocessing operations on the input video image data; In an alternative implementation manner of this embodiment, the input video image data is divided into several gait profile sequences according to categories.
[0023] Specifically, the input video image data is divided into several gait profile sequences according to the preset classification categories and the preset quantity of each category, and the sequence is denoted as , where is the sequence batch size, is the number of channels per frame of the sequence, is the sequence period, is the sequence height, is the sequence width.
[0024] S102. Perform spatio-temporal feature extraction processing on the video image data after the data preprocessing operations to obtain first-stage features; In an alternative implementation of this embodiment, spatio-temporal feature extraction processing is performed on the video image data after the data preprocessing operation based on the initial 3D convolutional layer to obtain the first-stage features of the video image data after the data preprocessing operation.
[0025] Specifically, a plurality of gait contour sequences formed from the video image data after the data preprocessing operation are input into the initial 3D convolutional layer for spatio-temporal feature extraction processing to extract the first-stage features of the video image data after the data preprocessing operation , and the expression is as follows: ; In the formula, is the first-stage feature, is a 3D convolutional operation with a kernel size of (3, 3, 3), a stride of (1, 1, 1), and a padding of (1, 1, 1), is the activation function.
[0026] S103. Perform dynamic mask slice feature extraction processing on the first-stage features to obtain second-stage features; In an alternative implementation of this embodiment, based on the constructed Dynamic Mask Slice Feature Module (DMSFM), the first-stage features extracted in step S102 are subjected to dynamic mask slice feature extraction processing to obtain second-stage features , and the expression is as follows:
[0027] Specifically, as Figure 3 and Figure 4 shown, Figure 3 shows the flowchart of performing dynamic mask slice feature extraction processing in Embodiment 1 of the present invention, Figure 4 shows the structural diagram of the dynamic mask slice feature module in Embodiment 1 of the present invention, including the following steps: S301. Perform max pooling processing and average pooling processing on the first-stage features to obtain the maximum value and the average value of the first-stage features in a preset time dimension; In an alternative implementation of this embodiment, given the input feature, that is, the first-stage feature , max pooling processing and average pooling processing are respectively used to calculate the maximum value and the average value of the first-stage feature in a preset time dimension.
[0028] S302. Calculate the difference between the maximum value and the average value, and generate a first-stage dynamic mask based on the difference; In an alternative implementation of this embodiment, calculate the first-stage features The maximum value and the mean value After calculating the difference, based on the difference After normalizing with the activation function, it is used as the basis for the first-stage dynamic mask The generation formula is as follows: ; In the formula, Is the first-stage dynamic mask.
[0029] S303. Based on the first-stage dynamic mask, transform the first-stage features into important features and unimportant features; In an alternative implementation of this embodiment, based on the first-stage dynamic mask , transform the first-stage features Into important features And unimportant features , and the transformation formula is as follows: ; ; In the formula, Represents element-wise multiplication.
[0030] S304. Horizontally divide the important features and unimportant features of the first-stage features into several first-stage feature slices, and apply independent 3D convolutional layers to process the feature components of each first-stage feature slice; In an alternative implementation of this embodiment, the important features Of the first-stage features And unimportant features Are horizontally divided into several first-stage feature slices, and independent 3D convolutional layers are applied to operate on the feature components of each first-stage feature slice. The expression is as follows: ; In the formula, Represents When, the important feature Of the Th slice, and When, the unimportant feature Of the Th slice, Represents the 3D convolutional operation of the Th slice.
[0031] S305. Aggregate and concatenate the outputs of the feature components of all the first-stage feature slices, and perform batch normalization and activation processing on the aggregated and concatenated first-stage feature slices to obtain second-stage features.
[0032] In an alternative implementation of this embodiment, on a preset high dimension aggregate and concatenate the outputs of the feature components of all the first-stage feature slices, including adding the respective parts of the important features and the unimportant features together, and then perform batch normalization processing and activation processing on all the aggregated and concatenated first-stage feature slices to obtain second-stage features , and the expression is as follows: ; In the formula, is the second-stage feature, is the batch normalization function, is the activation function, represents the feature concatenation operation on the preset high dimension dimension.
[0033] S104. Perform multi-scale temporal feature extraction processing on the second-stage features to obtain third-stage features; In an alternative implementation of this embodiment, based on the constructed Multi-scale Time Attention Module (MTAM), perform multi-scale temporal feature extraction processing on the second-stage features extracted in step S103 to obtain third-stage features , and the expression is as follows:
[0034] Specifically, as Figure 5 and Figure 6 show, Figure 5 shows the flowchart of performing multi-scale temporal feature extraction processing in Embodiment 1 of the present invention, Figure 6 shows the structural diagram of the multi-scale time attention module in Embodiment 1 of the present invention, including the following steps: S501. Horizontally divide the second-stage features into several second-stage feature slices, and perform max-pooling processing on the feature components of two time scales in each second-stage feature slice to respectively obtain the feature components of each second-stage feature slice at the first time scale and the second time scale; In an alternative implementation of this embodiment, given the input feature, that is, the second-stage feature , for the second-stage feature Horizontally divide it into several two-stage feature slices, and perform max pooling on the feature components of two time scales in each two-stage feature slice, where one of the two time scales is a short time scale and the other is a long time scale, and respectively obtain the feature components of each two-stage feature slice at the first time scale and the second time scale.
[0035] Specifically, the feature components of two time scales are fused in an element-wise summation manner, where the feature components of the first time scale are max-pooled with a kernel size of (3, 1, 1) and a stride of (3, 1, 1) to obtain the feature components of the first time scale , and the feature components of the second time scale are max-pooled with a sum size of (5, 1, 1) and a stride of (3, 1, 1) to obtain the feature components of the second time scale , where represents the th two-stage feature slice.
[0036] S502. Apply independent 1D convolutional layers to the feature components of each two-stage feature slice at the first time scale and the second time scale respectively to generate corresponding two-stage attention weights; In an alternative implementation of this embodiment, for the feature components of each two-stage feature slice at the first time scale and the feature components of the second time scale apply independent 1D convolutional layers respectively to generate corresponding two-stage attention weights , and the expression is as follows: ; In the formula, represents at time , the two-stage attention weight of the first time scale, and at time , the two-stage attention weight of the second time scale,
[0037] S503. Normalize the two-stage attention weights between the first time scale and the second time scale to obtain normalized two-stage attention weights; In an alternative implementation of this embodiment, for the two-stage attention weights the two-stage attention weight of the first time scale and the two-stage attention weight of the second time scale , are respectively applied to the feature components of the first time scale and the feature components of the second time scale between, through function for normalization to obtain the normalized two-stage attention weights.
[0038] S504. Based on element-wise multiplication, apply the normalized two-stage attention weights to perform weighted summation on the feature components of each two-stage feature slice at the first time scale and the second time scale, and add it to the feature components of the first time scale to obtain the output features of each two-stage feature slice; In an optional implementation manner of this embodiment, among the applied normalized two-stage attention weights, the two-stage attention weights of the first time scale perform element-wise multiplication on the feature components of the first time scale and apply the two-stage attention weights of the second time scale to perform element-wise multiplication on the feature components of the second time scale and then perform weighted summation on both of them and the feature components of the first time scale to obtain the output features of each two-stage feature slice , and the expression is as follows: ; In the formula, the output features of the represents element-wise multiplication, represents weighted summation.
[0039] S505. Concatenate the output features of all two-stage feature slices in the height dimension to obtain the three-stage features.
[0040] In an optional implementation manner of this embodiment, concatenate the output features of all two-stage feature slices of each sub-region in the height dimension to obtain the three-stage features , and the expression is as follows: ; In the formula, is the three-stage feature, represents the feature connection operation in the preset high dimension dimension.
[0041] S105. Perform quadratic dynamic mask slice feature extraction processing on the three-stage features to obtain the four-stage features; In an optional implementation manner of this embodiment, based on the constructed dynamic mask slice feature module, for the three-stage features extracted in step S104 Perform secondary dynamic mask slicing feature extraction processing to obtain four-stage features , and the expression is as follows:
[0042] Specifically, as Figure 7 shown, Figure 7 shows the flowchart of performing secondary dynamic mask slicing feature extraction processing in the first embodiment of the present invention, including the following steps: S701. Perform max pooling processing and average pooling processing on the three-stage features to obtain the maximum value and the average value of the three-stage features in a preset time dimension; In an optional implementation manner of this embodiment, given the input features, that is, the three-stage features , use max pooling processing and average pooling processing respectively to calculate the three-stage features in a preset time dimension for the maximum value and the average value .
[0043] S702. Calculate the difference between the maximum value and the average value, and generate a three-stage dynamic mask based on the difference; In an optional implementation manner of this embodiment, calculate the difference between the maximum value and the average value of the three-stage features in a preset time dimension, and then perform activation function normalization based on the difference as the basis for the three-stage dynamic mask , and the generation formula is as follows: ; In the formula, is the three-stage dynamic mask.
[0044] S703. Based on the three-stage dynamic mask, transform the three-stage features into important features and unimportant features; In an optional implementation manner of this embodiment, based on the three-stage dynamic mask , transform the three-stage features into important features and unimportant features , and the transformation formula is as follows: ; ; In the formula, represents element-wise multiplication.
[0045] S704. Horizontally divide the important and unimportant features of the three-stage features into several three-stage feature slices, and process the feature components of each three-stage feature slice by applying an independent 3D convolutional layer respectively; In an alternative implementation of this embodiment, the three-stage features of important features and unimportant features are horizontally divided into several three-stage feature slices, and the feature components of each three-stage feature slice are respectively operated on by applying an independent 3D convolutional layer. The expression is as follows: ; In the formula, represents when, the th slice of the important feature , and when, the th slice of the unimportant feature , represents the 3D convolution operation of the th slice.
[0046] S705. Aggregate and connect the outputs of the feature components of all three-stage feature slices, and perform batch normalization and activation processing on the aggregated and connected three-stage feature slices to obtain four-stage features.
[0047] In an alternative implementation of this embodiment, the outputs of the feature components of all three-stage feature slices are aggregated and connected in a preset high-dimensional dimension, including adding the respective parts of the important feature and the unimportant feature together, and then performing batch normalization processing and activation processing on all the aggregated and connected three-stage feature slices to obtain the four-stage feature . The expression is as follows: ; In the formula, is the four-stage feature, is the batch normalization function, is the activation function, represents the feature connection operation in the preset high-dimensional dimension.
[0048] S106. Perform pyramid adaptive feature fusion processing on the one-stage feature, two-stage feature, three-stage feature, and four-stage feature to obtain the gait feature after fusion processing; In an alternative implementation of this embodiment, based on the constructed Pyramid Adaptive Feature Fusion (PAFF) module, perform pyramid adaptive feature fusion processing on the first-stage features extracted in step S102, the second-stage features extracted in step S103, the third-stage features extracted in step S104, and the fourth-stage features extracted in step S105 to obtain the finally output gait features after fusion processing. The expression is as follows:
[0049] Specifically, as Figure 8 and Figure 9 shown, Figure 8 shows the flowchart of performing pyramid adaptive feature fusion processing in the first embodiment of the present invention, Figure 9 shows the structural diagram of the pyramid adaptive feature fusion module in the first embodiment of the present invention, including the following steps: S801. Respectively perform slice horizontal pooling processing on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain the first-stage features, second-stage features, third-stage features, and fourth-stage features after slice horizontal pooling processing; In an alternative implementation of this embodiment, based on the Slice Horizontal Pooling (SHP) module in the pyramid adaptive feature fusion module, first perform slice horizontal pooling processing on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain the first-stage features, second-stage features, third-stage features, and fourth-stage features after slice horizontal pooling processing.
[0050] Specifically, as Figure 10 shown, Figure 10 shows the flowchart of performing slice horizontal pooling processing in the first embodiment of the present invention, including the following steps: S1001. Horizontally divide the first-stage features, second-stage features, third-stage features, and fourth-stage features into several final-stage feature slices; In an alternative implementation of this embodiment, given the input feature , that is, the first-stage feature , the second-stage feature , the third-stage feature , and the fourth-stage feature , and horizontally divide them into several final-stage feature slices .
[0051] S1002. Perform a max pooling operation on each final-stage feature slice in the preset time dimension to extract the most important feature information of each final-stage feature slice; In an alternative implementation of this embodiment, for each final-stage feature slice perform a max pooling operation in a preset time dimension to extract the most important feature information of each final-stage feature slice .
[0052] It should be noted that the most important feature information is the feature information corresponding to the max pooling value.
[0053] S1003. Perform feature compression processing on the most important feature information of each final-stage feature slice based on a fully connected layer; In an alternative implementation of this embodiment, based on a fully connected layer perform feature compression processing on the most important feature information of each final-stage feature slice to obtain the most important feature information after feature compression processing of each final-stage feature slice .
[0054] S1004. Perform global fusion on the most important feature information after feature compression processing of all final-stage feature slices based on a generalized mean pooling layer, and perform splicing in the height dimension to generate stage-one features, stage-two features, stage-three features, and stage-four features after slice-level pooling processing.
[0055] In an alternative implementation of this embodiment, apply a generalized mean pooling layer (GeM) to perform global fusion on the most important feature information after feature compression processing of all final-stage feature slices .
[0056] In an alternative implementation of this embodiment, after performing global fusion, perform splicing in a preset height dimension to generate stage-one features, stage-two features, stage-three features, and stage-four features after slice-level pooling processing. The expression is as follows: ; In the formula, are the stage-one features, stage-two features, stage-three features, and stage-four features after slice-level pooling processing, represents a feature connection operation in a preset high-dimensional dimension, is the feature length, , corresponding to the stage-one, stage-two, stage-three, and stage-four features.
[0057] S802. Perform adaptive fusion processing on the stage-one features, stage-two features, stage-three features, and stage-four features after slice-level pooling processing to obtain gait features after fusion processing.
[0058] In an alternative implementation of this embodiment, based on the Adaptive Feature Fusion Module (AFM) in the pyramid adaptive feature fusion module, the one-stage feature, two-stage feature, three-stage feature, and four-stage feature after the slice horizontal pooling process are subjected to adaptive fusion processing to obtain the gait feature after the fusion processing, that is, they are combined using an adaptive fusion strategy. Each adaptive fusion module adaptively combines two input features by learning the fusion weights to highlight important information.
[0059] Specifically, as Figure 11 and 12 shown, Figure 11 shows the flowchart of the adaptive fusion processing in the first embodiment of the present invention, Figure 12 shows the structural diagram of the adaptive fusion module in the first embodiment of the present invention, including the following steps: S1101. Concatenate the one-stage feature, two-stage feature, three-stage feature, and four-stage feature after the slice horizontal pooling process pairwise in the feature length dimension; In an alternative implementation of this embodiment, the one-stage feature, two-stage feature, three-stage feature, and four-stage feature after the slice horizontal pooling process are concatenated pairwise in the feature length dimension and For example, for two input features and are concatenated, and the expression is as follows: ; In the formula, is the input feature after pairwise concatenation, is the feature concatenation operation of pairwise concatenation in the feature length dimension and .
[0060] S1102. Perform batch normalization on the pairwise concatenated features and apply a 1D convolutional layer for processing to generate an attention map; In an alternative implementation of this embodiment, perform batch normalization on the pairwise concatenated features and apply a 1D convolutional layer for processing to generate an attention map.
[0061] S1103. Perform activation processing on the attention map to obtain the normalized final-stage attention weights; In an alternative implementation of this embodiment, through the activation function performs activation processing on the attention map to obtain the normalized final-stage attention weights , the expression is as follows:
[0062] S1104. Based on element-wise multiplication, apply the normalized final-stage attention weights to the pairwise concatenated features respectively for weighted processing to obtain the weighted features; In an alternative implementation of this embodiment, based on element-wise multiplication, apply the normalized final-stage attention weights to the pairwise concatenated features respectively for weighted processing to obtain the weighted features , the expression is as follows: ; In the formula, represents the weighted processing of the element-wise multiplication operation.
[0063] S1105. Perform residual connection on the weighted features and the normalized final-stage attention weights, and output the gait features after fusion processing.
[0064] In an alternative implementation of this embodiment, perform residual connection on the weighted features and the original unweighted normalized final-stage attention weights to output the final gait features after fusion processing , the expression is as follows: ; In the formula, represents the residual connection.
[0065] S107. Perform training optimization processing on the gait features after fusion processing.
[0066] In an alternative implementation of this embodiment, the gait features F after fusion processing pass through a fully connected layer (FC) and batch normalization (BN), are mapped to the feature space, and are applied to the gait recognition task after training optimization processing.
[0067] Specifically, a combination of triplet loss and cross-entropy loss is used for training optimization processing, which can optimize the separability and classification performance of gait features simultaneously.
[0068] In an alternative implementation of this embodiment, on the widely used CASIA-B gait recognition dataset, comparative experiments were conducted with a variety of existing mainstream methods. The experiment adopted a large-scale sample evaluation protocol. The gait sequences of the first 74 subjects were used as the training set, and the gait sequences of the remaining 50 subjects were used as the test set. In the test phase, the NM#01-04 sequences were designated as the registration set for constructing the gait reference library, and the remaining NM#05-06, BG#01-02, and CL#01-02 sequences were used as the validation set to test the recognition performance under different scenarios.
[0069] In this experiment, the size of the gait silhouette was adjusted to 64×44 as the input. In the training phase, 30 frames were sampled from each video as the input, while in the test phase, all frames were used for inference. The batch size was set to , where represents the number of subjects, is the number of training samples for each subject in each batch. For the CASIA-B dataset, the optimizer used was Adam, with an initial learning rate of , and a momentum of 0.2. The batch size was set to and , trained for 120K iterations, and the learning rate was decayed to 0.1 of the original at the 70Kth iteration.
[0070] The main results are shown in Tables 1 and 2. Under the three conditions of NM, BG, and CL, the improvement compared to the leading GaitAMR method was 0.21%, 0.07%, and 2.9% respectively. Specifically, the recognition accuracies of this embodiment under these conditions reached 98.12%, 95.82%, and 89.59% respectively. Both CSTL and this embodiment extracted multi-scale spatio-temporal features, but their processing methods were different. CSTL fused the features of three time scales, namely frame-level, short-term, and long-term features. While the MTAM module proposed in this embodiment not only aggregated multi-scale spatio-temporal features but also particularly focused on the modeling of short-term features. The experimental results show that this embodiment can effectively extract discriminative gait features, especially outstanding under the most challenging CL condition. According to Table 2, the overall average Rank-1 accuracy of this embodiment on the CASIA-B dataset was 94.51%, an improvement of 1.06% compared to the 93.45% of the GaitAMR method. This improvement is mainly attributed to the following key factors: First, the dynamic mask slice feature module enhanced the expressive ability of gait detail features through an adaptive mask mechanism; second, the multi-scale temporal attention module improved the modeling accuracy of temporal features; finally, the pyramid adaptive feature fusion module fully integrated the feature information of each stage through the adaptive fusion of multi-level features, thus improving the overall recognition performance.
[0071] In addition, to evaluate the contribution of each module in this embodiment to gait recognition performance, we conducted a systematic ablation experiment on the CASIA-B dataset, and the results are shown in Table 3. The experiment first constructed a baseline method that only included the basic 3D convolutional layer. The average recognition rates of this method under the conditions of normal walking (NM), walking with a backpack (BG), and walking with a coat (CL) were 97.11%, 94.39%, and 84.66% respectively, and the overall average recognition rate was 92.05%. On this basis, after introducing the DMSFM module and the MTAM module respectively, the average recognition rates of the baseline method were increased to 92.93% and 92.88% respectively. Further, when the DMSFM and MTAM modules were applied jointly, the average recognition rate was increased to 93.73%. Finally, after adding the PAFF module, the average recognition rate of this embodiment reached 94.51%, with improvements in all walking conditions.
[0072] Table 1 Rank-1 accuracy (%) of the CASIA-B dataset at all viewpoints and different conditions
[0073] Table 2 Overall average rank-1 accuracy (%) on the CASIA-B dataset
[0074] Table 3 Accuracy (%) of different module combinations on the CASIA-B dataset
[0075] In summary, Embodiment 1 of the present invention provides a gait recognition method based on masked slice pyramid spatio-temporal feature fusion. By adopting dynamic masked slice feature extraction processing, multi-scale temporal feature extraction processing, and pyramid adaptive feature fusion processing, improvements are achieved in capturing fine-grained gait information, modeling temporal features, and fusing multi-level features respectively. Among them, in the dynamic masked slice feature extraction processing, after dynamically generating important features and unimportant features, they are segmented into several feature slices, and independent 3D convolution operations are applied to each slice, effectively improving the performance of extracting dynamic features in complex environments; in the multi-scale temporal feature extraction processing, by aggregating long-term and short-term temporal features of different time scales, highly discriminative temporal features are effectively extracted, while paying attention to global temporal features and key features in short time series; in the pyramid adaptive feature fusion processing, including slice-level scaling processing and adaptive fusion processing, through multi-level feature fusion, combining the advantages of low-level features and high-level features, the spatial and temporal features of gait recognition are fully modeled from a global perspective, and the overall accuracy of gait recognition is improved, with high application value.
[0076] Embodiment 2 Embodiment 2 of the present invention provides a gait recognition system for spatio-temporal feature fusion of masked slice pyramids. The gait recognition system for spatio-temporal feature fusion of masked slice pyramids is used to implement the gait recognition method for spatio-temporal feature fusion of masked slice pyramids described in Embodiment 1. The system includes a data preprocessing module, a first-stage feature generation module, a second-stage feature generation module, a third-stage feature generation module, a fourth-stage feature generation module, a feature fusion processing module, and a training optimization processing module.
[0077] In an alternative implementation of this embodiment, as Figure 13 shown Figure 13 Figure shows the architecture diagram of the gait recognition system for spatio-temporal feature fusion of masked slice pyramids in Embodiment 2 of the present invention, including the following modules: Data preprocessing module 10, which is used to perform data preprocessing operations on the input video image data; First-stage feature generation module 20, which is used to perform spatio-temporal feature extraction processing on the video image data after data preprocessing operations to obtain first-stage features; Second-stage feature generation module 30, which is used to perform dynamic masked slice feature extraction processing on the first-stage features to obtain second-stage features; Third-stage feature generation module 40, which is used to perform multi-scale temporal feature extraction processing on the second-stage features to obtain third-stage features; Fourth-stage feature generation module 50, which is used to perform secondary dynamic masked slice feature extraction processing on the third-stage features to obtain fourth-stage features; Feature fusion processing module 60, which is used to perform pyramid adaptive feature fusion processing on the first-stage features, second-stage features, third-stage features, and fourth-stage features to obtain gait features after fusion processing; Training optimization processing module 70, which is used to perform training optimization processing on the gait features after fusion processing.
[0078] In summary, Embodiment 2 of the present invention provides a gait recognition system for spatio-temporal feature fusion of a masked slice pyramid. The gait recognition system for spatio-temporal feature fusion of a masked slice pyramid is used to implement the gait recognition method for spatio-temporal feature fusion of a masked slice pyramid described in Embodiment 1. By adopting dynamic masked slice feature extraction processing, multi-scale temporal feature extraction processing, and pyramid adaptive feature fusion processing, improvements are achieved in capturing fine-grained gait information, modeling temporal features, and fusing multi-level features respectively. Among them, in the dynamic masked slice feature extraction processing, after dynamically generating important features and unimportant features, they are segmented into several feature slices, and an independent 3D convolution operation is applied to each slice, effectively improving the performance of extracting dynamic features in complex environments; in the multi-scale temporal feature extraction processing, by aggregating long-term and short-term temporal features of different time scales, highly discriminative temporal features are effectively extracted, while paying attention to global temporal features and key features in short time series; in the pyramid adaptive feature fusion processing, including slice level scaling processing and adaptive fusion processing, through multi-level feature fusion, combining the advantages of low-level features and high-level features, the spatial and temporal features of gait recognition are fully modeled from a global perspective, overall improving the accuracy of gait recognition, and having high application value.
[0079] The above has introduced in detail a gait recognition method and system for spatio-temporal feature fusion of a masked slice pyramid provided by the present invention. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0080] In addition, the above has introduced the embodiments of the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A gait recognition method based on mask slice pyramid spatiotemporal feature fusion, characterized in that: The method comprises: Performing data preprocessing operations on input video image data; Performing spatiotemporal feature extraction processing on the video image data after the data preprocessing operation to obtain first-stage features; Performing dynamic mask slicing feature extraction processing on the first-stage features to obtain second-stage features; Performing multi-scale time feature extraction processing on the second-stage features to obtain third-stage features; Performing secondary dynamic mask slicing feature extraction processing on the three-stage features to obtain four-stage features; Performing pyramid adaptive feature fusion processing on the first-stage features, the second-stage features, the third-stage features and the fourth-stage features to obtain gait features after fusion processing; The gait features after the fusion processing are trained and optimized.
2. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The data preprocessing operation on the input video image data includes: The input video image data is divided into several gait profile sequences according to categories.
3. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The performing of spatiotemporal feature extraction processing on the video image data after the data preprocessing operation to obtain the first-stage features includes: Based on the initial 3D convolutional layer, spatiotemporal feature extraction processing is performed on the video image data after the data preprocessing operation to obtain the first-stage features of the video image data after the data preprocessing operation.
4. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The step of performing dynamic mask slicing feature extraction processing on the first-stage features to obtain the second-stage features includes: Performing maximum pooling and average pooling on the first-stage features to obtain the maximum value and average value of the first-stage features in a preset time dimension; Calculating a difference between the maximum value and the mean value, and generating a first-stage dynamic mask based on the difference; Converting the first-stage features into important features and unimportant features based on the first-stage dynamic mask; The important features and the unimportant features of the first-stage features are horizontally divided into a plurality of first-stage feature slices, and the feature components of each first-stage feature slice are processed by an independent 3D convolution layer respectively; The outputs of the feature components of all the first-stage feature slices are aggregated and connected, and the first-stage feature slices after aggregation and connection are batch normalized and activated to obtain the second-stage features.
5. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The performing multi-scale time feature extraction processing on the second-stage features to obtain the third-stage features comprises: The two-stage feature level is divided into a plurality of two-stage feature slices, and the feature components of the two time scales in each two-stage feature slice are subjected to maximum pooling processing, so as to obtain the feature components of the first time scale and the feature components of the second time scale of each two-stage feature slice respectively; Apply independent 1D convolutional layers to process the feature components of the first time scale and the feature components of the second time scale of each two-stage feature slice, and generate the corresponding two-stage attention weights; Normalizing the two-stage attention weights between the first time scale and the second time scale to obtain normalized two-stage attention weights; Based on element-by-element multiplication, applying the normalized two-stage attention weights to weighted sum the feature components of the first time scale and the feature components of the second time scale of each two-stage feature slice, and adding them to the feature components of the first time scale to obtain output features of each two-stage feature slice; The output features of all second-stage feature slices are concatenated in the height dimension to obtain the third-stage features.
6. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The performing secondary dynamic mask slicing feature extraction processing on the three-stage features to obtain the four-stage features comprises: Performing maximum pooling and average pooling on the three-stage features to obtain the maximum value and average value of the three-stage features in a preset time dimension; Calculating a difference between the maximum value and the mean value, and generating a three-stage dynamic mask based on the difference; Converting the three-stage features into important features and unimportant features based on the three-stage dynamic mask; The important features and the unimportant features of the three-stage features are horizontally divided into a plurality of three-stage feature slices, and the feature components of each three-stage feature slice are processed by an independent 3D convolution layer respectively; The outputs of the feature components of all three-stage feature slices are aggregated and connected, and the three-stage feature slices after aggregation and connection are batch normalized and activated to obtain the four-stage features.
7. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 1, characterized in that: The step of performing pyramid adaptive feature fusion processing on the first-stage feature, the second-stage feature, the third-stage feature and the fourth-stage feature to obtain the gait feature after the fusion processing comprises: Performing slice-level pooling processing on the first-stage features, the second-stage features, the third-stage features, and the fourth-stage features respectively, to obtain the first-stage features, the second-stage features, the third-stage features, and the fourth-stage features after the slice-level pooling processing; The first-stage features, the second-stage features, the third-stage features and the fourth-stage features after the slice level pooling processing are adaptively fused to obtain the gait features after the fusion processing.
8. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 7, characterized in that: The first-stage features, the second-stage features, the third-stage features and the fourth-stage features are respectively subjected to slice-level pooling processing to obtain the first-stage features, the second-stage features, the third-stage features and the fourth-stage features after the slice-level pooling processing, including: Segmenting the first-stage feature, the second-stage feature, the third-stage feature, and the fourth-stage feature level into a plurality of final-stage feature slices; Perform a maximum pooling operation on each final stage feature slice in a preset time dimension to extract the most important feature information of each final stage feature slice; Based on the fully connected layer, the most important feature information of each feature slice in the final stage is compressed; Based on the generalized mean pooling layer, the most important feature information of all final stage feature slices after feature compression processing is globally fused and spliced in the height dimension to generate the first-stage features, second-stage features, third-stage features and fourth-stage features after slice level pooling processing.
9. The gait recognition method based on mask slice pyramid spatiotemporal feature fusion as claimed in claim 7, characterized in that: The adaptively fusing the first-stage features, the second-stage features, the third-stage features and the fourth-stage features after the slice level pooling processing to obtain the gait features after the fusion processing comprises: The first-stage features, second-stage features, third-stage features and fourth-stage features after the slice horizontal pooling processing are spliced in pairs in the feature length dimension; The concatenated features are batch normalized and processed using a 1D convolutional layer to generate an attention map. Activate the attention map to obtain a normalized final-stage attention weight; Based on element-by-element multiplication, the normalized final stage attention weights are applied to perform weighted processing on the concatenated features respectively to obtain weighted features; The weighted features are residually connected with the normalized final stage attention weights to output the fused gait features.
10. A gait recognition system based on mask slice pyramid spatiotemporal feature fusion, characterized in that: The masked slice pyramid spatiotemporal feature fusion gait recognition system is used to implement the masked slice pyramid spatiotemporal feature fusion gait recognition method according to claims 1 to 9, and the system comprises: A data preprocessing module, wherein the data preprocessing module is used to perform data preprocessing operations on input video image data; A first-stage feature generation module, wherein the first-stage feature generation module is used to perform spatiotemporal feature extraction processing on the video image data after the data preprocessing operation to obtain first-stage features; A second-stage feature generation module, the second-stage feature generation module is used to perform dynamic mask slicing feature extraction processing on the first-stage features to obtain second-stage features; A three-stage feature generation module, the three-stage feature generation module is used to perform multi-scale time feature extraction processing on the two-stage features to obtain three-stage features; A four-stage feature generation module, the four-stage feature generation module is used to perform secondary dynamic mask slicing feature extraction processing on the three-stage features to obtain four-stage features; A feature fusion processing module, wherein the feature fusion processing module is used to perform pyramid adaptive feature fusion processing on the first-stage feature, the second-stage feature, the third-stage feature and the fourth-stage feature to obtain gait features after fusion processing; A training optimization processing module is used to perform training optimization processing on the gait features after the fusion processing.