A cross-view gait recognition method based on interactive enhancement of skeleton spatiotemporal joint features
By designing a global spatiotemporal feature extraction network and a microspatiotemporal joint feature extraction network in cross-view gait recognition, and combining a hierarchical feature interaction enhancement fusion module, the problem of ignoring spatiotemporal correlation in the existing technology is solved, and a more efficient cross-view gait recognition accuracy is achieved.
Patent Information
- Application Number
- CN202410595928.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-05-14
AI Technical Summary
The existing cross-view gait recognition method ignores the spatial and temporal correlation between key points in the human body when extracting skeleton gait features, resulting in insufficient representation ability of the model.
A cross-view gait recognition method based on the interaction enhancement of the skeleton space-time joint feature interaction is designed. Through the global spatiotemporal feature extraction network and the microspatiotemporal joint feature extraction network, combined with the hierarchical feature interaction enhancement fusion module, multi-level features of the skeleton gait are extracted and fused.
The long-range global spatiotemporal characteristics and short-range spatiotemporal correlation information in the skeleton gait are effectively utilized, which improves the accuracy of gait recognition across perspectives and enhances the model's representation and generalization capabilities.
Smart Images

Figure CN118447576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-view gait recognition method based on skeleton spatiotemporal joint feature interactive enhancement, belonging to the technical field of deep learning and pattern recognition. Background Art
[0002] Gait recognition technology identifies an individual by analyzing the characteristics of their walking movements. Compared with mature biometric technologies such as facial and fingerprint recognition, it has many advantages. For example, it can use existing monitoring systems for long-distance recognition, does not require additional equipment or the cooperation of the person being identified, and is difficult to disguise. These characteristics make gait recognition have broad application potential in many fields, including security monitoring and daily attendance. However, gait recognition also faces challenges introduced by the high degree of freedom of the acquisition environment, such as interference from factors such as perspective, clothing, and carried items. Among them, changes in perspective can significantly affect recognition accuracy, because the difference in gait of the same person under different perspectives often exceeds the difference between different people under the same perspective. Therefore, the study of cross-perspective gait recognition is crucial for the practical application of gait recognition technology.
[0003] There are two main methods for cross-view gait recognition: model-based and appearance-based. Appearance-based methods usually use silhouette images of pedestrians as input and focus on the visual features of gait. Silhouette images are obtained by detecting, segmenting, cropping and binarizing the original video. These methods can be divided into template-based and contour set-based methods based on whether the contours are temporally compressed before feature extraction. Template-based methods simplify the model and calculation by integrating contour images into a single image, but this will lose temporal information. Contour set-based methods retain the temporal information of contour frames and process this information in feature space.
[0004] Model-based methods focus on analyzing the physical structure and movement of the human body and building a clear human body model. Skeleton sequence is a commonly used model that represents gait by estimating the coordinates of key points of the human body in the video. This compact data format is robust to viewpoints and other variables. Liao et al. used OpenPose to estimate the skeleton coordinates, combined with three artificial features based on prior knowledge, integrated the feature vectors of each frame to form a feature matrix, and then input it into CNN for recognition. Rao et al. proposed a self-supervised gait encoding method to learn gait representation by reconstructing skeleton sequences based on LSTM autoencoders. Different from the commonly used frame-level temporal feature extraction, Rashmi et al. subdivided a gait skeleton cycle into 8 events from the initial contact to the final swing, and extracted features from the gait events through LSTM. The above methods often use intermediate vectors to process graph structure data such as skeletons in non-Euclidean space. The development of graph convolutional networks (GCNs) makes it possible to extract gait features directly from skeleton sequences. Many skeleton-based cross-view gait recognition methods use spatiotemporal graph convolutional networks (ST-GCN) to extract gait features from skeletons. The spatiotemporal graph convolutional network consists of spatial graph convolutions for aggregating features between adjacent key points within a frame and one-dimensional temporal convolutions for modeling sequential temporal features. GaitGraph proposed by Teepe et al. adopts a ResGCN structure, adds residual connections to the ST-GCN block, and uses a bottleneck layer to reduce the feature size. They added multi-branch inputs containing speed and skeleton information to GaitGraph2, further improving the model performance. Wang et al. proposed a frame-level refinement network FTR-GC to adaptively learn specific topological structures in different frames and capture long-range dependencies between frames through self-attention.
[0005] However, there are still some problems to be solved in the above methods: First, the temporal and spatial features of the skeleton are extracted separately in different ways, which ignores the temporal and spatial correlation between the key points of the human body, that is, the connection between different key points in different frames. During walking, the correlation between different key points in short-range continuous frames is very important. For example, when lifting a leg, the movement of the knee joint in the previous frame will drive the movement of the ankle joint in the next frame. Ignoring this correlation will weaken the representation ability of the model. Secondly, in the model using a multi-branch network structure, different branches are independent of each other, and only the output features of the last layer of each branch are used for recognition. On the one hand, this will cause insufficient interaction between the features of different branches, resulting in the model being unable to utilize the fusion information of different features; on the other hand, the multi-level features of the network are not fully utilized, and the rich local information in the shallow network will be lost as the network deepens. Finally, when extracting local spatial features, body parts need to be manually divided according to the semantics of key points. Therefore, different division strategies need to be designed for different key point forms, which is not only time-consuming and labor-intensive, but also not conducive to model promotion.
[0006] Therefore, how to extract skeletal gait features that are view-invariant and efficiently utilize the long-range global spatiotemporal features and short-range spatiotemporal correlation information in the gait skeleton is the key to improving the accuracy of skeleton-based cross-view gait recognition. Summary of the invention
[0007] In view of the deficiencies in the prior art, the present invention provides a cross-view gait recognition method based on interactive enhancement of skeleton-based spatiotemporal joint features. Summary of the invention:
[0009] A cross-view gait recognition method based on skeleton spatiotemporal joint feature interactive enhancement, including skeleton data preprocessing, global spatiotemporal feature extraction network construction, micro-spatiotemporal joint feature extraction network construction, hierarchical feature interactive enhancement fusion module construction, overall framework training and cross-view gait recognition.
[0010] In order to obtain a rich gait skeleton representation, a skeleton data preprocessing module is designed. In order to model the global spatial relationship and long-range temporal relationship between key points, a global spatiotemporal feature extraction network based on multi-head self-attention graph convolution is designed. In order to model the local short-range spatiotemporal dependency between key points, a micro-spatiotemporal joint feature extraction network based on spatiotemporal joint graph convolution is constructed. In order to strengthen the interaction between the features extracted by the two networks, a hierarchical feature interaction enhancement fusion module is designed to enhance the features of different levels of the two networks and fully fuse them. In order to improve the discrimination ability of the entire framework structure, the triplet loss is used to train the entire model. Finally, the trained model is used for cross-view gait recognition.
[0011] The technical solution of the present invention is as follows:
[0012] A cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction includes the following steps:
[0013] A. Skeleton Data Preprocessing
[0014] Preprocess the gait skeleton sequence data, including skeleton data enhancement and skeleton descriptor calculation;
[0015] B. Global spatiotemporal feature extraction network construction
[0016] The global spatiotemporal feature extraction network is used to extract the global spatial relationship and long-range temporal relationship between key points. The global spatiotemporal features include global spatial relationship and long-range temporal relationship. For the skeleton sequence data obtained after preprocessing in step A, the self-attention adjacency matrix is calculated to extract the global spatial relationship, and then the long-range temporal relationship is modeled through large-kernel temporal convolution.
[0017] C. Micro-temporal and spatial joint feature extraction network construction
[0018] The micro-spatiotemporal joint feature extraction network is used to model the short-range spatiotemporal dependency between key points, namely the micro-spatiotemporal joint features. For the skeleton sequence data obtained after preprocessing in step A, several frames of skeleton graphs are aggregated into a spatiotemporal graph, and the tiny short-range dynamic features between key points are captured through the spatiotemporal joint graph convolution.
[0019] D. Construction of hierarchical feature interaction enhancement fusion module
[0020] The hierarchical feature interaction enhancement fusion module is used to fuse different features extracted by the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; the global spatiotemporal features and micro-spatiotemporal joint features of different levels obtained by step B and step C are spliced according to the channel dimension, and the global spatiotemporal features and micro-spatiotemporal joint features are fully fused through two sets of channel dimension adjustment networks and multi-dimensional attention enhancement fusion modules;
[0021] E. Overall framework training
[0022] The overall framework includes a global spatiotemporal feature extraction network, a micro-spatiotemporal joint feature extraction network, and a hierarchical feature interactive enhancement fusion module. The global spatiotemporal feature extraction network is stacked to extract global spatiotemporal features, which are then passed through the output layer as the final global spatiotemporal features; the micro-spatiotemporal joint feature extraction network is stacked to extract micro-spatiotemporal joint features, which are then passed through the output layer as the final micro-spatiotemporal joint features; the hierarchical feature interactive enhancement fusion module fully fuses the global spatiotemporal features and the micro-spatiotemporal joint features, which are then passed through the output layer as the final fused features;
[0023] Independently calculate the triplet loss of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module, and use the average of the losses of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module to supervise the training process of the overall framework;
[0024] F. Cross-view gait recognition
[0025] The gait skeleton sequences of the registration set and the query set are fed into the trained overall framework, and identity confirmation is achieved by comparing the similarity between the query sample features and the registration sample features.
[0026] Preferably, according to the present invention, in step A, the gait skeleton sequence data, i.e., the sequence composed of T-frame human skeleton images, is preprocessed, comprising the following steps:
[0027] a. Skeleton data enhancement: add Gaussian noise to the coordinates of each joint point and flip the skeleton left and right;
[0028] b. Calculation of skeleton descriptors: For the skeleton image sequence enhanced in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a skeleton gait representation.
[0029] Further preferably, for the enhanced skeleton image sequence in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a skeleton gait representation; including:
[0030] The human skeleton graph includes a set of nodes and edge set ε, expressed as Among them, the node set It includes N nodes representing key points of the human body; the edge set ε includes edges representing the connections between key points, using the adjacency matrix It means that if v i and v j There is a connection between i,j =1, otherwise A i,j = 0; the calculation of the skeleton descriptor takes a sequence of T-frame human skeleton images as input, and its node feature set The tensor is represented as Among them, x t,n is the key point v at time t in the human skeleton image of frame T n The C-dimensional feature vector is the coordinate of the key point in the original data; the skeleton sequence is represented by X and A;
[0031] Relative keypoint position set Its corresponding tensor is expressed as where r t,nCalculated by the following formula:
[0032] r t,n =x t,n -x t,* (1)
[0033] Among them, x t,* are the coordinates of the central key point;
[0034] The speed of movement is obtained by subtracting the coordinates of the same key point in different frames to calculate slow movement. and fast movement Two descriptors, using tensors as represents, where:
[0035]
[0036]
[0037] x t+1,n 、x t+3,n They are the key points v at time t+1 and t+2 in the human skeleton sequence. n The C-dimensional feature vector of
[0038] For bones, two descriptors, bone length and bone angle, are calculated; bone length is expressed by key point x t,n and the key point x connected to it t,adj The bone length set is expressed as the difference in coordinates. Its corresponding tensor is expressed as b t,n Calculated by the following formula:
[0039] b t,n =x t,n -x t,adj (4)
[0040] Bone Angle Set Its corresponding tensor is expressed as e t,n,c Calculated by the following formula:
[0041]
[0042] Among them, ‖·‖ represents the magnitude of the vector;
[0043] Finally, the tensors corresponding to the three types of descriptors are represented as X,X R ,X S ,X F ,X B ,X E Concatenate by channel dimension to get the skeleton gait representation F after enriching featuresin :
[0044] F in =Concat(X,X R ,X S ,X F ,X B ,X E ) (6)
[0045] in, C in =6C, Concat(·) means concatenation by channel dimension.
[0046] Preferably, according to the present invention, in step B, the global spatiotemporal feature extraction network is constructed, including:
[0047] The global spatiotemporal feature extraction network includes a global spatial feature extraction network and a long-range temporal feature extraction network;
[0048] c. Global spatial feature extraction network construction; for the input features of the first layer of the global spatial feature extraction network C t Represents the channel dimension of the input feature. First, After two independent linear layers, matrix multiplication, normalization and Softmax function, the self-attention adjacency matrix is obtained.
[0049]
[0050] in, is the learnable parameter matrix of the linear layer, C k Represents the channel dimension of the output tensor of the linear layer;
[0051] Calculate H-head self-attention
[0052] After obtaining the self-attention adjacency matrix, perform graph convolution operation:
[0053]
[0054] in, is the output of the global spatial feature extraction network, C l+1 represents the channel dimension of the output feature of the global spatial feature extraction network; σ(·) is the Mish activation function; is a learnable parameter matrix;
[0055] d. Construction of long-range temporal feature extraction network; Output features of the global spatial feature extraction network in step c Temporal features are extracted through temporal convolution with a convolution kernel size of 9×1:
[0056]
[0057] in, is the final output feature of the global spatiotemporal feature extraction network, and TCN(·) is the temporal convolution.
[0058] Preferably, according to the present invention, in step C, the micro-spatiotemporal joint feature extraction network is constructed, including:
[0059] e. Micro-spacetime subgraph construction; for the T-frame skeleton graph Human skeleton diagram sequence First, a time sliding window of size τ is set on the human skeleton sequence; the time sliding window slides on the human skeleton sequence with a step size of 1, and each sliding step generates a micro-spacetime subgraph. The micro-spacetime subgraph at time t in is a node set consisting of the nodes of the skeleton graph of the τ frame in the window, and its node features are expressed as tensors: Space-time edge Using the spatiotemporal adjacency matrix Indicates that A ST is given by τ 2 A block matrix composed of adjacency matrices:
[0060]
[0061] in, A is the original adjacency matrix, I is the identity matrix, i, j = 0, 1, ..., τ-1; the spatiotemporal adjacency matrix generalizes the connection relationship between key points in a frame to the time domain, and each key point is connected to the same key point in all frames in the time sliding window and its spatial neighbor key points;
[0062] Skeleton Diagram Its k-th hop adjacency matrix It is expressed as:
[0063]
[0064] in, is the adjacency matrix The element in row i and column j, d(v i ,v j ) represents the key point v in the skeleton graph i and v j The shortest path length between Correspondingly, the k-th hop time-space adjacency matrix is obtained
[0065]
[0066] f. Spatiotemporal joint graph convolution operation; define node features in step e and K-hop space-time adjacency matrix After that, a spatiotemporal joint graph convolution operation is performed in the sliding window at time t:
[0067]
[0068] in, Elements in yes Elements in are the input and output features of the spatiotemporal joint graph convolution operation in the l-th layer micro-spatiotemporal joint feature extraction network, is the learnable parameter matrix; σ(·) is the Mish activation function; Initialized to 0, dynamically strengthen or weaken any connection by adding it to the adjacency matrix; c is an element in δ;
[0069] g. Window dimension compression: concatenate the output features obtained by the spatiotemporal joint graph convolution operation in T groups of sliding windows to obtain the output spatiotemporal unified features. Feed it into a 3D convolution to compress the time window:
[0070]
[0071] Among them, reshape1(·) transforms the feature dimension from T×τN×C l Transform to C l ×T×τ×N, Conv3d(·) is a three-dimensional convolution with a kernel size of 1×τ×1 and an output feature dimension of C l+1 ×T×1×N; reshape2(·) changes the feature dimension from C l+1 ×T×1×N is transformed into T×N×C l+1 , and then get the final output feature of the l-th layer micro-temporal and spatial joint feature extraction network
[0072] Preferably, according to the present invention, in step D, the hierarchical feature interactive enhancement fusion module is constructed, comprising the following steps:
[0073] h. Multi-level feature channel dimension adjustment; for the intermediate layer features of the global spatiotemporal feature extraction network and the intermediate layer features of the micro-temporal and spatial joint feature extraction network Connect them according to the channel dimension respectively, input their respective channel dimension adjustment networks, and obtain the output feature F fs1 and F fs2 :
[0074]
[0075]
[0076] F fs1 =CDAN1(F G_inner ) (18)
[0077] F fs2 =CDAN3(F M_inner ) (19)
[0078] in, C inner is the sum of the number of feature channels in the middle layer; The output features of the network are adjusted for the channel dimension of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network, C fs is the number of feature channels after conversion, which is consistent with the number of output channels of the last layer of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; Concat(·) means concatenation by channel dimension; CDAN1(·) and CDAN2(·) are the channel dimension adjustment networks of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network respectively; for the middle-layer features of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network is used as its channel dimension adjustment network CDAN1(·); correspondingly, for the middle-layer features of the micro-spatiotemporal joint feature extraction network, the global spatiotemporal feature extraction network is used as its channel dimension adjustment network CDAN2(·);
[0079] i. Construction of multi-dimensional attention enhancement fusion module; for the output feature F of step h fs1 and F fs2 The multi-dimensional attention enhancement fusion module uses the time, space and channel dimension information of one input feature to enhance another input feature; the multi-dimensional attention enhancement fusion module includes three parallel attention branches, which respectively extract effective information from the time, space and channel dimensions and generate attention vectors; with F fs1 As the source feature, F fs2 As the target feature to be enhanced, for the source feature Time dimension attention weight α t Calculated by:
[0080] α t =reshape2(Sigmoid(Conv1d(AvgPool s (reshape1(F fs1 ))))) (20)
[0081] in, reshape1(·) transforms the feature dimension from T×N×C fs Convert to C fs ×T×N,AvgPool s (·) is the average pooling function in the spatial dimension, Conv1d(·) is the convolution kernel with size τ a The one-dimensional convolution is used to expand the temporal receptive field. Sigmoid(·) is the Sigmoid activation function. reshape2(·) converts the feature dimension from 1×T to T×1×1.
[0082] The attention weight α of the spatial dimension s Calculated by:
[0083] α s =reshape(Sigmoid(AvgPool t (F fs1 )W s )) (twenty one)
[0084] in, AvgPool t (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer, and reshape(·) transforms the feature dimension from N×1 to 1×N×1;
[0085] The attention weight α in the channel dimension c Calculated by:
[0086] α s =reshape(Sigmoid(σ((AvgPool st (F fs1 )W c1 ))W c2 )) (twenty two)
[0087] in, AvgPool st (·) is the average pooling function in time and space dimensions, are the learnable parameter matrices of the two fully connected layers, r is the dimension reduction factor used to reduce the number of parameters; σ(·) is the Mish activation function, and reshape(·) converts the feature dimension from 1×C fs Transformed to 1×1×C fs ;
[0088] The attention weights of time, space and channel dimensions are multiplied element by element, and the final multi-dimensional attention tensor W is obtained through the broadcast mechanism. md :
[0089] W md =Sigmoid(α t ⊙α s ⊙α c ) (twenty three)
[0090] in, ⊙ represents element-by-element multiplication, multi-dimensional attention tensor W md To enhance the target feature F fs2 :
[0091] F en1 =F fs2 ⊙W md +F fs2 (twenty four)
[0092] in, F fs1 Enhanced output features;
[0093] F fs2 As the source feature, F fs1 As the target feature to be enhanced, we get F fs2 Enhanced output features
[0094] Preferably, according to the present invention, step E, overall framework training, comprises the following steps:
[0095] j. Output features of the last layer of feature extraction network of global spatiotemporal feature extraction network and micro-spatiotemporal joint feature extraction network Interact with hierarchical features to enhance the output features of the fusion module Among them, C L =C fs , is the number of channels of the output feature; F en1 、F en2 Four independent and identically structured output layers are input respectively, and the output layer includes a temporal pooling layer, a spatial pooling layer, and a fully connected layer; The pooling process is expressed as:
[0096]
[0097]
[0098] in, Respectively represent the output features of the time pooling layer and the spatial pooling layer; MaxPool t (·) indicates the maximum pooling in the time dimension; MaxPool s (·) and AvgPool s(·) represent the maximum and average pooling of the spatial dimension respectively; Through a fully connected layer, the final output features of the global spatiotemporal feature extraction network are obtained
[0099]
[0100] in, C out is the number of channels of the output feature, FC(·) represents the fully connected layer;
[0101] The final output features of the micro-temporal joint feature extraction network and the hierarchical feature interactive enhancement fusion module are obtained and
[0102] The output features and Splicing as the final output feature F of the overall framework out :
[0103]
[0104] in,
[0105] Further preferably, the whole framework is trained using triplet loss, and the triplet loss of each output feature is calculated independently, and the triplet loss of the global spatiotemporal feature extraction network is calculated. The calculation method is as follows:
[0106]
[0107] Among them, N tri Represents the total number of triplets that can be formed in a batch, Extract the anchor sample gait features of the global spatiotemporal feature extraction network for the i-th triplet in the batch, is the gait feature of the positive sample with the same identity as the anchor sample, is the gait feature of the negative sample with different identity from the anchor sample, d(·) represents the Euclidean distance metric, and m represents the margin of triple loss;
[0108] The loss function of the micro-temporal and spatial joint feature extraction network and the hierarchical feature interaction enhancement fusion module is obtained and
[0109] Use the average of all losses as the final loss To supervise the model training process:
[0110]
[0111] Preferably, according to the present invention, the step F, cross-view gait recognition, comprises:
[0112] k. Send the registered set into the trained overall framework and convert the final output feature F in step j into out As the feature representation of each gait skeleton sequence, the feature database of the registration set is constructed using the feature representation of each gait skeleton sequence;
[0113] 1. Pre-process the query sample to be identified and send it into the trained overall framework to obtain the query sample features; calculate the Euclidean distance between the query sample features and all the features in the feature database of the registration set, and finally identify the identity of the query sample as the label of the feature in the registration set database with the smallest Euclidean distance to its features. By outputting the identity label of the query sample, the cross-view gait recognition process is completed.
[0114] The beneficial effects of the present invention are:
[0115] 1. The global spatiotemporal feature extraction network involved in the present invention combines the graph convolution based on the self-attention adjacency matrix with the large-kernel temporal convolution to dynamically model the global spatial relationship and long-range temporal relationship between key points.
[0116] 2. The micro-spatiotemporal joint feature extraction network involved in the present invention aggregates several frames of skeleton graphs into a spatiotemporal graph, and captures tiny short-range dynamic features between key points through spatiotemporal joint graph convolution.
[0117] 3. The hierarchical feature interactive enhancement fusion module involved in the present invention extracts the multi-level features of the two networks separately to form a separate fusion branch, so that the features of the two branches with the same dimension and different information are mutually enhanced, thereby summarizing complementary features with larger information content while avoiding negative impact on the original branch. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Figure 1 It is a schematic diagram of the structure of the global spatiotemporal feature extraction network in the present invention;
[0119] Figure 2 It is a schematic diagram of the structure of the micro-temporal and spatial joint feature extraction network in the present invention;
[0120] Figure 3 It is a structural schematic diagram of the hierarchical feature interactive enhancement fusion module in the present invention;
[0121] Figure 4 Schematic diagram of the structure of the multi-dimensional attention enhancement fusion module in the present invention;
[0122] Figure 5 This is the overall framework diagram of the cross-view gait recognition method based on interactive enhancement of skeleton spatiotemporal joint features proposed in the present invention. DETAILED DESCRIPTION
[0123] The present invention will be further described below by way of embodiments in conjunction with the accompanying drawings, but is not limited thereto.
[0124] Example 1
[0125] A cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction, such as Figure 5 As shown, the steps include:
[0126] A. Skeleton Data Preprocessing
[0127] Preprocess the gait skeleton sequence data, including skeleton data enhancement and skeleton descriptor calculation;
[0128] B. Global spatiotemporal feature extraction network construction
[0129] The global spatiotemporal feature extraction network is used to extract the global spatial relationship and long-range temporal relationship between key points. The global spatiotemporal features include global spatial relationship and long-range temporal relationship. For the skeleton sequence data obtained after preprocessing in step A, the self-attention adjacency matrix is calculated to extract the global spatial relationship, and then the long-range temporal relationship is modeled through large-kernel temporal convolution.
[0130] C. Micro-temporal and spatial joint feature extraction network construction
[0131] The micro-spatiotemporal joint feature extraction network is used to model the short-range spatiotemporal dependency between key points, namely the micro-spatiotemporal joint features. For the skeleton sequence data obtained after preprocessing in step A, several frames of skeleton graphs are aggregated into a spatiotemporal graph, and the tiny short-range dynamic features between key points are captured through the spatiotemporal joint graph convolution.
[0132] D. Construction of hierarchical feature interaction enhancement fusion module
[0133] The hierarchical feature interaction enhancement fusion module is used to fuse different features extracted by the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; the global spatiotemporal features and micro-spatiotemporal joint features of different levels obtained by step B and step C are spliced according to the channel dimension, and the global spatiotemporal features and micro-spatiotemporal joint features are fully fused through two sets of channel dimension adjustment networks and multi-dimensional attention enhancement fusion modules;
[0134] E. Overall framework training
[0135] The overall framework includes a global spatiotemporal feature extraction network, a micro-spatiotemporal joint feature extraction network, and a hierarchical feature interactive enhancement fusion module. The global spatiotemporal feature extraction network is stacked to extract global spatiotemporal features, which are then passed through the output layer as the final global spatiotemporal features; the micro-spatiotemporal joint feature extraction network is stacked to extract micro-spatiotemporal joint features, which are then passed through the output layer as the final micro-spatiotemporal joint features; the hierarchical feature interactive enhancement fusion module fully fuses the global spatiotemporal features and the micro-spatiotemporal joint features, which are then passed through the output layer as the final fused features;
[0136] The output layers of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module have the same structure, including pooling layers and fully connected layers. The output feature dimensions of the global spatiotemporal feature extraction branch and the micro-spatiotemporal feature extraction branch are consistent. The triplet loss of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module are calculated independently, and the average value of the loss of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module is used to supervise the training process of the overall framework; the output feature dimensions of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network are consistent. The names of each layer of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network are Layer-1, Layer-2, Layer-3, Layer-4, pooling layer, and fully connected layer, and their output feature dimensions are shown in Table 1.
[0137] Table 1
[0138] Layer Name Output feature dimension (T×N×C) Layer-1 60×17×64 Layer-2 60×17×64 Layer-3 60×17×128 Layer-4 60×17×256 Pooling Layer 1×256 Fully connected layer 1×128
[0139] F. Cross-view gait recognition
[0140] The gait skeleton sequences of the registration set and the query set are fed into the trained overall framework, and identity confirmation is achieved by comparing the similarity between the query sample features and the registration sample features.
[0141] Example 2
[0142] The difference between the cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction described in Example 1 is that:
[0143] In step A, the gait skeleton sequence data, i.e., a sequence of T-frame human skeleton images, is preprocessed, including the following steps:
[0144] a. Skeleton data enhancement: add Gaussian noise with a standard deviation of 0.3 to the coordinates of each joint point, and flip the skeleton left and right with a probability of 0.1 to improve the robustness and generalization ability of the model;
[0145] b. Calculation of skeleton descriptors: For the skeleton image sequence enhanced in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a more complete and accurate skeleton gait representation.
[0146] For the enhanced skeleton image sequence in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a more complete and accurate skeleton gait representation; including:
[0147] The human skeleton graph includes a set of nodes and edge set ε, expressed as Among them, the node set It includes N nodes representing key points of the human body; the edge set ε includes edges representing the connections between key points (i.e. bones), using the adjacency matrix It means that if v i and v j There is a connection between i,j =1, otherwise A i,j = 0; the calculation of the skeleton descriptor takes a sequence of T-frame human skeleton images as input, and its node feature set The tensor is represented as Among them, x t,n is the key point v at time t in the human skeleton image of frame T n The C-dimensional feature vector is the coordinate of the key point in the original data; the skeleton sequence is represented by X and A;
[0148] Regarding the key point positions, in addition to the key point coordinates of the original data, the present invention also calculates the relative key point positions to normalize the key point coordinates. Its corresponding tensor is expressed as where r t,n Calculated by the following formula:
[0149] r t,n =x t,n -x t,c (1)
[0150] Among them, x t,c are the coordinates of the central key point (nose);
[0151] The speed of movement is obtained by subtracting the coordinates of the same key point in different frames to calculate slow movement. and fast movement Two descriptors, using tensors as represents, where:
[0152]
[0153]
[0154] x t+1,n 、x t+3,n They are the key points v at time t+1 and t+2 in the human skeleton sequence. n The C-dimensional feature vector of
[0155] For bones, the present invention calculates two descriptors: bone length and bone angle; the bone length is expressed by the key point x t,n and the key point x connected to it t,adj The bone length set is expressed as the difference in coordinates. Its corresponding tensor is expressed as b t,n Calculated by the following formula:
[0156] b t,n =x t,n -x t,adj (4)
[0157] Bone Angle Set Its corresponding tensor is expressed as e t,n,c Calculated by the following formula:
[0158]
[0159] Among them, ‖·‖ represents the magnitude of the vector;
[0160] Finally, the tensors corresponding to the three types of descriptors are represented as X,X R ,X S ,X F ,X B ,X E Concatenate by channel dimension to get the skeleton gait representation F after enriching features in :
[0161] F in =Concat(X,X R ,X S ,X F ,X B ,X E ) (6)
[0162] in, C in =6C, Concat(·) means concatenation by channel dimension.
[0163] In step B, the global spatiotemporal feature extraction network is constructed, including:
[0164] like Figure 1As shown, the global spatiotemporal feature extraction network includes a global spatial feature extraction network and a long-range temporal feature extraction network;
[0165] c. Global spatial feature extraction network construction; for the input features of the first layer of the global spatial feature extraction network C S Represents the channel dimension of the input feature. First, After two independent linear layers, matrix multiplication, normalization and Softmax function, the self-attention adjacency matrix is obtained.
[0166]
[0167] in, is the learnable parameter matrix of the linear layer, C c Represents the channel dimension of the output tensor of the linear layer;
[0168] In order to enable the model to learn rich relationships from different subspaces, the H-head self-attention is calculated
[0169] After obtaining the self-attention adjacency matrix, perform graph convolution operation:
[0170]
[0171] in, is the output of the global spatial feature extraction network, C l+1 represents the channel dimension of the output feature of the global spatial feature extraction network; σ(·) is the Mish activation function; is a learnable parameter matrix;
[0172] d. Construction of long-range temporal feature extraction network; Output features of the global spatial feature extraction network in step c Temporal features are extracted through temporal convolution with a convolution kernel size of 9×1:
[0173]
[0174] in, is the final output feature of the global spatiotemporal feature extraction network, and TCN(·) is the temporal convolution.
[0175] In step C, the micro-temporal and spatial joint feature extraction network is constructed, such as Figure 2 As shown, including:
[0176] e. Micro-spacetime subgraph construction; for the T-frame skeleton graph Human skeleton diagram sequence First, a time sliding window of size τ is set on the human skeleton sequence; the time sliding window slides on the human skeleton sequence with a step size of 1, and each sliding step generates a micro-spacetime subgraph. The micro-spacetime subgraph at time t in is a node set consisting of the nodes of the skeleton graph of the τ frame in the window, and its node features are expressed as tensors: Space-time edge Using the spatiotemporal adjacency matrix Indicates that A ST is given by τ 2 A block matrix composed of adjacency matrices:
[0177]
[0178] in, A is the original adjacency matrix, I is the identity matrix, i, j = 0, 1, ..., τ-1; the spatiotemporal adjacency matrix generalizes the connection relationship between key points in a frame to the time domain, and each key point is connected to the same key point in all frames in the time sliding window and its spatial neighbor key points; for example, for any key point in the spatiotemporal adjacency matrix The value of the elements on the diagonal is 1, indicating that the same key points of the skeleton image of the i-th frame and the skeleton image of the j-th frame are connected. The other elements with a value of 1 indicate that the key points of the skeleton image of the i-th frame are connected to the key points at the corresponding spatial adjacent positions in the skeleton image of the j-th frame.
[0179] In order to model the multi-scale spatiotemporal relationship between key points, the present invention further introduces a K-hop spatiotemporal adjacency matrix. Its k-th hop adjacency matrix It is expressed as:
[0180]
[0181] in, is the adjacency matrix The element in row i and column j, d(v i ,v j ) represents the key point v in the skeleton graph i and v j The shortest path length between Correspondingly, the k-th hop time-space adjacency matrix is obtained
[0182]
[0183] f. Spatiotemporal joint graph convolution operation; define node features in step e and K-hop space-time adjacency matrix After that, a spatiotemporal joint graph convolution operation is performed in the sliding window at time t:
[0184]
[0185] in, Elements in yes Elements in are the input and output features of the spatiotemporal joint graph convolution operation in the l-th layer of micro-spatiotemporal joint feature extraction network, respectively. In order to reduce the number of parameters, the channel dimension remains unchanged here, and the channel dimension is changed in the window dimension compression of the subsequent step g; is the learnable parameter matrix; σ(·) is the Mish activation function;
[0186] In order to improve the flexibility of the model, the present invention designs two learnable parameters for the spatiotemporal joint graph convolution: and Add it to formula (13) to get the final spatiotemporal joint graph convolution formula:
[0187]
[0188] in, Elements in yes Elements in are the input and output features of the spatiotemporal joint graph convolution operation in the l-th layer of micro-spatiotemporal joint feature extraction network, respectively. In order to reduce the number of parameters, the channel dimension remains unchanged here, and the channel dimension is changed in the window dimension compression of the subsequent step g; is the learnable parameter matrix; σ(·) is the Mish activation function; Initialized to 0, dynamically strengthen or weaken any connection by adding it to the adjacency matrix; k is an element in δ; its role is to assign a weight to each hop adjacency matrix. Considering that the correlations between key points with different distances are different, the importance of adjacency matrices with different hops is also different. Adding δ can dynamically learn this importance.
[0189] g. Window dimension compression: concatenate the output features obtained by the spatiotemporal joint graph convolution operation in T groups of sliding windows to obtain the output spatiotemporal unified features. Feed it into a 3D convolution to compress the time window:
[0190]
[0191] Among them, reshape1(·) transforms the feature dimension from T×τN×C l Transform to C l×T×τ×N, Conv3d(·) is a three-dimensional convolution with a kernel size of 1×τ×1 and an output feature dimension of C l+1 ×T×1×N; reshape2(·) changes the feature dimension from C l+1 ×T×1×N is transformed into T×N×C l+1 , and then get the final output feature of the l-th layer micro-temporal and spatial joint feature extraction network
[0192] In step D, the hierarchical feature interaction enhancement fusion module is constructed, such as Figure 3 As shown, the steps include:
[0193] h. Multi-level feature channel dimension adjustment; for the intermediate layer features of the global spatiotemporal feature extraction network and the intermediate layer features of the micro-temporal and spatial joint feature extraction network Connect them according to the channel dimension respectively, input their respective channel dimension adjustment networks, and obtain the output feature F fs1 and F fs2 :
[0194]
[0195]
[0196] F fs1 =CDAN1(F G_inner ) (18)
[0197] F fs2 =CDAN2(F M_inner ) (19)
[0198] in, C inner is the sum of the number of feature channels in the middle layer; The output features of the network are adjusted for the channel dimension of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network, C fs is the number of feature channels after conversion, which is set to 256 here, consistent with the number of output channels of the last layer of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; Concat(·) means concatenation by channel dimension; CDAN1(·) and CDAN2(·) are the channel dimension adjustment networks of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network, respectively; for the middle-layer features of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network is used as its channel dimension adjustment network CDAN1(·); correspondingly, for the middle-layer features of the micro-spatiotemporal joint feature extraction network, the global spatiotemporal feature extraction network is used as its channel dimension adjustment network CDAN2(·);
[0199] i. Construction of multi-dimensional attention enhancement fusion module; for the output feature F of step h fs1 and F fs2 ,The multi-dimensional attention enhancement fusion module uses the temporal, spatial and channel dimensional information of one input feature to enhance another input feature; Figure 4 As shown in Figure 2, the multi-dimensional attention enhancement fusion module includes three parallel attention branches, which extract effective information from the time, space and channel dimensions and generate attention vectors; fs1 As the source feature, F fs3 As the target feature to be enhanced, for the source feature Time dimension attention weight α t Calculated by:
[0200] α t =reshape2(Sigmoid(Conv1d(AvgPool s (reshape1(F fs1 ))))) (20)
[0201] in, reshape1(·) transforms the feature dimension from T×N×C fs Convert to C fs ×T×N,AvgPool s (·) is the average pooling function in the spatial dimension, Conv1d(·) is the convolution kernel with size τ a The one-dimensional convolution is used to expand the temporal receptive field. Sigmoid(·) is the Sigmoid activation function. reshape2(·) converts the feature dimension from 1×T to T×1×1.
[0202] The attention weight α of the spatial dimension s Calculated by:
[0203] α s =reshape(Sigmoid(AvgPool t (F fs1 )W s )) (twenty one)
[0204] in, AvgPool t (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer, and reshape(·) transforms the feature dimension from N×1 to 1×N×1;
[0205] The attention weight α in the channel dimension c Calculated by:
[0206] α s =reshape(Sigmoid(σ((AvgPool st (F fs1 )W c1 ))W c2 )) (twenty two)
[0207] in, AvgPool st (·) is the average pooling function in time and space dimensions, are the learnable parameter matrices of the two fully connected layers, r is the dimension reduction factor used to reduce the number of parameters; σ(·) is the Mish activation function, and reshape(·) converts the feature dimension from 1×C fs Transformed to 1×1×C fs ;
[0208] The attention weights of time, space and channel dimensions are multiplied element by element, and the final multi-dimensional attention tensor W is obtained through the broadcast mechanism. md :
[0209] W md =Sigmoid(α t ⊙α s ⊙α c ) (twenty three)
[0210] in, ⊙ represents element-by-element multiplication, multi-dimensional attention tensor W md To enhance the target feature F fs2 :
[0211] F en1 =F fs2 ⊙W md +F fs2 (twenty four)
[0212] in, F fs1 Enhanced output features;
[0213] Through the above method, F fs2 As the source feature, F fs1 As the target feature to be enhanced, we get F fs2 Enhanced output features
[0214] Step E, overall framework training, includes the following steps:
[0215] j. Figure 5As shown in the figure, the output features of the last layer of feature extraction network of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network are Interact with hierarchical features to enhance the output features of the fusion module Among them, C L =C fs , is the number of channels of the output feature; F en1 、F en2 Four independent and identically structured output layers are input respectively, and the output layer includes a temporal pooling layer, a spatial pooling layer, and a fully connected layer; The pooling process is expressed as:
[0216]
[0217]
[0218] in, Respectively represent the output features of the time pooling layer and the spatial pooling layer; MaxPool t (·) indicates the maximum pooling in the time dimension; MaxPool s (·) and AvgPool s (·) represent the maximum and average pooling of the spatial dimension respectively; Through a fully connected layer, the final output features of the global spatiotemporal feature extraction network are obtained
[0219]
[0220] in, C out is the number of channels of the output feature, FC(·) represents the fully connected layer;
[0221] The final output features of the micro-spatiotemporal joint feature extraction network and the hierarchical feature interactive enhancement fusion module are obtained by the same method as above and
[0222] The output features and Splicing as the final output feature F of the overall framework out :
[0223]
[0224] in,
[0225] For the design of loss function, the present invention adopts triplet loss to train the whole framework. In order to fully retain the discriminative information of features, the triplet loss of each output feature is calculated independently. The triplet loss of the global spatiotemporal feature extraction network The calculation method is as follows:
[0226]
[0227] Among them, N tri Represents the total number of triplets that can be formed in a batch, Extract the anchor sample gait features of the global spatiotemporal feature extraction network for the i-th triplet in the batch, is the gait feature of the positive sample with the same identity as the anchor sample, is the gait feature of the negative sample with different identity from the anchor sample, d(·) represents the Euclidean distance metric, and m represents the margin of triple loss;
[0228] The loss function of the micro-spatiotemporal joint feature extraction network and the hierarchical feature interaction enhancement fusion module is obtained by the same method as above and
[0229] Use the average of all losses as the final loss To supervise the model training process:
[0230]
[0231] Step F, cross-view gait recognition, includes:
[0232] k. Send the registered gait skeleton sequence into the trained overall framework, and convert the final output feature F in step j into out As the feature representation of each gait skeleton sequence, the feature database of the registration set is constructed using the feature representation of each gait skeleton sequence;
[0233] 1. Pre-process the query sample to be identified (the perspective is different from the registration set) and send it into the trained overall framework to obtain the query sample features; calculate the Euclidean distance between the query sample features and all the features in the feature database of the registration set, and finally identify the identity of the query sample as the label of the feature in the registration set database with the smallest Euclidean distance to its features. By outputting the identity label of the query sample, the cross-perspective gait recognition process is completed.
[0234] In this embodiment, the number of hops K of the spatiotemporal adjacency matrix of the micro-spatiotemporal unified feature extraction network is set to 2, and the time window size τ is set to 3; the convolution kernel size τ of the one-dimensional convolution of the temporal attention branch in the multi-dimensional attention enhancement fusion module is aThe weight attenuation factor is set to 2e-5, and the channel attention branch dimension reduction factor r is set to 4. The Adam optimizer is used for training, and the weight decay is set to 2e-5. The OneCycleLR learning rate adjustment strategy is adopted, and the initial, maximum and final learning rates are set to 1e-5, 1e-3 and 1e-8 respectively, and the proportion of the learning rate increase is set to 0.475. In the training phase, 60 frames of skeleton images are used as input, and in the test phase, all frames are used. The model is trained for 40K iterations, and the batch size is set to (4, 32), that is, each batch contains 4 subjects, each subject contains 32 gait sequences, and the Rank-1 accuracy is selected to measure the accuracy of the model's gait recognition performance.
[0235] In order to verify the advancedness of the cross-view gait recognition method based on skeleton spatiotemporal joint feature interactive enhancement proposed in the present invention, the present invention is compared with 7 skeleton-based gait recognition methods on the CASIA-B dataset, including PoseGait, JointsGait, SDHF-GCN, GaitGraph, MSGG, CycleGait and GPGait. The dataset collects gait data of 124 subjects in a laboratory setting, and each subject contains 110 gait sequences taken from 11 different perspectives (0°, 18°, ..., 180°). For each perspective, each subject has 10 gait sequences of 3 conditions, namely 6 normal sequences (NM#1-6), 2 backpack sequences (BG#1-2) and 2 coat sequences (CL#1-2). The training and testing of the present invention on this dataset follows the most commonly used large sample training (LT) setting, that is, all sequences of the first 74 subjects are used as training sets, and the sequences of the remaining 50 subjects are used as test sets. In the test set, the first four normal walking gait sequences (NM#1-4) of each subject in each perspective are the registration set, and the remaining sequences NM#5-6, BG#1-2 and CL#1-2 are used as query sets under normal, backpack and clothing conditions, respectively. Consistent with the mainstream skeleton-based gait recognition method, this embodiment uses 17 joint skeleton data estimated by HRNet from RGB video for training and testing. Table 2 lists the cross-view gait recognition Rank-1 accuracy (%) of the present invention and other 7 advanced gait recognition methods under normal, backpack and coat conditions, respectively, where the accuracy refers to the average Rank-1 accuracy of each query perspective excluding other registered perspectives of the same perspective.
[0236] Table 2
[0237]
[0238]
[0239] It can be seen from Table 2 that the method of the present invention achieves the best cross-view recognition effect under all walking conditions, with the average recognition accuracy reaching 95.6%, 89.0% and 88.7% respectively.
Claims
1. A cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction, characterized in that: The steps include: A. Skeleton Data Preprocessing Preprocess the gait skeleton sequence data, including skeleton data enhancement and skeleton descriptor calculation; B. Global spatiotemporal feature extraction network construction The global spatiotemporal feature extraction network is used to extract the global spatial relationship and long-range temporal relationship between key points. The global spatiotemporal features include global spatial relationship and long-range temporal relationship. For the skeleton sequence data obtained after preprocessing in step A, the self-attention adjacency matrix is calculated to extract the global spatial relationship, and then the long-range temporal relationship is modeled through large-kernel temporal convolution. C. Micro-temporal and spatial joint feature extraction network construction The micro-spatiotemporal joint feature extraction network is used to model the short-range spatiotemporal dependency between key points, namely the micro-spatiotemporal joint features. For the skeleton sequence data obtained after preprocessing in step A, several frames of skeleton graphs are aggregated into a spatiotemporal graph, and the tiny short-range dynamic features between key points are captured through the spatiotemporal joint graph convolution. D. Construction of hierarchical feature interaction enhancement fusion module The hierarchical feature interaction enhancement fusion module is used to fuse different features extracted by the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; The global spatiotemporal features and micro-spatiotemporal joint features of different levels obtained in step B and step C are spliced according to the channel dimension, and the global spatiotemporal features and micro-spatiotemporal joint features are fully fused through two sets of channel dimension adjustment networks and multi-dimensional attention enhancement fusion modules; E. Overall framework training The overall framework includes a global spatiotemporal feature extraction network, a micro-spatiotemporal joint feature extraction network, and a hierarchical feature interaction enhancement fusion module. The global spatiotemporal feature extraction network is stacked to extract global spatiotemporal features, which are then passed through the output layer as the final global spatiotemporal features; the micro-spatiotemporal joint feature extraction network is stacked to extract micro-spatiotemporal joint features, which are then passed through the output layer as the final micro-spatiotemporal joint features. The hierarchical feature interaction enhancement fusion module fully fuses the global spatiotemporal features and micro-spatiotemporal joint features, and passes through the output layer as the final fusion feature; Independently calculate the triplet loss of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module, and use the average of the losses of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network, and the hierarchical feature interaction enhancement fusion module to supervise the training process of the overall framework; F. Cross-view gait recognition The gait skeleton sequences of the registration set and the query set are fed into the trained overall framework, and the identity is confirmed by comparing the similarity between the query sample features and the registration sample features; In step B, the global spatiotemporal feature extraction network is constructed, including: The global spatiotemporal feature extraction network includes a global spatial feature extraction network and a long-range temporal feature extraction network; c. Global spatial feature extraction network construction; for the input features of the first layer of the global spatial feature extraction network C l Represents the channel dimension of the input feature. First, After two independent linear layers, matrix multiplication, normalization and Softmax function, the self-attention adjacency matrix is obtained. in, is the learnable parameter matrix of the linear layer, C k Represents the channel dimension of the output tensor of the linear layer; Calculate H-head self-attention After obtaining the self-attention adjacency matrix, perform graph convolution operation: in, is the output of the global spatial feature extraction network, C l+1 represents the channel dimension of the output feature of the global spatial feature extraction network; σ(·) is the Mish activation function; is a learnable parameter matrix; d. Construction of long-range temporal feature extraction network; Output features of the global spatial feature extraction network in step c Temporal features are extracted through temporal convolution with a convolution kernel size of 9×1: in, is the final output feature of the global spatiotemporal feature extraction network, TCN(·) is the temporal convolution; In step C, the micro-temporal and spatial joint feature extraction network is constructed, including: e. Micro-spacetime subgraph construction; for the T-frame skeleton graph Human skeleton diagram sequence First, a time sliding window of size τ is set on the human skeleton sequence; the time sliding window slides on the human skeleton sequence with a step size of 1, and each sliding step generates a micro-spacetime subgraph. The micro-spacetime subgraph at time t in is a node set consisting of the nodes of the skeleton graph of the τ frame in the window, and its node features are expressed as tensors: Space-time edge Using the spatiotemporal adjacency matrix Indicates that A ST is given by τ 2 A block matrix composed of adjacency matrices: in, A is the original adjacency matrix, I is the identity matrix, i, j = 0, 1, ..., τ-1; the spatiotemporal adjacency matrix generalizes the connection relationship between key points in a frame to the time domain, and each key point is connected to the same key point in all frames in the time sliding window and its spatial neighbor key points; Skeleton Diagram Its k-th hop adjacency matrix It is expressed as: in, is the adjacency matrix The element in row i and column j, d(v i , v j ) represents the key point v in the skeleton graph i and v j The shortest path length between Correspondingly, the k-th hop time-space adjacency matrix is obtained f. Spatiotemporal joint graph convolution operation; define node features in step e and K-hop space-time adjacency matrix After that, a spatiotemporal joint graph convolution operation is performed in the sliding window at time t: in, Elements in yes Elements in are the input and output features of the spatiotemporal joint graph convolution operation in the l-th layer micro-spatiotemporal joint feature extraction network, is the learnable parameter matrix; σ(·) is the Mish activation function; Initialized to 0, dynamically strengthen or weaken any connection by adding it to the adjacency matrix; k is an element in δ; g. Window dimension compression: concatenate the output features obtained by the spatiotemporal joint graph convolution operation in T groups of sliding windows to obtain the output spatiotemporal unified features. Feed it into a 3D convolution to compress the time window: Among them, reshape1(·) transforms the feature dimension from T×τN×C l Transform to C l ×T×τ×N, Conv3d(·) is a three-dimensional convolution with a kernel size of 1×τ×1 and an output feature dimension of C l+1 ×T×1×N; reshape2(·) changes the feature dimension from C l+1 ×T×1×N is transformed into T×N×C l+1 , and then get the final output feature of the l-th layer micro-temporal and spatial joint feature extraction network 2. According to claim 1, a cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction is characterized in that: In step A, the gait skeleton sequence data, i.e., a sequence of T-frame human skeleton images, is preprocessed, including the following steps: a. Skeleton data enhancement: add Gaussian noise to the coordinates of each joint point and flip the skeleton left and right; b. Calculation of skeleton descriptors: For the skeleton image sequence enhanced in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a skeleton gait representation.
3. The cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction according to claim 1 is characterized in that: For the enhanced skeleton image sequence in step a, three types of skeleton descriptors related to joint positions, movement speeds and bones are calculated, and the three types of skeleton descriptors are combined to form a skeleton gait representation; including: The human skeleton graph includes a set of nodes and edge set ε, expressed as Among them, the node set It includes N nodes representing key points of the human body; the edge set ε includes edges representing the connections between key points, using the adjacency matrix It means that if v i and v j There is a connection between i,j =1, otherwise A i,j = 0; the calculation of the skeleton descriptor takes a sequence of T-frame human skeleton images as input, and its node feature set The tensor is represented as Among them, x t,n is the key point v at time t in the human skeleton image of frame T n The C-dimensional feature vector is the coordinate of the key point in the original data; the skeleton sequence is represented by X and A; Relative keypoint position set Its corresponding tensor is expressed as where r t,n Calculated by the following formula: r t,n =x t,n -x t,c (1) Among them, x t,c are the coordinates of the central key point; The speed of movement is obtained by subtracting the coordinates of the same key point in different frames to calculate slow movement. and fast movement Two descriptors, using tensors as represents, where: x t+1,n 、x t+2,n They are the key points v at time t+1 and t+2 in the human skeleton sequence. n The C-dimensional feature vector of ; For bones, two descriptors, bone length and bone angle, are calculated; bone length is expressed by key point x t,n and the key point x connected to it t,adj The bone length set is expressed as the difference in coordinates. Its corresponding tensor is expressed as b t,n Calculated by the following formula: b t,n =x t,n -x t,adj (4) Bone Angle Set Its corresponding tensor is expressed as e t,n,c Calculated by the following formula: Among them, ||·|| represents the modulus of the vector; Finally, the tensors corresponding to the three types of descriptors are represented as X, X R , X s , X F , X B , X E Concatenate by channel dimension to get the skeleton gait representation F after enriching features in : F in =Concat(X,X R ,X S ,X F ,X B ,X E ) (6) in, C in =6C, Concat(·) means concatenation by channel dimension.
4. The cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction according to claim 1 is characterized in that: In step D, the hierarchical feature interaction enhancement fusion module is constructed, including the following steps: h. Multi-level feature channel dimension adjustment; For the intermediate layer features of the global spatiotemporal feature extraction network and the intermediate layer features of the micro-temporal and spatial joint feature extraction network Connect them according to the channel dimension respectively, input their respective channel dimension adjustment networks, and obtain the output feature F fs1 and F fs2 : F fs1 =CDAN1(F G_inner ) (18) F fs2 =CDAN2(F M_inner ) (19) Among them, F G_inner , C inner is the sum of the number of feature channels in the middle layer; The output features of the network are adjusted for the channel dimension of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network, C fs is the number of feature channels after conversion, which is consistent with the number of output channels of the last layer of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network; Concat(·) means concatenation by channel dimension; CDAN1(·) and CDAN2(·) are the channel dimension adjustment networks of the global spatiotemporal feature extraction network and the micro-spatiotemporal joint feature extraction network respectively; for the middle-layer features of the global spatiotemporal feature extraction network, the micro-spatiotemporal joint feature extraction network is used as its channel dimension adjustment network CDAN1(·); correspondingly, for the middle-layer features of the micro-spatiotemporal joint feature extraction network, the global spatiotemporal feature extraction network is used as its channel dimension adjustment network CDAN2(·); i. Construction of multi-dimensional attention enhancement fusion module; for the output feature F of step h fs1 and F fs2 The multi-dimensional attention enhancement fusion module uses the time, space and channel dimension information of one input feature to enhance another input feature; the multi-dimensional attention enhancement fusion module includes three parallel attention branches, which respectively extract effective information from the time, space and channel dimensions and generate attention vectors; with F fs1 As the source feature, F fs2 As the target feature to be enhanced, for the source feature Time dimension attention weight α t Calculated by: α t =reshape2(Sigmoid(Conv1d(AvgPool s (reshape1(F fs1 ))))) (20) in, reshape1(·) transforms the feature dimension from T×N×C fs Convert to C fs ×T×N,AvgPool s (·) is the average pooling function in the spatial dimension, Conv1d(·) is the convolution kernel with size τ a The one-dimensional convolution is used to expand the temporal receptive field. Sigmoid(·) is the Sigmoid activation function. reshape2(·) converts the feature dimension from 1×T to T×1×1. The attention weight α of the spatial dimension s Calculated by: α s =reshape(Sigmoid(AvgPool t (F fs1 )W s )) (21) in, AvgPool t (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer, and reshape(·) transforms the feature dimension from N×1 to 1×N×1; The attention weight α in the channel dimension c Calculated by: α s =reshape(Sigmoid(σ((AvgPool st (F fs1 )W c1 ))W c2 )) (22) in, AvgPool st (·) is the average pooling function in time and space dimensions, are the learnable parameter matrices of the two fully connected layers, r is the dimension reduction factor used to reduce the number of parameters; σ(·) is the Mish activation function, and reshape(·) converts the feature dimension from 1×C fs Transformed to 1×1×C fs ; The attention weights of time, space and channel dimensions are multiplied element by element, and the final multi-dimensional attention tensor W is obtained through the broadcast mechanism. md : W md =Sigmoid(α t ⊙α s ⊙α c ) (23) in, ⊙ represents element-by-element multiplication, multi-dimensional attention tensor W md To enhance the target feature F fs2 : F en1 =F fs2 ⊙W md +F fs2 (24) in, F fs1 Enhanced output features; F fs2 As the source feature, F fs1 As the target feature to be enhanced, we get F fs2 Enhanced output features 5. The cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction according to claim 1, characterized in that: The step E, overall framework training, includes the following steps: j. Output features of the last layer of feature extraction network of global spatiotemporal feature extraction network and micro-spatiotemporal joint feature extraction network Interact with hierarchical features to enhance the output features of the fusion module Among them, C L =C fs , is the number of channels of the output feature; F en1 、F en2 Four independent and identically structured output layers are input respectively, and the output layer includes a temporal pooling layer, a spatial pooling layer, and a fully connected layer; The pooling process is expressed as: in, Respectively represent the output features of the time pooling layer and the spatial pooling layer; MaxPool t (·) indicates the maximum pooling in the time dimension; MaxPool s (·) and AvgPool s (·) represent the maximum and average pooling of the spatial dimension respectively; Through a fully connected layer, the final output features of the global spatiotemporal feature extraction network are obtained in, C out is the number of channels of the output feature, FC(·) represents the fully connected layer; The final output features of the micro-temporal joint feature extraction network and the hierarchical feature interactive enhancement fusion module are obtained and The output features and Splicing as the final output feature F of the overall framework out : in, 6. The cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction according to claim 1 is characterized in that: The whole framework is trained using triplet loss, and the triplet loss of each output feature is calculated independently, and the triplet loss of the global spatiotemporal feature extraction network is calculated. The calculation method is as follows: Among them, N tri Represents the total number of triplets that can be formed in a batch, Extract the anchor sample gait features of the global spatiotemporal feature extraction network for the i-th triplet in the batch, is the gait feature of the positive sample with the same identity as the anchor sample, is the gait feature of the negative sample with different identity from the anchor sample, d(·) represents the Euclidean distance metric, and m represents the margin of triple loss; The loss function of the micro-temporal and spatial joint feature extraction network and the hierarchical feature interaction enhancement fusion module is obtained and Use the average of all losses as the final loss To supervise the model training process:
7. A cross-view gait recognition method based on multi-scale skeleton spatiotemporal feature extraction according to any one of claims 1-6, characterized in that: The step F, cross-view gait recognition, comprises: k. Send the registered set into the trained overall framework and convert the final output feature F in step j into out As the feature representation of each gait skeleton sequence, the feature database of the registration set is constructed using the feature representation of each gait skeleton sequence; 1. Pre-process the query sample to be identified and send it into the trained overall framework to obtain the query sample features; calculate the Euclidean distance between the query sample features and all the features in the feature database of the registration set, and finally identify the identity of the query sample as the label of the feature in the registration set database with the smallest Euclidean distance to its features. By outputting the identity label of the query sample, the cross-view gait recognition process is completed.
Citation Information
Patent Citations
Skeleton behavior recognition method based on spatial-temporal feature enhanced graph convolutional network
CN114882421A
Person re-identification method and apparatus for fusing global features with ladder-shaped local features
WO2024021394A1