Gait recognition method based on three-branch input network and channel topological graph convolution
By using three-branch input network and channel topology map convolution in gait recognition technology, the topological features between bone points are captured, and the problem of limited recognition accuracy caused by ignoring the correlation of bone points in the existing technology is solved, and higher gait recognition accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510325757.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing gait recognition technology ignores the correlation between bone points, resulting in limited recognition accuracy.
The method based on three-branch input network and channel topology map convolution is adopted to capture the topological features between bone points through joint position, joint velocity and bone position as inputs, and aggregate with channel advanced features as output features for gait recognition.
By considering the correlation between bone points, the accuracy and stability of gait recognition are improved.
Smart Images

Figure CN120126219A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition based on artificial intelligence, and particularly relates to a gait recognition method based on a three-branch input network and channel topological graph convolution. Background Art
[0002] Gait recognition is a technology for recognizing human motion states and identity information by identifying human gait characteristics such as stride frequency and stride length. Its advantage lies in that it does not require contact sensors and allows for long-distance and non-invasive human posture identity recognition, and has many applications in the fields of medical treatment, security, smart home, etc.
[0003] In recent years, human bone modeling methods have shown excellent robustness and efficiency in gait recognition tasks. Some emerging technologies, such as graph convolutional networks and deep learning, have also made significant progress in this field, providing better solutions for the accuracy and robustness of gait recognition.
[0004] Since the human skeleton is a complex structure with unique characteristics, attention is usually focused on the bone points with obvious displacements, and basic gait characteristics can be obtained through these points. However, the existing technologies all ignore the correlation between bone points. In practice, different motion patterns will have different effects on the spatial correlation of nodes. If these correlations are ignored, it will lead to different degrees of feature loss problems, resulting in limited recognition accuracy. Summary of the Invention
[0005] In order to overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a gait recognition method based on a three-branch input network and channel topological graph convolution to improve the gait recognition accuracy by combining the correlation between bone points.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A gait recognition method based on a three-branch input network and channel topological graph convolution, comprising:
[0008] Step 1, obtaining a human bone sequence, constructing a bone graph G=(V, E, X') with joints as vertices and bones as edges, and obtaining the joint movement speed through the sequence time relationship, where V is the vertex set, E is the edge set, and X' is the feature set of vertices;
[0009] Step 2, using the joint position, joint speed, and bone position as the three-way inputs of the three-branch input network respectively, performing feature extraction on each of them, and then splicing the three output features to obtain a fused feature X;
[0010] Step 3: Using the fused feature X as the input of the channel topology module, extract channel features and sample features and convert them into high-level features, capture the topological features between joints related to different types of gaits, and aggregate the topological features with the high-level features of the corresponding channels as the output feature Z;
[0011] Step 4: After passing the output feature Z through the pooling layer, use the fully connected layer for gait recognition.
[0012] Further, in the skeleton graph, V = {v 1 , v 2 ,......, v N}, N is the number of joints, E is represented as the adjacency matrix I ∈ R N×N , and its element a ij reflects the correlation strength between the i-th joint v i and the j-th joint v j . The neighborhood of v i is represented as N(v i ) = {v j ∣ a ij ≠ 0}, and the feature representation of v i is x i ∈ R C , X’ ∈ R N×C , which is a matrix containing the C-dimensional features of all N joints.
[0013] Further, in the three-branch input network, each branch is composed of a BN layer and three convolutional layers M1, M2, and M3 connected in sequence; where M1 is a basic block, including a spatial convolution and a temporal convolution for extracting preliminary features, M2 adds a bottleneck block on the basis of M1 to reduce the computational complexity and retain important information, and M3 adds a residual structure and an SC-Attention block on the basis of M2 to focus on more important features.
[0014] Further, the channel topology module includes two cascaded units, each unit includes a channel topology graph convolutional network, the number of channels of the two units is different, and the second unit adds an SC-Attention block.
[0015] Further, the SC-Attention block performs channel and spatial weighting on the input features respectively to obtain channel features and spatial features, and adds the spatial features and the channel features as the output features.
[0016] Furthermore, each unit of the channel topology module further includes a temporal convolutional network. The input of the channel topology graph convolutional network is the fused feature X, and the output, together with the fused feature X, serves as the input of the temporal convolutional network. The input and output of the temporal convolutional network are concatenated to obtain the output feature Z.
[0017] Furthermore, the channel topology graph convolutional network consists of three functional parts: channel topology modeling function, feature transformation function, and channel feature aggregation function;
[0018] Its input is the fused feature X ∈ R N×C , and the output feature Z' ∈ R N×C is given by the following formula:
[0019] Z' = I(B(X), Y'(G(X), E'))
[0020] where I ∈ R N×N represents a fixed adjacency matrix based on the human bone structure, which is the matrix representation form of the edge set E. Both B(X) and G(X) represent feature transformation operations on the input, extracting different feature dimension information from the input; E' is a fixed feature matrix or parameter set, independent of the input fused feature X, and y' is a fusion function. Y'(G(X), E) represents combining the result of G(X) with E' to generate a new feature representation.
[0021] Furthermore, the implementation method of the channel topology modeling function is as follows:
[0022] A specific channel correlation K is obtained by performing a series of linear transformations on the input fused feature X, and then K is introduced into the adjacency matrix I, finally obtaining the refined channel topology Y ∈ R N×N×C ;
[0023] The implementation method of the feature transformation function is as follows:
[0024] When modeling the channel topology, a feature transformation function is used to transform the input fused feature X into a higher-level feature, and the expression is as follows:
[0025]
[0026] where is the transformed feature and W is the weight matrix;
[0027] The channel feature aggregation function:
[0028] According to the refined channel topology Y, the refined topology Y of each channel is obtained c , and using Y c and Construct a channel graph for each channel to achieve feature aggregation in a channel manner.
[0029] Further, the specific channel correlation K is obtained by the following method:
[0030] First, for a pair of vertices (v i , v j ) with corresponding features (x i , x j ), perform linear transformations Γ(x i ) and Φ(x j );
[0031] Second, calculate the distance d i along the channel dimension between x j and x ij ;
[0032] Then, perform a non-linear transformation on d ij , and use the obtained value w ij as the association weight of the specific channel between v i and v j to construct a dynamic topology matrix or represent the topological relationship of the vertex pair;
[0033] Finally, increase the channel dimension using the linear transformation ξ to obtain the specific channel correlation K, expressed as:
[0034] k ij = {ξ(G(Γ(x i ), Φ(x j ))}
[0035] where k ij is a vector belonging to K, reflecting the channel-specific topological relationship between the vertex pair (v i , v j ), k ij is not forced to be symmetric, G(Γ(x i ), Φ(x j ) is a modeling function, G(Γ(x i ), Φ(x j )) = σ(Γ(x i ) - Φ(x j ))), and σ(·) is an activation function.
[0036] Further, introducing K into the adjacency matrix I to obtain the refined channel topology Y, expressed as: Y = I + a·K; where a is an adjustable scaling factor used to control the influence degree of K on the finally obtained Y.
[0037] Compared with the prior art, the present invention takes joint positions, joint velocities, and bone positions as inputs, captures the topological features between joints under different types of gait correlations, and aggregates them with the high-level features of corresponding channels as output features for gait recognition. By considering the correlations between bone points, the accuracy and stability of gait recognition are improved. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0039] Figure 2 It is a schematic diagram of the topological structure of bone points with obvious displacements.
[0040] Figure 3 It is a structural diagram of the three-branch input network of the present invention.
[0041] Figure 4 It is the convolutional framework of the channel topological graph of the present invention. Detailed Embodiment
[0042] The following will describe in detail the implementation manner of the present invention in conjunction with the drawings and embodiments.
[0043] The present invention is a gait recognition method based on a three-branch input network and convolutional channel topological graph, which uses a three-branch input network and convolutional channel topological mapping to capture feature information in the spatio-temporal and channel dimensions. At the same time, in order to prevent the model from overly focusing on redundant features and reducing efficiency, the present invention also designs a spatial-channel focusing mechanism to help better capture key features. The overall framework of the present invention is as Figure 1 shown. The detailed description is as follows.
[0044] Step 1: Obtain the human bone sequence, construct a bone graph, and obtain the joint movement speed through the sequence time relationship.
[0045] The present invention uses human joint information as input features. Specifically, in one embodiment, HRNet is used to extract human joint points in the CASIA-B dataset, obtaining a human bone sequence composed of 17 joint points, and using the bone information for gait recognition. In the bone graph, joints are used as vertices and bones are used as edges to obtain the human bone representation. The bone graph is represented as G=(V, E, X'), as Figure 2 shown. Where V is the vertex set, which can be expressed as V={v 1 , v 2 ,......, v N}, and N is the number of joints. E is the edge set, which can be expressed as the adjacency matrix I∈R N×N , and its element a ij reflects the connection between the i-th joint v i and the j-th joint v jThe relevance strength, v i The neighborhood of is denoted as N(v i ) = {v j |a ij ≠ 0}, v i The feature of is denoted as x i ∈ R C . X’ is the feature set of vertices, X’ ∈ R N×C , which is a matrix containing the C-dimensional features of all N joints.
[0046] In a specific embodiment, first estimate the human pose, then perform topological optimization on the human skeleton, and use the optimized bone topology structure to combine the correlation weight and feature extraction mechanism to provide a more reliable feature representation for downstream pattern analysis and task applications. Figure 2 The different colors of the dashed lines in are used to distinguish different topological structures, and the thickness of the solid lines represents different degrees of correlation.
[0047] Step 2, construct an identification model.
[0048] The identification model of the present invention mainly consists of a three-branch input network and a channel topology module. Figure 2 The green part in shows the internal structure of the channel topology module, including CTCN, TCN, and the residual structure. The content executed by the three-branch input network and the channel topology module will be described separately below.
[0049] Step 21, take the joint position, joint velocity, and bone position as the three-way inputs of the three-branch input network respectively, perform feature extraction on each, and then splice the three output features to obtain the fused feature X.
[0050] It is worth mentioning that the current model-based methods should not only focus on one type of input feature. In model-based gait recognition, the present invention uses an improved three-branch input network as the first part of the model, as Figure 3 shown.
[0051] The input features of the model are divided into three categories: 1) joint position; 2) joint velocity; 3) bone position.
[0052] The three-branch input network of the present invention as a whole consists of three branches. Each branch is sequentially connected by a BN layer and three convolutional layers (M1, M2, M3).
[0053] To facilitate the processing of raw data, a basic block M1 is used in the first part of each branch. It contains a spatial convolution and a temporal convolution to extract preliminary features. Then, a bottleneck block and a residual structure are added to the basic block to obtain the M2 module, which is used to reduce the computational complexity and retain important information. Based on M2, a residual structure and an SC-Attention block are added to obtain the M3 module, which is used to focus on more important features. Finally, the features of these three branches are concatenated, and the resulting fused feature X is used as the input for the next part.
[0054] In the three-branch input network of the present invention, by capturing the bone positions, the model can capture the spatial layout information between joints, more accurately reflecting the walking characteristics of an individual. By considering the joint velocities, the dynamic changes during the gait can be captured, which helps to distinguish different gait phases and increase the sensitivity to subtle gait differences. In contrast, the comprehensive consideration of bone positions enables the model to integrate static and dynamic information, enhancing the overall understanding of gait patterns.
[0055] Step 22: Using the fused feature X as the input to the channel topology module, extract channel features and sample features and transform them into high-level features, capture the topological features between joints related to different types of gaits, and aggregate the topological features with the high-level features of the corresponding channels as the output feature Z.
[0056] The structure of the channel topology module is as Figure 2 and 4 , including two cascaded units C1 and C2. Each unit includes a channel topology graph convolutional network (CTCN module). The number of channels of the two units is different, and C2 adds an SC-Attention block based on C1. First, the extracted channel and sample features are converted into high-level features, and the static channel topology of the sample features is performed through an adjacency matrix to capture the topological features between bone points related to different types of gait features. Then, these topological features are aggregated with the high-level features of the corresponding channels as the output feature Z.
[0057] Furthermore, each unit (C1, C2) of the channel topology module of the present invention also includes a temporal convolutional network (TCN module). The input of the channel topology graph convolutional network (CTCN module) is the fused feature X, and the output, together with the fused feature X, is used as the input of the temporal convolutional network (TCN module). The input and output of the temporal convolutional network (TCN module) are concatenated to obtain the output feature Z.
[0058] The channel topology graph convolutional network (CTCN module) is a key part of the present invention. It mainly consists of three functional parts: channel topology modeling function, feature transformation function, and channel feature aggregation function; given its input as the fused feature X ∈ R M×C , the output feature Z’ ∈ R N×C is given by the following formula:
[0059] Z’ = I(B(X), Y’(G(X), E’))
[0060] where I ∈ R N×N is the aforementioned adjacency matrix, that is, the matrix representation of the edge set E, which numerically records the strength or weight of the node connection relationship. Therefore, I is a fixed adjacency matrix based on the human bone structure, defining the fixed connection relationship between the joint points in the bone graph. This adjacency matrix is "static" because it does not change with the samples or dynamic features, but is a fixed encoding of the predefined skeleton topology relationship. Static shared topology means using a fixed and unchanging adjacency matrix I to represent the relationship between nodes. This topological structure remains unchanged throughout the training and inference processes. It defines a fixed connection pattern for the channels or nodes in the network, that is, when calculating the channel topology graph convolution, it does not change dynamically with the input samples. This topological structure can be understood as a basic framework, depicting the mutual correlation between channels.
[0061] Both B(X) and G(X) represent operations for feature transformation on the input, which are a basic convolutional layer or feature extraction module, responsible for generating the preliminarily processed feature representation. Their role is to extract different feature dimension information from the input respectively, providing a basis for subsequent topological modeling and feature fusion. Among them, G(X) represents another feature transformation method different from B(X) for the input, and is complementary to B(X). It can be used to capture different feature dimension information.
[0062] E’ is a fixed feature matrix or parameter set, which is independent of the input fusion feature X and may contain the graph structure characteristics of the model or other inherent information. Y’ is a fusion function, and Y’(G(X), E’) represents the fusion of G(X) and E’, meaning that E’ is no longer a simple edge set at this time, but an additional, static feature or parameter, used to combine G(X) to generate the final result. By integrating the output features of G(X) and the external information contained in E’, Y’(G(X), E) generates a new feature representation. This representation is high-dimensional and contains the joint information of both, used for further topological modeling or feature classification.
[0063] The following details each function.
[0064] Channel topology modeling function:
[0065] Channel topology modeling is the most important part of the graph convolutional network. In this invention, a static adjacency matrix I is used to implement the topology of all channels. Specifically, a specific channel correlation K is obtained by performing a series of linear transformations on the input fused feature X, and then K is introduced into the static adjacency matrix I to finally obtain the refined channel topology Y ∈ R N×N×C . Among them, the channel topology Y is an actual tensor, representing the topological structure relationship between channels, and is a specific feature matrix generated through specific calculations.
[0066] Before calculating the specific channel correlation K, two linear transformation operations Γ and Φ are first performed to reduce the complexity of the model and save computational costs. In addition, since there is no fixed expression for the correlation between channels in existing research, a correlation modeling function G() is used to simulate the channel correlation between vertices. Assume there is a pair of vertices (v i , v j ), and their corresponding features are (x i , x j ), then the modeling function can be expressed as:
[0067] G(Γ(x i ), Φ(x j )) = σ(Γ(x i ) - Φ(x j ))
[0068] Where σ(·) is the activation function, Γ(x i ) represents the output after applying a linear transformation to the feature x i of vertex v i , and Φ(x j ) represents the output after applying a linear transformation to the feature x j of vertex v j ; Γ(x i ) and Φ(x j ) respectively perform affine transformations on the vertex features through different weight matrices and biases, expressed as:
[0069] Γ(x i ) = W Γ x i + b Γ
[0070] Φ(x j ) = W Φ x j + b Φ
[0071] Among them, W Φ and W Γ are both weight matrices, and b Φ and b Γ are both bias vectors.
[0072] Essentially, when a pair of vertex features are linearly transformed twice by Γ and φ, x can be calculated. i and x j The distance along the channel dimension. Then the distance is non-linearly transformed, and the results of these transformations can be used to characterize the channel-specific topological relationship between vertex pairs (v i , v j ).
[0073] During the non-linear transformation process, the distance is not directly used for characterization. Instead, the distance values are converted into weights or representations with a certain non-linear relationship through a mapping function. This non-linear transformation can capture complex patterns in the distance and map them to a space more suitable for describing the topological relationship between vertices.
[0074] The implementation method of non-linearly transforming the distance in the present invention is as follows:
[0075] First, calculate the Euclidean distance between vertex pairs:
[0076] d ij = ||x i - x j || 2
[0077] Then, perform a non-linear mapping on d ij through the activation function σ().
[0078] w ij = σ(d ij )
[0079] The purpose of the non-linear transformation here is to amplify or suppress certain distance patterns, map the distances in the linear space to the non-linear space, so that the weights of the approximate relationships are more in line with the actual topological relationships. The value w ij after non-linear transformation is used as the association weight of the specific channel between v i and v j to construct a dynamic topological matrix or represent the topological relationship of these vertex pairs. After obtaining the topological relationship based on the modeling function, the channel dimension is increased by using the linear transformation ξ to obtain the specific channel correlation K ∈ Y N×N×C' , which can be expressed as:
[0080] k ij = {ξ(G(Γ(x i ), Φ(x j ))} i,j ∈ {1,2,..., N}
[0081] where j ij ∈ R C′ is a vector belonging to K. k ijThe role is to reflect the channel-specific topological relationship between vertex pairs (v i , v j ). The vector k ij can be regarded as a representation characterizing the correlation between vertex pairs. By calculating k ij , the channel correlation between v i and v j can be captured. ξ is an operation for performing a linear transformation, and its main purpose is to map the result after a non-linear transformation to a higher channel dimension. The input received by ξ is the relational value calculated after being transformed by Γ and Φ, and then a linear mapping is performed on the input relational value. The formula form can be written as:
[0082] ξ(G(Γ(x i ), Φ(x j ))) = W · G(Γ(x i ), Φ(x j )) + b
[0083] where W is the weight matrix, which determines how to expand the low-dimensional relational value to the high-dimensional channel space. b is the bias vector, which is used to adjust the transformed value. Through the linear transformation ξ, the relational result based on Γ(x i ) and Φ(x j ) (obtained through the modeling function G) is further mapped to a higher channel dimension. Together, they complete the feature transformation process from the original vertex features to the channel-specific topological relationship, providing a richer information representation for subsequent graph convolution.
[0084] Note that the vectors k in K are not forced to be symmetric, so k ij ≠ k ji . This property increases the flexibility of correlation modeling and allows handling a wider range of topological relationship situations. By allowing asymmetry, the connection and association patterns between different channels can be depicted more accurately, thus better modeling the channel correlation between vertex pairs. Finally, the channel correlation K is used to refine the adjacency matrix I, and the shared channel topology Y is obtained as follows:
[0085] Y = γ(K, I) = I + a · K
[0086] where γ(K, I) represents a combined operation function that combines the static adjacency matrix I and the channel-specific correlation K. This function adjusts the original static adjacency matrix I so that it not only considers the static topological structure but also dynamically incorporates the channel correlation information, thus obtaining a new channel topology Y. a is an adjustable scaling factor, which is used to control the influence degree of the channel correlation matrix K on the final topology Y. The larger a is, the greater the contribution of the channel correlation to the final result; the smaller a is, the more the result depends on the static adjacency matrix I.
[0087] Feature transformation function:
[0088] When modeling the channel topology, a feature transformation function is used to transform the fused feature X into a more advanced feature. Its expression is:
[0089]
[0090] where is the transformed feature and W ∈ R N×C′ is the weight matrix. Other transformations, such as multi-layer perceptrons, can also be used here.
[0091] Channel feature aggregation function:
[0092] After obtaining the refined channel topology Y and the advanced feature , feature aggregation is achieved in a channel-wise manner. That is, based on the channel topology and features, the summary and representation enhancement of features are completed in each channel dimension. Specifically, the channel topology graph convolution constructs a channel graph for each channel. The channel topology graph convolution uses the refined topology Y c ∈ R N×N and the advanced feature to construct a channel graph for each channel. Here, Y c corresponds to the c-th channel, where c ∈ {1,..., C′}. Each channel graph reflects the vertex relationship under a specific type of motion feature. Therefore, feature aggregation is performed on each channel graph, and the final output Z is obtained by concatenating the output features of all channel graphs, which is formulated as:
[0093]
[0094] where || is the concatenation operation. During the whole process, the inference of the channel-specific correlation K depends on the input samples. Therefore, the proposed channel topology graph convolution is a dynamic graph convolution that adapts to different input samples. For each channel c, the channel-specific correlation part corresponding to this channel can be extracted from the channel correlation matrix and combined with the static topology to obtain the refined topology Y c of this channel, which means that the edge weights (representing the strength of the topological relationship) between vertex pairs have been determined. The graph structure of each channel can be simply understood as a "weighted adjacency matrix", that is, Y c . This matrix directly defines the graph structure under this channel, and the relationship and weights between vertices are characterized by Y c . Through Y c , the vertex set V and edge set E of each channel are defined. The vertices are fixed. The weights of the edges come from the element values of Y c : if Y cIf [i, j] > 0, then there exists a weighted edge from v i to v j with a weight of Y c [i, j].
[0095] It cannot be directly seen from the above calculation formula of Z that the inference of the channel-specific correlation K depends on the input samples. However, from the calculation process of the channel-specific correlation K, it can be known that different input samples will lead to different Ks, thus affecting the entire inference process. Therefore, the proposed channel topology graph convolution is a dynamic graph convolution that adapts to changes with different input samples.
[0096] In addition, to help learn the features of spatial positions and channel dimensions, the present invention designs an SC-Attention attention mechanism based on spatial and channel dimensions to improve the expression ability and generalization ability of the model. In the present invention, the main role of the SC-Attention block is to perform channel and spatial weighting on the input features respectively, obtain channel features and spatial features, and add the spatial features and channel features as output features. The specific implementation principle and method are as follows:
[0097] The SC-Attention block can multiply the channels containing important information by larger weights to obtain channel features. Specifically, channel information is extracted from the input features, and channel weights are adaptively generated through the attention mechanism. The channel weights represent the importance of each channel position in the feature map. Multiply the generated channel weights by the input features, that is, perform attention calculation on the input features. The obtained channel features have the same number of input channels, and each channel has only one pixel, and its pixel value represents the weight, that is, the importance of each channel. Due to the broadcast mechanism, all elements (one element represents one position) on the same input channel will be multiplied by the same channel weight, thus finally obtaining a set of channel feature maps with adjusted channel importance. They retain the information of the key channels and at the same time reduce the influence of redundant features.
[0098] Similarly, the SC-Attention block can multiply the space containing important information by a larger weight to obtain spatial features. Specifically, it extracts spatial location information from the input features and generates spatial weights through the attention mechanism, with a dimension of N×1×H×W. The spatial weights represent the importance of each spatial position in the feature map. Multiplying the generated spatial weights by the input features, i.e., performing attention calculation on the input features, the obtained spatial features have the same dimension. Since the dimension of the spatial weights matches the height and width of the input features, the broadcasting mechanism weights the feature values at each spatial position, thus highlighting the information at certain important positions and weakening the influence of unimportant regions. After weighting, a new set of feature maps is obtained, where the weights of each spatial position have been adjusted according to their importance. This set of weighted feature maps reflects the importance distribution of the input features in the spatial dimension, providing a more discriminative spatial information representation for subsequent feature fusion.
[0099] Subsequently, the spatial features are added to the channel features to obtain the final output features, and the expression is as follows:
[0100] X out = W channel ⊙X in + W pixel ⊙X in
[0101] where X in ∈R N×C×H×W represents the input features, that is, the feature map before channel weighting and spatial weighting, W chann ∈R N ×C×1×1 represents the channel weights, W pixel ∈R N×1×H×W represents the spatial weights, and ⊙ represents tensor broadcast multiplication.
[0102] It is worth mentioning that the SC-Attention block can adaptively adjust the importance of different channels, highlight key features, and suppress redundant information, thereby enhancing the model's attention to important features. On the other hand, the spatial attention mechanism helps the model assign different weights to different spatial positions in the feature map to capture important spatial structures and context information. This integrated attention mechanism enables the SC-Attention block to extract more discriminative and distinguishable features, thus improving the performance of the model in the gait recognition task. With fewer parameters, it effectively reduces the computational complexity and storage requirements of the model.
[0103] Step 4, after passing the output feature Z through the pooling layer, use the fully connected layer for gait recognition.
[0104] In summary, in the present invention, a three-branch input network is first used to process the gait sequence, and then non-shared topological modeling is performed. This modeling parameterizes the adjacency matrix I and uses it as the topology of the graph, while providing the general correlation between joints, inferring the correlation of channels from each sample, and representing the degree of correlation between the skeletal points within each channel. In addition, the present invention introduces the SC-Attention block to calculate the importance of each spatial position and channel to better capture key features, thereby improving the accuracy and robustness of the model.
Claims
1. A gait recognition method based on a three-branch input network and channel topology graph convolution, characterized in that: include: Step 1, obtain the human skeleton sequence, take the joints as vertices and the bones as edges, construct the skeleton graph G = (V, E, X'), and obtain the joint movement speed through the sequence time relationship, where V is the vertex set, E is the edge set, and X' is the feature set of the vertex; Step 2, using joint position, joint velocity and bone position as the three inputs of the three-branch input network, respectively, extracting features and concatenating the three output features to obtain the fusion feature X; Step 3, using the fusion feature X as the input of the channel topology module, extracting channel features and sample features and converting them into high-level features, capturing the topological features between joints related to different types of gaits, and aggregating the topological features with the high-level features of the corresponding channels as the output feature Z; Step 4: After the output feature Z passes through the pooling layer, the fully connected layer is used for gait recognition.
2. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 1 is characterized in that: In the skeleton graph, V = {v1, v2, ..., v N }, N is the number of joints, and E is represented by the adjacency matrix I∈R N×N , whose element a ij Reflects the i-th joint v i With the jth joint v j The strength of the correlation, v i The neighborhood of i )={v j ∣a ij ≠0}, v i The feature is represented by x i ∈R C , X'∈R N×C , is a matrix containing the C-dimensional features of all N joints.
3. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 1 is characterized in that: In the three-branch input network, each branch is composed of a BN layer and three convolutional layers M1, M2, and M3 connected in sequence; M1 is a basic block, including a spatial convolution and a temporal convolution, which is used to extract preliminary features. M2 adds a bottleneck block on the basis of M1 to reduce the computational complexity and retain important information. M3 adds a residual structure and an SC-Attention block on the basis of M2 to focus on more important features.
4. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 3 is characterized in that: The channel topology module includes two cascaded units, each unit includes a channel topology graph convolutional network, the number of channels of the two units is different, and the second unit adds an SC-Attention block.
5. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 3 or 4, characterized in that: The SC-Attention block performs channel and spatial weighting on the input features respectively to obtain channel features and spatial features, and adds the spatial features to the channel features as output features.
6. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 4 is characterized in that: Each unit of the channel topology module also includes a time domain convolutional network. The input of the channel topology graph convolutional network is the fusion feature X, and the output and the fusion feature X are used as the input of the time domain convolutional network. The input and output of the time domain convolutional network are concatenated to obtain the output feature Z.
7. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 5 is characterized in that: The channel topology graph convolutional network consists of three functional parts: channel topology modeling function, feature conversion function and channel feature aggregation function; Its input is the fusion feature X∈R N×C , output feature Z'∈R N×C Given by: Z'=I(B(X),Y'(G(X),E')) Where I∈R N×N It represents a fixed adjacency matrix based on the human skeleton structure, which is a matrix representation of the edge set E. B(X) and G(X) both represent feature transformation operations on the input, respectively extracting different feature dimension information from the input; E' is a fixed feature matrix or parameter set, independent of the input fusion feature X, Y' is a fusion function, and Y'(G(X), E) means combining the result of G(X) with E' to generate a new feature representation.
8. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 7 is characterized in that: The channel topology modeling function is implemented as follows: The specific channel correlation K is obtained by performing a series of linear transformations on the input fusion feature X, and then K is introduced into the adjacency matrix I to finally obtain the refined channel topology Y∈R N×N×C ; The feature conversion function is implemented as follows: When modeling the channel topology, the feature conversion function is used to convert the input fusion feature X into a higher-level feature. The expression is as follows: That is the transformed feature, W is the weight matrix; The channel feature aggregation function: According to the refined channel topology Y, obtain the refined topology Y of each channel c , using Y c and A channel graph is constructed for each channel to achieve feature aggregation in a channel manner.
9. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 8, characterized in that: The specific channel correlation K is obtained by the following method: First, for the corresponding feature (x i ,x j ) of a pair of vertices (v i ,v j ), performs a linear transformation Γ(x i ) and Φ(xj); Next, calculate x i and x j The distance d along the channel dimension ij ; Then, d ij Perform a nonlinear transformation and obtain the value w ij As v i With v j The association weights of specific channels between them are used to construct a dynamic topological matrix or represent the topological relationship between vertex pairs; Finally, the linear transformation ξ is used to increase the channel dimension and obtain the specific channel correlation K, which is expressed as: k ij ={ξ(G(Γ(x i ),Φ(x j ))} where k ij is a vector belonging to K, reflecting the vertex pair (v i ,v j ), k ij Symmetry is not enforced, G(Γ(x i ),Φ(x j ) is the modeling function, G(Γ(x i ),Φ(x j ))=σ(Γ(x i )-Φ(x j )), σ(·) is the activation function.
10. The gait recognition method based on three-branch input network and channel topology graph convolution according to claim 8 or 9, characterized in that: The introduction of K into the adjacency matrix I to obtain the refined channel topology Y is expressed as: Y=I+a·K; wherein a is an adjustable proportional factor for controlling the influence of K on the final Y.
Citation Information
Patent Citations
End-to-end human behavior recognition method and model based on skeleton nodes
CN114613013A
Human skeleton behavior recognition method and system based on multi-scale residual image convolutional network
CN114743273A
SAR image water body extraction method based on multi-scale residual attention model
CN115937707A