Gait emotion recognition method and system based on multi-scale cross-space-time directed space-time graph

By constructing multi-scale cross-space-time directed space-time graphs and using adaptive graph convolution blocks, the problem of signal redundancy during feature extraction in gait emotion recognition is solved, and higher recognition accuracy is achieved.

CN116740802BActive Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310496658.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-05-06
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The existing gait emotion recognition methods fail to effectively consider the dependence between one frame and another during feature extraction, resulting in signal redundancy and recognition accuracy degradation.

Method used

By constructing a multi-scale cross-space-time directed space-time graph, a multi-scale adaptive graph convolution block is used to extract the feature connection between each frame node and its multi-hop node, and the node features are directionally updated in the time dimension.

Benefits of technology

The time-dimensional features in the gait data are effectively extracted, reducing feature redundancy and improving the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740802B_ABST
    Figure CN116740802B_ABST
Patent Text Reader

Abstract

The present invention discloses a gait emotion recognition method and system based on a multi-scale cross-time and space directed spatiotemporal graph, including three multi-scale spatiotemporal directed adaptive graph convolution networks, each of which is preceded by a normalization layer, followed by an activation function layer and a discard layer. After the human skeleton graph passes through the three multi-scale spatiotemporal directed adaptive graph convolution networks, feature data is extracted from the input data, and then a global average pooling layer is executed to perform global feature average pooling on the extracted features, and finally four emotion classifications are performed through a softmax layer. The present invention adopts the method of constructing a directed spatiotemporal graph, taking into account the direction of time flow, and using graph deep learning to extract the connection between nodes in the time dimension, and updating the node features in a direction to achieve better recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a gait emotion recognition method and system based on a multi-scale cross-spacetime directed spatiotemporal graph. Background Art

[0002] Emotion recognition has become one of the research hotspots in the field of human-computer interaction and affective computing. Among them, emotion recognition based on non-verbal information such as speech, images, and physiological signals has received widespread attention. However, these recognition methods all require people to consciously provide data sources when expressing emotions, so there are many limitations. In contrast, gait emotion recognition does not require people to consciously express emotions, but judges emotional states by analyzing walking postures, which is more natural and non-intrusive. Therefore, gait emotion recognition has become an important research direction in the field of emotion recognition in recent years. With the development of deep learning, non-verbal feature extraction has developed from a single facial emotion pattern to multiple patterns, namely multimodal emotion recognition. However, efficient multimodal emotion recognition requires high-quality and sufficient data, which brings high computing resource costs. As an alternative, emotion recognition from gait can classify human emotions based on easily accessible gait data. Unlike multi-modal emotion recognition, gait data can be simply obtained from videos, thus avoiding privacy issues.

[0003] Gait emotion recognition has broad application prospects, especially in the fields of smart medical care, smart monitoring, smart home, virtual reality, etc. For example, gait emotion recognition can be used to monitor the patient's mental state, monitor their emotional changes in real time, and provide necessary medical assistance in a timely manner; it can be used to monitor the emotional state of criminal suspects, prevent and solve cases; it can be used in smart home systems to provide more intelligent and personalized services; it can also be used in the field of virtual reality to make virtual reality systems more real and natural. It can be seen that gait emotion recognition has broad application prospects and has important social and economic value.

[0004] In traditional gait recognition methods, the features of each frame are extracted and embedded to represent a vector of a certain dimension, which can not only reduce the amount of calculation but also retain useful information. Gait recognition is performed using Dense Trajectories and its improved version iDT, machine learning classifiers such as SVM and random forest, and gait recognizers with neural networks. [Liu J, Shahroudy A, Perez M, et al. Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding [J]. IEEE transactions on pattern analysis and machine intelligence, 2019, 42 (10): 2684-2701.] The gait recognizer with a neural network achieves high-precision performance on benchmark datasets such as Human3.6M, Kinetics, and NTU RGB-D. [Reference. [3]. Yan S, Xiong Y, Lin D. Spatial temporal graph convolutional networks for skeleton-based action recognition [C]. Proceedings of the AAAI conference on artificial intelligence. 2018, 32 (1).] In ST-GCN, graph data is analogized to image data, and a graph convolutional neural network based on graph convolutional neural network is designed. Based on this, an adaptive graph convolutional neural network (AST-GCN) is proposed, which enables multiple learning subnetworks to perform adaptive learning on different graphs. In DGNN, graph neural networks are applied to gait recognition, where node features are updated and directed. [Reference [2]. Liu Z, Zhang H, Chen Z, et al. Disentangling and unifying graph convolutions for skeleton-based action recognition [C]. Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2020: 143-152.] In G3D, nodes establish undirected connections and are updated by considering the features of other nodes.Karg uses a PCA-based classifier, Crenn uses a support vector machine for emotional features, and Venture uses the autocorrelation matrix between joint angles for similarity-based classification. Daoudi represents joint motion as a symmetric positive definite matrix and performs nearest neighbor classification, LSTM uses a recurrent neural network to classify emotions, and STEP builds an autoencoder based on a graph convolutional neural network to improve the accuracy of classification tasks. HAP performs classification tasks by adding semi-supervised autoencoders to different parts of the human body.

[0005] There are key challenges in recognizing emotion from gait. Some methods only perform feature extraction through spatial or temporal modules without considering the dependency between one frame and another. In fact, when a pair of nodes are strongly correlated, one node should contain the features of the other node to reflect the connection between them. The connection between nodes in multiple frames is undirected, while gait is actually directional in the time direction. If the feature extraction is not updated in the direction of the time series, the nodes between different frames will interfere with each other and signal redundancy will appear, which will affect the update of node features in each frame. Summary of the invention

[0006] In each sample sequence, the spatial feature attributes of each frame node should not only depend on the current spatial position, but also be affected by the neighboring frame nodes before and after. For example, a person's left foot action depends on his leg lifting action at the previous moment and his leg lowering action at the next moment. The actions in the previous and next time periods are related to the current moment. Therefore, how to effectively extract this correlation is the key to improving the accuracy of emotion recognition.

[0007] The present invention directly stacks multiple time convolution blocks or connects nodes undirectedly to complete the extraction of time dimension features. The present invention constructs a directed space-time graph, takes into account the direction of time flow, and uses graph deep learning to extract the connection between nodes in the time dimension, and updates the node features in a direction to achieve better recognition accuracy.

[0008] The present invention solves the above technical problems and adopts a technical solution: a gait emotion recognition method based on a multi-scale cross-temporal directed space-time graph, comprising the following steps:

[0009] Construct a human skeleton graph, where each joint represents a node and each bone represents an edge.

[0010] A multi-scale cross-temporal and spatial information directed aggregation method is used to extract features. A cross-temporal and spatial method is used to obtain the connection between each frame node and frame nodes at a long distance. A multi-scale method is used to obtain the connection between each frame node and distant nodes.

[0011] According to the direction of time flow, all frame nodes are placed into an overall structure to construct a directed space-time graph.

[0012] Multi-scale adaptive graph convolution blocks are used to adaptively extract the feature connections between each frame node and its multi-hop nodes.

[0013] Softmax is used for sentiment classification.

[0014] The present invention also provides a gait emotion recognition system based on a multi-scale cross-time and space directed spatiotemporal graph, comprising three multi-scale spatiotemporal directed adaptive graph convolutional networks, each of which is preceded by a normalization layer, followed by an activation function layer and a discard layer, after which the human skeleton graph passes through the three multi-scale spatiotemporal directed adaptive graph convolutional networks, feature data is extracted from the joint data, and then a global average pooling layer is executed to perform global feature average pooling on the extracted features, and finally four emotion classifications are performed through a softmax layer;

[0015] The multi-scale spatiotemporal directed adaptive graph convolutional network aggregates multi-frame node information in a directed manner, updates each frame graph node, and extracts features of each frame graph node and its m-order neighbors, and then adaptively fuses them through a 1×1 adaptive graph convolution block, thereby fusing features of various scales.

[0016] The multi-scale spatiotemporal directed adaptive graph convolutional network aggregates the features of nodes of different scales in multiple frames, and constructs multiple directed spatiotemporal graphs according to different time spans and the number of frames that each frame node is connected to its neighboring frames.

[0017] A computer-readable storage medium stores executable instructions for implementing the above-mentioned gait emotion recognition method based on a multi-scale cross-spacetime directed spatiotemporal graph when executed by a processor.

[0018] The above technical solution has the following beneficial effects:

[0019] (1) A multi-scale cross-temporal and spatial information directed aggregation method is proposed to extract features in a directed manner. By using the cross-temporal and spatial method, the connection between each frame node and the frame nodes with a long distance can be obtained. At the same time, by using the multi-scale method, the connection between each frame node and the distant nodes can be obtained.

[0020] (2) According to the direction of time flow, a directed space-time graph is constructed. By constructing a directed space-time graph, all frame nodes are placed into an overall structure. The nodes are updated using the graph deep learning method, which can take into account the time flow and avoid the redundancy of feature information. This method can extract features in the time dimension.

[0021] (3) Use multi-scale adaptive graph convolution blocks to adaptively extract the connections between nodes in multiple frames. By using multi-scale graph convolution blocks, the feature connections between nodes in each frame and their multi-hop nodes can be adaptively extracted, and the adaptive graph convolution blocks can learn different parameters for different graph structure data.

[0022] (4) A multi-scale cross-spacetime directed adaptive spatiotemporal graph convolutional network (MSDAST-GCN) is proposed, in which the update of node feature data takes into account the influence of nodes in other frames with long and short time distances on nodes in the current frame. This network incorporates our proposed multi-scale cross-spacetime information directed aggregation method and directed spatiotemporal graph to extract the connection between long-distance time and multi-scale node features. This method avoids the feature redundancy problem in previous methods. At the same time, the node features after the temporal features are updated are learned using multi-scale adaptive graph convolution blocks, which ultimately improves the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a node connection diagram for any τ = 3 frames;

[0024] Figure 2 is a node connection diagram for any τ = 5 frames;

[0025] Figure 3 This is the overall model diagram;

[0026] Figure 4 Schematic diagram of the MSDAST-GCN module structure;

[0027] Figure 5 Learn feature maps with different dilation rates and scales for the MSDAST-GCN module. DETAILED DESCRIPTION

[0028] In the MSDAST-GCN recognition network of the present invention, the input data is a representation of the human skeleton. The specific data form is represented by a graph, each joint represents a node, and the skeleton represents an edge. The input data of the present invention begins to enter the directed aggregation module. In order to aggregate the time module features, each frame and its adjacent multiple neighbor frames are used to construct a directed space-time graph. When in order to obtain the influence of the multi-hop nodes of the neighbor frame on the current frame, the adjacency matrix representing the multi-hop neighbors can be used to construct a multi-scale directed space-time graph, which is multi-scale directed aggregation. In order to reduce the time complexity and obtain the connection between the nodes of the distant frames at the same time, a multi-scale cross-time and space directed aggregation is constructed in a cross-time manner. When k is set to 1 and d is set to 1, it is directed aggregation; when k is greater than 1 and d is set to 1, it is multi-scale directed aggregation; when k and d are greater than 1 at the same time, it is multi-scale cross-time and space directed aggregation. k represents the relationship between each node and its k-hop neighbor nodes in the adjacent frames when constructing the directed spatiotemporal graph, and d represents the number of spanned frames when each frame node is connected to the other frame nodes when constructing the directed spatiotemporal graph. When d=1, it means there is no span. After extracting the temporal features, the output features of the aggregated network with different settings of k and d are input into the multi-scale adaptive graph convolution blocks with different settings of m. m represents the use of adaptive graph convolution blocks to learn the feature connection between each frame graph node and its m-hop neighbor nodes. Then, the output features of the multi-scale adaptive graph convolution blocks with different learning scales are fused, and the fused feature data is feature classified to identify four different emotions.

[0029] Construct a human skeleton graph. The human skeleton graph is represented by G = (V, E), where V = {v1, v2, ..., v N} represents the joint represented by N nodes. E represents the set of edges of the bone, represented by the adjacency matrix A∈R N×N Capture. The adjacency matrix A is shown in formula (1.1). As a graph sequence, there exists a node feature set X = {x t,n ∈R C t,n∈1≤t≤T,1≤n≤N},X∈R T ×N×C , where T is the total number of time frames, N represents the number of nodes per frame, and R C Represents the characteristics of each node, R T×N×C Represents all feature data of each input sample, x t,n =X t,n is the node v in time frame t n The C-dimensional feature vector of the input data, t represents the frame of the input data, and n represents the nth node. The input gait can be described by A and X, where X t ∈R N×C are all node features in each frame at time t, X t,n is the nth node feature in each frame.

[0030]

[0031] d(v i ,v j ) represents the shortest distance between nodes i and j.

[0032] Directed Aggregation

[0033] In each update, the frame data is selected according to the time series direction, and the previous and next frames are used to update the current frame data. Other graph nodes are not connected to the graph node, and the node does not point to the node it is connected to. The update equation is as shown in Equation (1.2). Please note that nodes within a frame will not update each other.

[0034]

[0035] In the lth layer of the graph neural network, the output features are represented as The node input feature is represented as is the learning weight Θ of the lth layer in the neural network (l) Represents graph neural network. δ is the normalization function, which is used to normalize the updated nodes.

[0036] Multi-hop neighbors

[0037] Define a multi-scale adjacency matrix, where the shortest distance k between nodes i and j is given by equation (1.3). The size of k is the same as the distance size. By default, the distance between each node is 1, and the k-hop neighbor represents the distance between node i and node j is k.

[0038]

[0039] Among them, d(v i ,v j ) represents the shortest distance between nodes i and j, This is a diagram of the structure of human joints in their natural state. is the new adjacency matrix obtained after k jumps. Indicates whether nodes i and j are connected when the distance between nodes must be k. Reset the k-hop weight between two nodes to 1 so that n-hop nodes and k-hop nodes (k≠n) have the same initial weight when features are aggregated.

[0040] In order to efficiently extract node features in each frame, ST-GCN proposes a partitioning strategy based on 1-hop neighbors. In order to extract remote node features, a partitioning strategy based on m-hop neighbors is used. The partitioning strategy of the node hop neighbor set is defined as formula (1.4):

[0041]

[0042] Among them, r(v i ) represents the distance from node i to the centroid in all frames, and j is i’s m-hop neighbor. [A m ] i,j It means that the i-th node is connected to the j-th node.

[0043] Directed aggregation across time and space:

[0044] The present invention proposes a multi-scale cross-time directed aggregation update method to extract the connection relationship between nodes in different frames. Figure 1 As shown in Figure 1, each input sample contains T time series frames, and each frame contains V = 21 nodes. A directed spatiotemporal graph G is constructed over the entire time series. T =(V T ,E T ), the number of nodes is V T =T×V. The side is E T Indicates that its number is half the number of connected nodes, and the aggregated adjacency matrix Splice into As shown in formula (1.5). τ represents the number of frames connected to any frame node. In the constructed graph, it is not an undirected graph, but a bidirectional edge with different weights. This can not only reduce the complexity of the model, but also effectively extract the connection relationship between nodes and edges in the time dimension. In order to obtain the characteristic connection between distant frame nodes, the distant frame nodes can be connected in a directed manner. The cross-time expansion rate d is used to set the cross-time scale, such as Figure 1 shown.

[0045]

[0046] In each frame data, N represents the number of nodes. Represents the connection between node i and node j in the constructed directed space-time graph, where node i is in the i÷T frame and node j is in the j÷T frame.

[0047] The update function is the node feature update function h v In addition, two aggregation functions are used to extract features gv from source nodes and target nodes. t ,gv s , such as formula (1.6), formula (1.7), and formula (1.8).

[0048]

[0049]

[0050]

[0051] Indicates the extracted source node pointing to the i-node, Indicates the target node pointed to by the extracted i-node. Represents the source node pointing to the current node, Indicates the target node pointed to by the current node. v′ i Represents the updated features of node i.

[0052] Multi-scale directed aggregation across time and space:

[0053] We have previously shown how to construct a graph of the entire sequence, including directed connections between nodes in each frame and their multi-frame neighbors before and after them. However, we did not consider the connections between nodes in each frame and their multi-frame multi-hop neighbors before and after them. To address this issue, we can use multi-scale directed connections to connect nodes in each frame with their multi-frame multi-hop neighbors before and after them, which can capture more cross-temporal and spatial features and can be represented with fewer parameters, such as Figure 2 . Multi-scale directed space-time graph matrix It is composed of multiple original k-hop adjacency matrices spliced ​​together, as shown in formula (1.9).

[0054]

[0055] in is the multi-scale adjacency matrix of each input frame. I is an identity matrix.

[0056] MSDAST-GCN Network:

[0057] All updated node data passes through the adaptive graph convolution block to perform adaptive learning on each frame of the graph. The directed aggregation and adaptive graph convolution block as a whole is called the spatiotemporal directed adaptive graph convolution block. Its learning equation is as follows (1.10):

[0058]

[0059] In the formula, A n Represents an n-subgraph using a partitioning strategy, where the subset of each node is its one-hop neighbor node. n is a learnable parameter with an initial value of 1 and its shape and size are similar to A n Same. v is the number of subsets, usually set to 3. θn and W φn are the parameters of each embedding function θ and φ respectively. C n Represents adaptive learning parameters, used for adaptive learning of different features, is the input node feature, W nis the adaptive graph convolution network parameter. Formula (1.10), each graph adaptively learns and updates the feature strength according to the expansion rate d and the aggregation size k of each frame.

[0060] In order to consider the association of distant nodes. For example, clapping is related to happy emotions. A MSDAST-GCN block is proposed to extract and aggregate long-distance neighbor feature information according to the multi-scale segmentation strategy of each feature map as shown in Equation (1.11).

[0061]

[0062] Among them A n,m It is a subgraph of the construction graph when the multi-scale partitioning strategy is m. n,m is a learnable parameter B n,m The initial value is 0, and A n,m The same shape and size. is the input node feature, N m is the number of subsets, which is generally set to 3. The multi-scale adaptive learning method extracts and updates neighbors of each node at different distances. Due to the partitioning strategy, the feature extraction method can use different aggregation methods for different types of neighbors.

[0063] The n-hop neighbor features of the graph nodes in each frame are extracted and learned by formula (1.11). Since neighbors of different distances have different effects on each node, information of different scales is fused through a convolutional block of size 1×1.

[0064]

[0065] Where M is the maximum neighbor distance, set to 4. Θ l is the convolution block parameter of size 1×1 in layer l. σ(·) is the activation function. After adaptively learning features of different scales, these features can be adaptively fused according to formula (1.12).

[0066] From a global perspective, the model consists of 3 MSDAST-GCNs, each of which is preceded by a normalization layer, followed by an activation function layer and a dropout layer to prevent overfitting. After the data passes through the three MSDAST-GCNs, feature data is extracted from the joint data, and then a global average pooling layer is performed to perform global feature average pooling on the extracted features. Finally, four sentiment classifications are performed through a softmax layer. An MSDAST-GCN consists of multiple MDAST-GCNs with different void rates d and sliding windows τ. By setting different time spans, feature extraction can be performed on both short-distance and long-distance time scale feature data at the same time. The overall model is as follows Figure 3By setting different sliding window sizes, each frame node in the sliding window can be affected by more other frame nodes.

[0067] In the whole model, the normalization layer is represented by symbol ①, the MSDAST-GCN layer is represented by symbol ②, the activation layer is represented by symbol ③, the Dropout layer is represented by number ④, the average pooling layer is represented by number ⑤, and the Softmax layer is represented by number ⑥, which is finally classified into four emotions. The model has three MSDAST-GCN layers with 64, 128 and 256 channels respectively.

[0068] Multi-scale spatiotemporal directed adaptive graph convolutional network (MDAST-GCN), such as Figure 4 As shown in . It is constructed through multi-scale adaptive graph convolution blocks and multi-scale aggregate adjacency matrices. Adaptive fusion is performed through 1×1 convolution blocks to fuse features of various scales. Figure 4 As shown, this module aggregates multi-frame node information in a directed manner, updates each frame graph node, and extracts features of each frame graph node and its m-order neighbors.

[0069] The MSDAST-GCN module aggregates the features of nodes of different scales in multiple frames, and constructs multiple directed spatiotemporal graphs according to different time spans and the number of frames that each node is connected to its neighboring frames, and adaptively fuses the features through 1×1 convolution blocks. Figure 5 As shown, MSDAST-GCN can learn features with different dilation rates and scales.

[0070] The present invention uses the Emotion-Gait dataset for experiments, which collects various 3D posture sequences including 4 emotion labels. As shown in Table 1, the method of the present invention improves the recognition performance of all emotion categories compared with the previous method. Specifically, compared with the HAP method, the accuracy of the method of the present invention on the happy, sad, angry and normal categories is improved by 1.4%, 0.4%, 2.9% and 7.5%, respectively, while the average precision is improved by 3.55%.

[0071] Table 1 Comparison with other algorithms

[0072]

[0073] Each component of MSDAST-GCN is analyzed separately and the results are analyzed in a MAP manner. Note that the basic experiments do not use directed aggregation or multi-scale directed spatiotemporal adaptive graph convolution blocks.

[0074] The performance of the proposed multi-scale directed aggregation module under different aggregation scales and dilation rates is analyzed by setting the scale of the multi-scale spatiotemporal directed aggregation adaptive graph convolution block to a fixed value m = 1. Table 2 shows the average accuracy of the four emotions under different dilation rates and different aggregation scales. The accuracy of the baseline method is higher than that of the state-of-the-art HAP method. [7] The proposed algorithm is 3.7% lower than the baseline method and HAP by adding the directed aggregation module to the baseline method, with the expansion rate d = 1 and the aggregation scale fixed to 1. [7] The methods achieved MAP improvements of 4.8% and 1.1%, respectively. This shows that the node feature expression can be effectively enhanced by directionally aggregating the temporal neighboring joint features of nodes in different time frames of the node in the current frame. When d=2, MAP is improved by 0.7% compared with d=1. It can be seen that the cross-time aggregation method can effectively aggregate the features of distant time nodes. When k=2, MAP is improved by 0.5% compared with k=1. There is a correlation between the multi-hop neighbors of joint nodes in different time frames. When the aggregation scale is further expanded to k=4, MAP is improved by 0.3% compared with k=2, indicating that increasing the aggregation scale can improve the recognition accuracy. However, the influence of distant neighbors is smaller than that of nearby neighbors, but the complexity of the model is doubled, and a certain trade-off needs to be made.

[0075] Table 2. Benchmark experiments compared with aggregation methods with different aggregation scales k and expansion ratios d.

[0076]

[0077] The performance of MSDAST-GCN is demonstrated by changing the scale m of MSDAST-GCN with a constant dilation rate d=1 and aggregation scale k=1. Table 3 shows the MAP of MSDAST-GCN at scales m=1, 2, 3, and 4, as well as the comparative experiments when m=1. When m=2, the 1- and 2-hop neighbor features of each graph are learned, and the new features obtained by aggregation are processed by adaptive fusion. Compared with the baseline experiment, MAP is improved by 0.4%, indicating that there is a correlation between the multi-hop nodes of the spatiotemporal graph of each frame. When m=3 and m=4, MAP is improved by 1.3% and 1.4% respectively compared with the baseline method (first row). These two ablation experiments show that there is a correlation between the n-hop neighbors of the joint nodes in each frame. The MAP difference between the two is only 0.1%, indicating that the correlation is constantly weakening.

[0078] Table 3. Comparison of benchmark experiments with aggregation methods with different aggregation scales k and expansion ratios d.

[0079]

[0080]

[0081] [1].Yan S,Xiong Y,Lin D.Spatial temporal graph convolutional networksfor skeleton-based action recognition[C].Proceedings of the AAAI conferenceon artificial intelligence.2018,32(1).

[0082] [2].Shi L,Zhang Y,Cheng J,et al.Skeleton-based action recognitionwith directed graph neural networks[C].Proceedings of the IEEE / CVF conferenceon computer vision and pattern recognition.2019:7912-7921.

[0083] [3].Liu Z,Zhang H,Chen Z,et al.Disentangling and unifying graphconvolutions for skeleton-based action recognition[C].Proceedings of theIEEE / CVF conference on computer vision and pattern recognition.2020:143-152.

[0084] [4].Randhavane T,Bhattacharya U,Kapsaskis K,et al.Identifyingemotions from walking using affective and deep features[J].arXiv preprintarXiv:1906.11884,2019.

[0085] [5].Bhattacharya U,Mittal T,Chandra R,et al.Step:Spatial temporalgraph convolutional networks for emotion perception from gaits[C].Proceedingsof the AAAI Conference on Artificial Intelligence.2020,34(02):1342-1350.

[0086] [6].Chen Z,Li S,Yang B,et al.Multi-scale spatial temporal graphconvolutional network for skeleton-based action recognition[C] / / Proceedingsof the AAAI conference on artificial intelligence.2021,35(2):1113-1122.

[0087] [7].Bhattacharya U,Roncal C,Mittal T,et al.Take an emotion walk:Perceiving emotions from gaits using hierarchical attention pooling andaffective mapping[C].Computer Vision–ECCV 2020:16th European Conference,Glasgow,UK,August 23–28,2020,Proceedings,Part X.Cham:Springer InternationalPublishing,2020:145-163.

Claims

1. A gait emotion recognition method based on a multi-scale cross-temporal directed spatiotemporal graph, characterized in that: The following steps are involved: Construct a human skeleton graph, where each joint represents a node and each bone represents an edge; A multi-scale cross-temporal and spatial information directed aggregation method is used to extract features. A cross-temporal and spatial method is used to obtain the connection between each frame node and frame nodes at a long distance. A multi-scale method is used to obtain the connection between each frame node and distant nodes. According to the direction of time flow, all frame nodes are put into an overall structure to construct a directed space-time graph; Use multi-scale adaptive graph convolution blocks to adaptively extract the feature connections between each frame node and its multi-hop nodes; Use softmax for sentiment classification; The method of extracting features by using multi-scale cross-temporal and cross-spatial information directed aggregation specifically includes: in each update, frame data is selected according to the time series direction, and the previous and next frames are used to update the current frame data, and the update equation is as follows: In the lth layer of the graph neural network, the output features are represented as The node input feature is represented as Θ l represents graph neural network, δ is the normalization function; Define a multiscale adjacency matrix: Among them, d(v i ,v j ) represents the shortest distance between nodes i and j, This is a diagram of the structure of human joints in their natural state. is the new adjacency matrix obtained after k jumps, Indicates that the k-hop neighbor node of i is j, and resets the k-hop weight between the two nodes to 1, so that when features are aggregated, the n-hop node and the k-hop node have the same initial weight, k≠n; In order to extract remote node features, a partitioning strategy based on m-hop neighbors is used. The partitioning strategy of the node hop neighbor set is defined as: Among them, r(v i ) represents the distance from node i to the center of gravity in all frames, r(v j ) represents the distance from node j to the centroid in all frames, j is the m-hop neighbor of i, [A m ] i,j It means that the i-th node is connected to the j-th node; A multi-scale cross-time directed aggregation update method is used to extract the connection relationship between nodes in different frames, and two aggregation functions are used to extract features gv from source nodes and target nodes. t ,gv s , Indicates the extracted source node pointing to the i-node, Indicates the target node pointed to by the extracted i-node. Represents the source node pointing to the current node, Indicates the target node pointed to by the current node, v′ i represents the updated feature of node i, h v Represents the node feature update function; Use multi-scale directed connections to connect nodes in each frame with their previous and next multi-frame multi-hop neighbors, a multi-scale directed spatio-temporal graph matrix From multiple original k-hop adjacency matrices Stitched together, in is the multi-scale adjacency matrix of each input frame, and I is an identity matrix.

2. The gait emotion recognition method based on a multi-scale cross-spacetime directed spatiotemporal graph according to claim 1 is characterized in that: The human skeleton graph is represented by G = (V, E), where V = {v1, v2, ..., v N } represents the joint represented by N nodes, and E represents the set of edges of the bones, which is captured by the adjacency matrix A, which is shown as follows: d(v i ,v j ) represents the shortest distance between nodes i and j; As a graph sequence, there exists a node feature set X = {x t,n ∈R C |t,n∈1≤t≤T,1≤n≤N} means X∈R T×N×G , where T is the total number of time frames, N represents the number of nodes per frame, and R C Represents the characteristics of each node, R T×N×C Represents all feature data of each input sample, x t,n =X t,n is the node v in time frame t n The C-dimensional feature vector of , t represents the frame of the input data, and n represents the nth node; The input gait is described by A and X.

3. The gait emotion recognition method based on a multi-scale cross-spacetime directed spatiotemporal graph according to claim 1 is characterized in that: A directed space-time graph G is constructed over the entire time series T =(V T ,E T ), the number of nodes is V T =T×V, T means the input sample contains T time series frames, V means the number of nodes in each frame, E T Represents an edge.

4. The gait emotion recognition method based on a multi-scale cross-spacetime directed spatiotemporal graph according to claim 1, characterized in that: The multi-scale adaptive graph convolution block performs adaptive learning on each frame of the graph, and the learning equation is as follows: In the formula, A n represents the n subgraphs using the partitioning strategy, B n is a learnable parameter, C n represents the adaptive learning parameter, N v is the number of subsets, is the input node feature, W n are the adaptive graph convolutional network parameters; According to the multi-scale segmentation strategy of each feature map, the long-distance neighbor feature information is extracted and aggregated as follows: Among them A n,m is a subgraph of the construction graph when the multi-scale partitioning strategy is m, B n,m is a learnable parameter, is the input node feature, Nm is the number of subsets; The multi-scale adaptive learning method extracts and updates neighbors of each node at different distances.

5. The gait emotion recognition method based on a multi-scale cross-spacetime directed spatiotemporal graph according to claim 4 is characterized in that: Information of different scales is fused through convolution blocks of size 1×1; Where M is the maximum neighbor distance, Θ l are the parameters of the 1×1 convolutional block in layer l, and σ(·) is the activation function.

6. A gait emotion recognition system based on a multi-scale cross-temporal directed space-time graph, characterized by: It includes three multi-scale spatiotemporal directed adaptive graph convolutional networks. Each multi-scale spatiotemporal directed adaptive graph convolutional network is preceded by a normalization layer, followed by an activation function layer and a discard layer. After the human skeleton image passes through the three multi-scale spatiotemporal directed adaptive graph convolutional networks, feature data is extracted from the joint data, and then a global average pooling layer is executed to perform global feature average pooling on the extracted features. Finally, four emotion classifications are performed through the softmax layer; The multi-scale spatiotemporal directed adaptive graph convolutional network aggregates multi-frame node information in a directed manner, updates each frame graph node, and extracts features of each frame graph node and its m-order neighbors, and then adaptively fuses them through a 1×1 adaptive graph convolution block, thereby fusing features of various scales; The multi-scale spatiotemporal directed adaptive graph convolutional network uses a multi-scale cross-temporal and spatial information directed aggregation method to extract features, specifically including: in each update, frame data is selected according to the time series direction, and the previous and next frames are used to update the current frame data, and the update equation is as follows: In the lth layer of the graph neural network, the output features are represented as The node input feature is represented as Θ l represents graph neural network, δ is the normalization function; Define a multiscale adjacency matrix: Among them, d(v i ,v j ) represents the shortest distance between nodes i and j, This is a diagram of the structure of human joints in their natural state. is the new adjacency matrix obtained after k jumps, Indicates that the k-hop neighbor node of i is j, and resets the k-hop weight between the two nodes to 1, so that when features are aggregated, the n-hop node and the k-hop node have the same initial weight, k≠n; In order to extract remote node features, a partitioning strategy based on m-hop neighbors is used. The partitioning strategy of the node hop neighbor set is defined as: Among them, r(v i ) represents the distance from node i to the center of gravity in all frames, r(v j ) represents the distance from node j to the centroid in all frames, j is the m-hop neighbor of i, [A m ] i,j It means that the i-th node is connected to the j-th node; A multi-scale cross-time directed aggregation update method is used to extract the connection relationship between nodes in different frames, and two aggregation functions are used to extract features gv from source nodes and target nodes. t ,gv s , Indicates the extracted source node pointing to the i-node, Indicates the target node pointed to by the extracted i-node. Represents the source node pointing to the current node, Indicates the target node pointed to by the current node, v′ i represents the updated feature of node i, h v Represents the node feature update function; Use multi-scale directed connections to connect nodes in each frame with their previous and next multi-frame multi-hop neighbors, a multi-scale directed spatio-temporal graph matrix From multiple original k-hop adjacency matrices Stitched together, in is the multi-scale adjacency matrix of each input frame, and I is an identity matrix.

7. The gait emotion recognition system based on a multi-scale cross-spacetime directed spatiotemporal graph according to claim 6, characterized in that: The multi-scale spatiotemporal directed adaptive graph convolutional network aggregates the features of nodes of different scales in multiple frames, and constructs multiple directed spatiotemporal graphs according to different time spans and the number of frames connected between each frame node and neighbor frames; adaptive learning is performed on each frame of the graph, and the learning equation is as follows: In the formula, A n represents the n subgraphs using the partitioning strategy, B n is a learnable parameter, C n represents the adaptive learning parameter, N v is the number of subsets, is the input node feature, W n are the adaptive graph convolutional network parameters; According to the multi-scale segmentation strategy of each feature map, the long-distance neighbor feature information is extracted and aggregated as follows: Among them A n,m is a subgraph of the construction graph when the multi-scale partitioning strategy is m, B n,m is a learnable parameter, is the input node feature, Nm is the number of subsets; The multi-scale adaptive learning method extracts and updates the neighbors of each node at different distances; Information of different scales is fused through convolution blocks of size 1×1; Where M is the maximum neighbor distance, Θ l are the parameters of the 1×1 convolutional block in layer l, and σ(·) is the activation function.

8. A computer-readable storage medium, characterized in that: Executable instructions are stored, which are used to implement the gait emotion recognition method based on a multi-scale cross-spacetime directed space-time graph as described in any one of claims 1 to 5 when executed by a processor.