Human body behavior recognition method based on fine-grained data space-time diagram neural network
By constructing a behavior graph and adding virtual nodes, combined with the graph Transformer model, the shortcomings of existing models in space-time feature extraction and time-dependent feature extraction are solved, and more efficient human behavior recognition effect is achieved.
Patent Information
- Application Number
- CN202410425402.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing human behavior recognition models have shortcomings in spatial and temporal feature extraction and temporal-dependent feature extraction, especially the lack of non-Euclidean spatial information and long-range-dependent feature extraction capabilities, resulting in limited model performance.
The spatiotemporal graph neural network method based on fine-grained data is adopted to construct the behavioral graph and its adjacency matrix, add virtual nodes, combine the graph Transformer model to fusion of spatial and temporal features, and use GAT and Transformer encoder for end-to-end deep learning to achieve efficient extraction and classification of spatiotemporal features.
It improves the accuracy and robustness of human behavior recognition, solves the limitations of existing models in non-Euclidean spatial information and time-dependent feature extraction, and achieves better classification performance.
Smart Images

Figure CN120277452A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning human behavior recognition, and in particular to a human behavior recognition method based on a spatio-temporal graph neural network of fine-grained data. Background Technique
[0002] With the rapid development of microelectronics technology and pervasive computing, multi-sensor systems based on smartphones, bracelets, and smartwatches have become ubiquitous in daily life. This trend has promoted the wide application of pervasive computing applications relying on sensor data in fields such as smart homes, medical assistance, and sports training. Human activity recognition (HAR) takes the human behavior data detected by sensors as input, and outputs specific behavior categories after being processed by a model.
[0003] Human activity recognition (HAR) benefits from research in other fields, such as natural language processing and computer vision. With the progress of artificial intelligence in various fields, HAR models and methods are constantly evolving, and their performance is also continuously improving. Human behavior recognition models can be roughly divided into two categories: traditional machine learning methods and deep learning methods. Traditional machine learning methods use domain knowledge to manually extract features from human behavior data, and then use machine learning algorithms to classify these features. The advantage of traditional machine learning methods lies in the better interpretability of their manually designed features and models. However, it is also limited by the domain knowledge of researchers, thus limiting the performance of the model. On the other hand, deep learning methods provide automatic feature extraction capabilities, making the model have better robustness and generalization. In addition, the end-to-end training mode makes model training more convenient. Therefore, deep learning methods have received extensive attention and become the mainstream of research in this field.
[0004] In the human behavior recognition task, multi-channel sensor data often cascades along the channel axis. This data structure is similar to the representation form of single-channel images. In addition, multi-channel perceptual data can also be interpreted as sequential feature inputs along the time dimension, which has similarity with the natural language processing method (text data is also processed sequentially), paving the way for applying many classic model ideas of computer vision and natural language processing to the field of human behavior recognition.
[0005] As a classic model in image classification, CNN was first applied to feature extraction and activity classification in HAR tasks. Basic models for sequence data classification, such as LSTM and its variants, have also been applied to HAR tasks, that is, multiple time steps of multi-sensor perception data are input sequentially for classification. Recently, the Transformer model, which has achieved excellent performance in multiple fields such as computer vision (CV) and natural language processing (NLP), has also been applied to the HAR field. However, the feature extraction ability of a single model for a single category is limited. Therefore, using multiple basic models serially or in parallel for feature extraction can obtain more ideal results. et al. connected CNN and LSTM sequentially, making full use of the spatial feature extraction ability of CNN and the temporal feature extraction ability of LSTM. Han et al. further improved the extraction of spatial features, using a graph convolutional network (GCN) for spatial feature extraction in a non-Euclidean space, and using temporal convolution and LSTM for temporal feature extraction. Qian et al. proposed fusing statistical features for human behavior recognition on the basis of parallelly extracting spatio-temporal features using CNN and LSTM. The hybrid model extracts and fuses features from the spatio-temporal perspective and achieves better classification performance.
[0006] A fundamental assumption in spatio-temporal modeling is that the future information of a node depends on its own historical information and the historical information of its neighbors. Therefore, how to capture both spatial and temporal dependencies simultaneously has become a major challenge. Currently, there are mainly two directions in spatio-temporal modeling research. One method is to concatenate the spatio-temporal features extracted in parallel, and the other method is to sequentially extract spatio-temporal features. Commonly used methods are to use 2D convolutional neural networks (CNNs) or graph convolutional networks (GCNs) to extract spatial features, and use recurrent neural networks or one-dimensional temporal convolutions to obtain temporal features. In fact, CNNs, LSTMs, and Transformers have been used to extract features from both temporal and spatial dimensions. Although they have brought the advantage of introducing spatio-temporal features into the model, these methods face two significant drawbacks. First, these studies have demonstrated the effectiveness of spatio-temporal models. However, extracting features only from the perspective of spatio-temporal models, without considering data, models, and features as a unified entity, and without integrating human prior knowledge into the data, will result in the lack of non-Euclidean space information in the perceptual data. In this case, the model can only extract Euclidean space features based on the existing data, thus losing the more semantically rich non-Euclidean space features. Second, the current spatio-temporal modeling research has poor performance in extracting temporal dependency features. RNN-based methods have time-consuming iterative propagation and gradient explosion / vanishing problems when capturing long sequences, while the Transformer-based method has the advantages of parallel computing and stable gradients. This method uses many layers of self-attention to enhance the expressive ability to extract more meaningful classification features. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a human behavior recognition method based on a spatio-temporal graph neural network for fine-grained data in view of the above-mentioned deficiencies of the prior art, so as to realize the recognition of human behaviors.
[0008] To solve the above technical problem, the technical solution adopted by the present invention is as follows: A human behavior recognition method based on a spatio-temporal graph neural network for fine-grained data, comprising the following steps:
[0009] Step 1: Collect human behavior data and construct a behavior graph and its adjacency matrix;
[0010] Step 1.1: Collect human behavior data and construct a behavior graph;
[0011] Use a wearable sensor system to collect human behavior perception data and construct a behavior graph G=(V, E, A), where V = v1, v2,..., v N represents the set of data features collected by N sensor nodes in the wearable sensor system, E is the set of edges composed of ordered node pairs, and A is the adjacency matrix, where the relationship values are shown in the following formula:
[0012]
[0013] Among them, (v i , v j ) represents a directed edge from sensor node v i to sensor node v j .
[0014] Set the sensor state characteristics of N sensor nodes at each time point t as is the sensor state of the i-th sensor node at time point t, and C is the number of sensor channels; if each sensor has only one channel, that is, C = 1, then the channel-based adjacency matrix and the sensor-based adjacency matrix have the same structure and can be represented by the same behavior graph;
[0015] Step 1.2: Construct the adjacency matrix of the behavior graph based on prior knowledge;
[0016] When constructing the adjacency matrix, to simplify the graph structure, the node types in the behavior graph are defined as sensors and channels, and the device-level types are ignored; therefore, the sensor-level adjacency matrix and the channel-level adjacency matrix are defined;
[0017] The construction of the sensor-level adjacency matrix takes the sensor as the smallest unit of the nodes in the graph, so that a homogeneous graph can be constructed based on the nodes, where each node corresponds to a different position on the human body where the sensor is worn; a fully connected graph is used to construct the local correlation; during the calculation process, the sensor data is concatenated into a one-dimensional vector as the feature representation of the node; then the processed human behavior perception data generates a two-dimensional adjacency matrix and a two-dimensional sensor node state characteristic representing N sensors with C channels each, and the time series length of each channel is T;
[0018] The construction of the channel-level adjacency matrix takes the channel as the smallest unit of the nodes in the graph, and uses a fully connected graph to construct the adjacency matrix of multiple channels within the same sensor; the construction of the channel-level adjacency matrix requires expanding all channels to obtain a two-dimensional adjacency matrix and a two-dimensional sensor node state characteristic
[0019] Step 2: Add virtual nodes to further construct the behavior graph and its adjacency matrix;
[0020] First, add N' sensor-level virtual nodes to the behavior graph, which will cause the dimension of the adjacency matrix to become The dimension of the sensor node state characteristic becomes On the above corresponding adjacency matrix A, a global virtual node at the graph level is further introduced, which will add an additional 2*(N*C + N′) directed edges. The sensor node features are The initial value of the virtual sensor node features is the average of all adjacent child nodes; for the convenience of subsequent model description, the number of nodes in the behavior graph finally formed by adding virtual nodes is uniformly denoted as M, the node feature dimension, i.e., the time sampling dimension, is denoted as T, and the adjacency matrix is simplified as The sensor node state features are denoted as
[0021]
[0022] Step 3: Build a spatio-temporal graph Transformer model, and use GAT, Transformer encoder, and fully connected module to build an end-to-end deep learning model for human behavior recognition;
[0023] The spatio-temporal graph Transformer model uses GAT to aggregate spatial information from the collected multi-channel human behavior perception data; then the spatially aggregated information is fed into the Transformer encoder, and the spatio-temporal features in the spatially aggregated information are extracted by using the multi-head self-attention mechanism of the Transformer encoder; finally, the fully connected layer is used to classify the spatio-temporal features of various behaviors.
[0024] The spatio-temporal graph Transformer model consists of a spatial encoder layer, a spatio-temporal encoder layer, and a classifier; the spatial encoder layer takes the labeled sensor data as input, the sensor data is organized in the form of a graph, the sensors are used as nodes, the behavior perception data read by the sensors is used as node state features, and the relationship between sensors is used as the adjacency matrix of the graph. The embedded output of the spatial encoder is used as a high-level representation of the behavior features; the spatio-temporal encoder layer uses the Transformer architecture to extract spatio-temporal fusion features from the fused embedded features of time, space, and content information; the classifier classifies the extracted spatio-temporal fusion features and is trained using the pre-labeled behavior labels to enable the model to complete the behavior recognition task.
[0025] Step 3.1: Build a spatial encoder;
[0026] The spatial encoder is composed of two layers of GAT stacked and uses a residual layer for skip connection. The spatial encoder uses GAT to learn the spatial correlation relationship of sensor data and captures the spatial dependence between sensors through the attention mechanism; the form of the attention mechanism used by GAT assigns different attention weights to the neighbors within one hop; GAT performs representation learning on the state features of various sensor nodes for the spatial feature representation learning;
[0027] First, the temporal features sensed by the sensor nodes and the spatial features represented by the adjacency matrix are fed into the GAT network, and a learnable weight matrix is used to map the sensor node features into a high-dimensional space and share the weights of the matrix. As the output of the GAT, T′ is the dimension of the mapped features.
[0028] The importance of neighbor node j to node i is weighted using the attention weights, as shown in the following formula:
[0029]
[0030] where e ij represents the importance weight of neighbor node j to node i, u is the linear transformation matrix, T represents the matrix transpose, || represents vector concatenation, represents the set of neighbors of i;
[0031] After obtaining the attention weights, the softmax function is used to normalize them, as shown in the following formula:
[0032]
[0033] where α ij is the normalized attention weight coefficient of neighbor node j of i;
[0034] Finally, the normalized attention weight coefficients are used to weight-aggregate the adjacent nodes, the ReLU is used as the non-linear activation function, and the aggregated sensor node features are converted into the output of the first layer of GAT, as shown in the following formula:
[0035]
[0036] In the spatial encoder, two GAT layers are stacked, and the residual connection is used to accelerate the convergence speed and enable the training of a deeper model;
[0037] V″ = GAT(V) + V = V′ + V (5)
[0038] where V is the output of the first layer of GAT, V′ is the output of the second layer of GAT, V″ is the sum of the output of the first layer of GAT and the output of the second layer of GAT as the residual fusion output of the spatial encoder layer, and this output is defined as the content embedding In addition, V″ is linearly transformed through the Linear function as the spatial embedding As shown in the following formula:
[0039]
[0040] where H dimis the feature dimension after the linear transformation of V″ through the Linear function;
[0041] Finally, embed the content output by the spatial encoder layer into C emb and the spatial embedding S emb as the input of the spatio-temporal encoder layer;
[0042] Step 3.2: Construct the spatio-temporal encoder;
[0043] The input of the spatio-temporal encoder includes the content embedding the spatial embedding and the time embedding Combine these three embeddings to form a spatio-temporal fusion embedding feature, and then use it as the input of the spatio-temporal encoder layer; the content embedding C emb and the spatial embedding S emb are the outputs of the spatial encoding layer, and the time embedding feature T emb is the feature expression of the time series information of the behavior perception data;
[0044] After obtaining the three types of embeddings, the spatio-temporal encoder layer uses a variant of the Transformer encoder to extract the spatio-temporal information in the spatio-temporal fusion embedding feature for behavior classification; the Transformer encoder consists of a multi-head self-attention module, a normalization residual connection module, and a feed-forward network module;
[0045] Before performing feature fusion, the spatio-temporal encoder uses a linear transformation to map the content embedding to a high-dimensional space, as shown in the following formula:
[0046]
[0047] In equation (7), the feature dimension of the content embedding C emb is mapped from M to H dim through the linear model weight coefficient W1; to ensure the fusion of spatio-temporal features, all embedding feature dimensions are H dim ×T′ after the transformation;
[0048] Next, use the Transformer encoder to fuse the time and space features and extract the spatio-temporal fusion features;
[0049] First, after performing the dimension transformation using (7), fuse the content embedding C emb with the time embedding T emb as shown in the following formula:
[0050]
[0051] Secondly, after fusing the Emb features using the multi-head self-attention layer MultiAttn(·) in the time dimension, further fuse the spatial feature S emb , and at the same time add a residual mechanism to ensure the stable convergence of the deep network, as shown in the following formula:
[0052]
[0053] where λ is an adjustable parameter that adjusts the relative contribution of the embedding vector in space;
[0054] Then use layer normalization to normalize the fused features corresponding to the same measurement time;
[0055] Finally, use FFN(·) to further transform the feature Emb′, and at the same time use the residual connection and layer normalization mechanism to ensure the stability of training; finally output a spatio-temporal fused feature representation O
[0056] Step 3.3: Construct a classifier for human behavior classification;
[0057] Use two fully connected layers to form a classifier for human behavior classification; the output of the spatio-temporal encoder is flattened into a one-dimensional vector and sent to the fully connected layer, and the number of neurons in the output layer of the classifier corresponds to the specific classification in the human behavior data.
[0058] The beneficial effects produced by adopting the above technical solutions are as follows: A human behavior recognition method based on a spatio-temporal graph neural network for fine-grained data provided by the present invention
[0059] On the one hand, solve the spatio-temporal modeling problem of human behavior data through a fine-grained spatial structure data construction method, use the graph data structure to store the data, and realize the modeling and storage of Euclidean space data. For the Euclidean space modeling of human behavior perception data, combined with the time and space structure, using the human skeleton graph as the basis for the spatial relationship between sensor nodes, design local fully connected and global fully connected to enhance the connectivity of node message passing, and add global virtual nodes to reduce the propagation distance between multi-hop nodes, so as to learn the global expression of human behavior perception data.
[0060] On the other hand, propose a spatio-temporal graph Transformer network (STGT) based on fine-grained data, which is a new deep learning model for human behavior recognition. The introduction of the STGT model utilizes a new data organization method proposed by the present invention to solve the limitations of existing spatial feature extraction methods. This model enhances the spatial features in the time feature extraction module and realizes effective spatio-temporal feature extraction. Description of the Drawings
[0061] Figure 1Schematic diagram of abstract modeling of human behavior perception data provided by an embodiment of the present invention. Among them, (a) is a multi-sensor spatial perception structure diagram, and (b) is a behavior diagram with added sensor virtual nodes;
[0062] Figure 2 Schematic diagram of adding virtual nodes provided by an embodiment of the present invention;
[0063] Figure 3 Schematic diagram of the spatio-temporal graph Transformer model structure provided by an embodiment of the present invention. Detailed implementation manners
[0064] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0065] In this embodiment, a human behavior recognition method based on a fine-grained data spatio-temporal graph neural network includes the following steps:
[0066] Step 1: Use a wearable sensor system to collect human behavior data and construct a behavior graph and its adjacency matrix;
[0067] Step 1.1: Collect human behavior data and construct a behavior graph;
[0068] Use a wearable sensor system to collect human behavior perception data and construct a behavior graph G=(V, E, A), where V = v1, v2,..., v N represents the set of data features collected by N sensor nodes in a wearable sensor system (such as a three-axis accelerometer sensor, a gyroscope sensor, a magnetometer sensor, etc.). E is the set of edges composed of ordered node pairs, and A is the adjacency matrix, where the relationship values are shown in the following formula:
[0069]
[0070] Among them, (v i , v j ) represents a directed edge from sensor node v i to sensor node v j ;
[0071] Set the sensor state features of N sensor nodes at each time point t as is the sensor state of the i-th sensor node at time point t, and C is the number of sensor channels (i.e., the variable of the reading, such as the acceleration along the x-axis, y-axis, and z-axis); if each sensor has only one channel, that is, C = 1, then the adjacency matrix based on channels and the adjacency matrix based on sensors have the same structure and can be represented by the same behavior graph;
[0072] Human activity recognition (HAR) takes the human activity data detected by sensors as input and outputs specific activity categories after being processed by a model. As Figure 1 shown, data is collected through inertial sensors worn on different parts of the human body, continuously recording the perceptual information related to human activities. These data can be constructed and preprocessed based on prior knowledge and thus used as the input of the model. Compared with vision sensors, wearable inertial sensors have advantages in terms of low power consumption, portability, enhanced privacy, and reduced data and computing resource requirements. Step 1.2: Construct the adjacency matrix of the activity graph based on prior knowledge;
[0073] As Figure 1 (a) shows, wearable inertial sensors capture data for each channel. The process of capturing wearable human activity data may involve multiple devices, each device contains multiple sensors, and each sensor contains multiple channels; when constructing the adjacency matrix, to simplify the graph structure, the node types in the activity graph are defined as sensors and channels, ignoring the device-level types; therefore, the sensor-level adjacency matrix and the channel-level adjacency matrix are defined;
[0074] The construction of the sensor-level adjacency matrix takes the sensor as the smallest unit of the nodes in the graph, so that a homogeneous graph can be constructed based on the nodes, where each node corresponds to a different position on the human body where the sensor is worn; as Figure 1 (b) shows, node connections are made based on the human bone structure. Multiple sensors within the same device cannot discuss specific spatial correlations, so a fully connected graph is used to construct local correlations; this method does not consider multi-channel modeling within the sensor, and the sensor data is concatenated into a one-dimensional vector as the feature representation of the node during the calculation process; then the processed human activity perceptual data generates a two-dimensional adjacency matrix and a two-dimensional sensor node state feature representing N sensors, each sensor has C channels, and the time series length of each channel is T;
[0075] The construction of the channel-level adjacency matrix takes the channel as the smallest unit of the nodes in the graph and uses a fully connected graph to construct the adjacency matrix of multiple channels within the same sensor; the construction of the channel-level adjacency matrix requires expanding all channels to obtain a two-dimensional adjacency matrix and a two-dimensional sensor node state feature
[0076] Step 2: Add virtual nodes to further construct the activity graph and its adjacency matrix;
[0077] When constructing the adjacency matrix based on prior knowledge, the two-dimensional sensor data in the original Euclidean space is reconstructed into graph-based non-Euclidean data. However, the graph neural network model based on the propagation mechanism is limited by the distance between nodes and may not be able to learn the features of nodes that are too far apart. Therefore, this problem can be effectively alleviated by adding virtual nodes. As Figure 2 shown, the adjacent edges between the virtual nodes based on sensors and other virtual sensor nodes connect all channels inside the sensor, providing an abstract representation of the sensor nodes. First, N' sensor-level virtual nodes are added to the behavior graph, which will result in the dimension of the adjacency matrix becoming The dimension of the sensor node state feature becomes On the basis of the above corresponding adjacency matrix A, a global virtual node at the graph level is introduced, which will add an additional 2*(N*C + N') directed edges, and the sensor node feature is By constructing the adjacent edges between the global virtual node and all graph nodes, all nodes can be reached within 2 hops, meeting the requirements of the model for graph information propagation. The initial value of the virtual sensor node feature is the average value of all adjacent child nodes; for the convenience of subsequent model description, the number of nodes in the behavior graph finally formed by adding virtual nodes is uniformly denoted as M, the node feature dimension, i.e., the time sampling dimension, is denoted as T, and the adjacency matrix is simplified as The sensor node state feature is denoted as
[0078] Step 3: Construct a spatio-temporal graph Transformer (STGT) model, and use GAT, Transformer encoder, and fully connected module to build an end-to-end deep learning model for human behavior recognition;
[0079] The spatio-temporal graph Transformer (STGT) model uses GAT to aggregate spatial information from the collected multi-channel human behavior perception data; then the spatially aggregated information is fed into the Transformer encoder, and the spatio-temporal features in the spatially aggregated information are extracted by using the multi-head self-attention mechanism of the Transformer encoder; finally, a fully connected layer is used to classify the spatio-temporal features of various behaviors;
[0080] The spatio-temporal graph Transformer (STGT) model is as Figure 3As shown in the figure, it consists of a spatial encoder layer, a spatio-temporal encoder layer, and a classifier; the spatial encoder layer takes the labeled sensor data as input. The sensor data is organized in the form of a graph, with sensors as nodes, the behavior perception data read by the sensors as node state features, and the relationships between sensors as the adjacency matrix of the graph. The embedded output of the spatial encoder serves as a high-level representation of the behavior features; the spatio-temporal encoder layer uses the Transformer architecture to extract spatio-temporal fusion features from the fusion-embedded features of time, space, and content information; the classifier classifies the extracted spatio-temporal fusion features and is trained using pre-labeled behavior tags to enable the model to complete the behavior recognition task;
[0081] Step 3.1: Construct a spatial encoder;
[0082] The spatial encoder is composed of two layers of GAT stacked and uses a residual layer for skip connection. The spatial encoder uses GAT to learn the spatial correlation relationship of sensor data and captures the spatial dependence between sensors through the attention mechanism; the form of the attention mechanism adopted by GAT assigns different attention weights to the neighbors within one hop, and the attention mechanism is applied to all connections in the model in a shared manner; GAT performs representation learning of spatial features for various sensor node state features for spatial feature representation learning;
[0083] First, the temporal features sensed by the sensor nodes and the spatial features represented by the adjacency matrix are fed into the GAT network, while fusing the spatial features of adjacent sensor nodes and improving the robustness of the features; in order to learn a better spatial feature representation, a learnable weight matrix is used to map the sensor node features to a high-dimensional space and share the weights of the matrix, as the output of GAT, and T′ is the dimension of the mapped features;
[0084] The importance of neighbor node j to node i is weighted using the attention weight, as shown in the following formula:
[0085]
[0086] where, e ij represents the importance weight of neighbor node j to node i, u is the linear transformation matrix, T represents matrix transpose, || represents vector concatenation, represents the set of neighbors of i;
[0087] After obtaining the attention weight, the softmax function is used to normalize it. This normalization process can also ensure that the coefficients between different sensor channel readings can be scaled to reflect the probability distribution weights. In formula (3), α ijis the normalized attention weight coefficient of neighbor node j of i, as shown in the following formula:
[0088]
[0089] Finally, the adjacent nodes are weighted and aggregated using the normalized attention weight coefficient, ReLU is used as the non-linear activation function, and the aggregated sensor node features are converted into the output of the first layer of GAT:
[0090]
[0091] Stack two GAT layers in the spatial encoder, use residual connections to accelerate the convergence speed, and be able to train deeper models;
[0092] V″ = GAT(V) + V = V′ + V (14)
[0093] where V is the output of the first layer of GAT, V′ is the output of the second layer of GAT, and V″ is the sum of the output of the first layer of GAT and the output of the second layer of GAT as the residual fusion output of the spatial encoder layer, and this output is defined as the content embedding In addition, V″ is linearly transformed through the Linear function as the spatial embedding As shown in the following formula:
[0094]
[0095] where H dim is the feature dimension after linearly transforming V″ through the Linear function;
[0096] Finally, the content embedding C output by the spatial encoder layer emb and the spatial embedding S emb are used as the input of the spatio-temporal encoder layer;
[0097] Step 3.2: Construct a spatio-temporal encoder;
[0098] The input of the spatio-temporal encoder includes the content embedding the spatial embedding and the time embedding These three embeddings are combined to form a spatio-temporal fusion embedding feature, which is then used as the input of the spatio-temporal encoder layer; the content embedding C emb and the spatial embedding S emb are the outputs of the spatial encoding layer, and the time embedding feature T emb is the feature expression of the time series information of the behavior perception data;
[0099] After obtaining the three types of embeddings, the spatio-temporal encoder layer uses a variant of the Transformer encoder to extract spatio-temporal information in the spatio-temporal fusion embedding features for behavior classification; the Transformer encoder consists of a multi-head self-attention module, a normalization residual connection module, and a feed-forward network module;
[0100] Due to the limited number of sensor channels, the embedding vectors have a low dimension when representing each time step. To solve this problem, before feature fusion, the spatio-temporal encoder uses a linear transformation to map the content embedding into a high-dimensional space, as shown in the following formula:
[0101]
[0102] In Equation (7), the feature dimension of the content embedding C emb is mapped from M to H dim through the linear model weight coefficient W1; to ensure the fusion of spatio-temporal features, after the transformation, all embedding feature dimensions are H dim ×T′;
[0103] After completing the dimensionality conversion of the content embedding C emb , the model has three embedding features C e ′ mb , T emb , and S emb with the same shape, which represent the high-dimensional space representations of content, time, and space information respectively; next, use the Transformer encoder to fuse time and space features and extract spatio-temporal fusion features;
[0104] First, after performing dimensionality transformation using (7), fuse the content embedding C emb with the time embedding T emb , as shown in formula (8):
[0105]
[0106] Second, after fusing the Emb features using the multi-head self-attention layer MultiAttn(·) in the time dimension, further fuse the space feature S emb , and at the same time add a residual mechanism to ensure the stable convergence of the deep network, as shown in the following formula:
[0107]
[0108] where λ is an adjustable parameter that adjusts the relative contribution of the embedding vector in space;
[0109] Then, use layer normalization to normalize the fusion features corresponding to the same measurement time, as shown in the following formula:
[0110]
[0111] Equation (10) represents the specific operation for any input in LayerNorm. Among them, E and Var represent the mean and variance of each time step j respectively, ε is a very small constant used to prevent the denominator from being zero, and γ and β are trainable parameters;
[0112] Finally, use FFN(·) to further transform the feature Emb′, and at the same time use the residual connection and layer normalization mechanism to ensure the stability of training; finally output a spatio-temporal fusion feature representation O, as shown in the following formula:
[0113]
[0114] In Equation (11), FFN accepts an input with a specific dimension, maps it to a higher-dimensional hidden space, applies the GELU activation function, and then maps it back to the original dimension as the output. This process allows FFN to learn the complex non-linear relationship between the input and output features.
[0115] FFN(input) = FC2(GELU(FC1(input))) (21)
[0116] Step 3.3: Construct a classifier for human behavior classification;
[0117] Use two fully connected layers to form a classifier for human behavior classification; the output of the spatio-temporal encoder is flattened into a one-dimensional vector and sent into the fully connected layer, and the number of neurons in the output layer of the classifier corresponds to the specific classification in the human behavior data.
[0118] This embodiment uses four widely used open-source datasets: UCI-HAR, MotionSense, Shoaib, and WISDM to evaluate the performance of the STGT model. Table 1 gives the specific details of each dataset, providing comprehensive information about these datasets.
[0119] Table 1 Dataset Summary
[0120] Name Subject Sampling Frequency Channel Sample Behavior Category Sliding Window Size UCI-HAR 30 50 9 1318272 6 128 Motion 24 50 12 1412865 6 120 Shoaib 10 50 9 63000 7 120 WISDM 36 20 3 1098208 6 120
[0121] Among them, the UCI-HAR dataset records daily behavioral activities. It consists of 6 behavioral categories from 30 subjects, and 9 channels are collected by 3 sensors. The data is sampled at a frequency of 50Hz. In this embodiment, this dataset contains approximately 1,318,272 data points. The training set and the test set are also pre-divided into proportions of 80% and 20% respectively.
[0122] The Motion dataset records daily behavioral activities. It contains 6 behavioral categories of 24 subjects, where 4 sensors collect 12 channels. The data is sampled at a frequency of 50 Hz. In this embodiment, the dataset contains approximately 1,412,865 data points. The data is divided according to the subjects to evaluate the generalization ability of the model. 80% of the subjects are assigned to the training set, and 10% are used for testing and validation.
[0123] The Shoaib dataset records daily behavioral activities. It consists of 7 behavioral categories of 10 subjects, where 3 sensors collect 9 channels. The data is sampled at a frequency of 50 Hz. In this embodiment, the dataset contains approximately 63,000 data points. Other settings are the same as those of the Motion dataset. Only the samples of each category in this dataset are balanced.
[0124] The WISDM dataset records daily behavioral activities. It contains 6 behavioral categories of 36 subjects, where 1 sensor collects 3 channels. The data is sampled at a frequency of 20 Hz. In this embodiment, the dataset contains approximately 1,098,208 data points. Other settings are the same as those of the Motion dataset.
[0125] This embodiment uses the following evaluation metrics for evaluation:
[0126]
[0127] Among them, TP and TN represent the true positive rate and the true negative rate, and FP and FN represent the false positive rate and the false negative rate.
[0128]
[0129]
[0130] This embodiment uses the macro-average F1-score (MacroF1-Score) as an indicator to compare the performance of the method of the present invention with other methods. For this purpose, the F1-score of each category is calculated according to Equation (16), as shown in the following formula:
[0131]
[0132] Among them, |C| represents the number of classes, the F1-score of each class is given the same weight, regardless of the number of their instances, and i = 1,..., C represents the set of classes used in the experiment.
[0133] The micro F1-score (MicroF1-Score) is different from the macro F1-score and is weighted by the number of samples in different categories, as shown in the following formula:
[0134]
[0135] Among them, N c is the number of samples of category C, and N total is the total number of samples.
[0136] In this embodiment, experiments are conducted on the above dataset on a tower server configured with an i9-9900k 3.6GHz CPU, 32GB RAM, and Windows10 (x64) operating system, and the proposed STGT model is implemented and performance-evaluated using the Python language based on the Pytorch and PytorchGeometric frameworks.
[0137] The batch size for model training is set to 64, and the maximum number of training epochs is 100. The model training is optimized using the Adam optimizer, with a learning rate of 1e-3 during training, except for the Shoaib dataset, whose learning rate is 1e-4 and the weight decay is 1e-3. All experiments are run on a 2080Ti GPU.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A human behavior recognition method based on a spatio-temporal graph neural network for fine-grained data, characterized in that: It includes the following steps: Step 1: Collect human behavior data and construct a behavior graph and its adjacency matrix; Step 2: Add virtual nodes and further construct a behavior graph and its adjacency matrix; Step 3: Construct a spatio-temporal graph Transformer model, and use GAT, Transformer encoder, and fully connected module to build an end-to-end deep learning model for human behavior recognition; The spatio-temporal graph Transformer model uses GAT to aggregate spatial information from the collected multi-channel human behavior perception data; then sends the spatially aggregated information into the Transformer encoder to extract spatio-temporal features from the spatially aggregated information by using the multi-head self-attention mechanism of the Transformer encoder; finally, uses a fully connected layer to classify the spatio-temporal features of various behaviors.
2. The human behavior recognition method based on the spatio-temporal graph neural network of fine-grained data according to claim 1, wherein: The specific method of Step 1 is as follows: Step 1.1: Collect human behavior data and construct a behavior graph; The human behavior perception data is collected by using a wearable sensor system, and a behavior graph G=(V, E, A) is constructed, where V = {v1, v2, …, v} N represents the set of data features collected by N sensor nodes in the wearable sensor system, E is the set of edges composed of ordered node pairs, and A is the adjacency matrix, where the relationship values are shown in the following formula: Among them, (v i , v j ) represents a directed edge from sensor node v i to sensor node v j ; Set the sensor state characteristics of N sensor nodes at each time point t as is the sensor state of the i-th sensor node at time point t, and C is the number of sensor channels; if each sensor has only one channel, that is, C = 1, then the channel-based adjacency matrix and the sensor-based adjacency matrix have the same structure and can be represented by the same behavior graph; Step 1.2: Construct the adjacency matrix of the behavior graph based on prior knowledge; When constructing the adjacency matrix, to simplify the graph structure, the node types in the behavior graph are defined as sensors and channels, and the device-level types are ignored; therefore, a sensor-level adjacency matrix and a channel-level adjacency matrix are defined. The construction of the sensor-level adjacency matrix takes sensors as the smallest unit of nodes in the graph, enabling the construction of an isomorphic graph based on the nodes, where each node corresponds to a different position on the human body where a sensor is worn; a fully connected graph is used to construct local correlations; during the calculation process, the sensor data is concatenated into a one-dimensional vector as the feature representation of the node; then the processed human behavior perception data generates a two-dimensional adjacency matrix and a two-dimensional sensor node state feature It represents that there are C channels for each of the N sensors, and the time series length of each channel is T; Channel-level adjacency matrix construction takes channels as the smallest unit of nodes in the graph and constructs the adjacency matrix of multiple channels within the same sensor using a fully connected graph; the construction of the channel-level adjacency matrix requires expanding all channels to obtain a two-dimensional adjacency matrix and two-dimensional sensor node state features 3. The human behavior recognition method based on the spatio-temporal graph neural network of fine-grained data according to claim 2, wherein: The specific method of Step 2 is as follows: First, add N' sensor-level virtual nodes to the behavior graph, which will cause the dimension of the adjacency matrix to become The dimension of the sensor node state feature becomes Then, introduce a graph-level global virtual node to the corresponding adjacency matrix A above, which will add an additional 2*(N*C+N') directed edges. The sensor node feature is The initial value of the virtual sensor node feature is the average of all adjacent child nodes; for the convenience of subsequent model description, uniformly denote the number of nodes in the behavior graph finally formed by adding virtual nodes as M, the node feature dimension, i.e., the time sampling dimension, as T, and the adjacency matrix is simplified as The sensor node state feature is represented as 4. The human behavior recognition method based on a fine-grained data spatio-temporal graph neural network according to claim 3, characterized in that: The spatio-temporal graph Transformer model in Step 3 consists of a spatial encoder layer, a spatio-temporal encoder layer, and a classifier; the spatial encoder layer takes the labeled sensor data as input, the sensor data is organized in a graph form, the sensors are used as nodes, the behavior perception data read by the sensors is used as the node state features, and the relationship between sensors is used as the adjacency matrix of the graph, and the embedded output of the spatial encoder is used as the high-level representation of the behavior features; the spatio-temporal encoder layer uses the Transformer architecture to extract spatio-temporal fusion features from the fused embedded features of time, space, and content information; the classifier classifies the extracted spatio-temporal fusion features and is trained using pre-labeled behavior tags to enable the model to complete the behavior recognition task.
5. A human behavior recognition method based on a fine-grained data spatio-temporal graph neural network according to claim 4, characterized in that: The spatial encoder is stacked with two layers of GAT and uses a residual layer for skip connection. The spatial encoder uses GAT to learn the spatial correlation relationship of sensor data and captures the spatial dependence between sensors through the attention mechanism; The form in which GAT uses the attention mechanism is to assign different attention weights to neighbors within one hop; GAT performs representation learning on the spatial features of the state characteristics of various sensor nodes for spatial feature representation learning; First, the temporal features sensed by the sensor nodes and the spatial features represented by the adjacency matrix are fed into the GAT network, and a learnable weight matrix is used to map the sensor node features to a high-dimensional space and share the weights of the matrix. As the output of the GAT, T′ is the dimension of the mapped features. The importance of neighbor node j to node i is weighted using the attention weight, as shown in the following formula: Among them, e ij represents the importance weight of neighbor node j to node i, u is a linear transformation matrix, T represents matrix transpose, || represents vector concatenation, represents the neighbor set of i; After obtaining the attention weight, it is normalized using the softmax function, as shown in the following formula: Among them, α ij is the normalized attention weight coefficient of neighbor node j of i; Finally, the adjacent nodes are weighted and aggregated using the normalized attention weight coefficient, ReLU is used as the non-linear activation function, and the aggregated sensor node features are converted into the output of the first layer of GAT, as shown in the following formula: In the spatial encoder, two GAT layers are stacked, and residual connections are used to accelerate the convergence speed and enable the training of deeper models; V″ = GAT(V) + V = V′ + V (5) Among them, V is the output of the first-layer GAT, V' is the output of the second-layer GAT, and V'' is the sum of the output of the first-layer GAT and the output of the second-layer GAT, which is defined as the residual fusion output of the spatial encoder layer. This output is defined as the content embedding. In addition, V'' is linearly transformed through the Linear function as the spatial embedding. As shown in the following formula: Among them, H dim is the feature dimension after the linear transformation of V″ through the Linear function; Finally, the content output by the spatial encoder layer is embedded into C emb and the spatial embedding S emb as the input to the spatio-temporal encoder layer.
6. The human behavior recognition method based on a fine-grained data spatio-temporal graph neural network according to claim 5, characterized in that: The input of the spatio-temporal encoder includes a content embedding a spatial embedding and a temporal embedding These three embeddings are combined to form a spatio-temporal fusion embedding feature, which is then used as the input to the spatio-temporal encoder layer; the content embedding C emb and the spatial embedding S emb are the outputs of the spatial encoding layer, and the temporal embedding feature T emb is the feature representation of the temporal information of the behavior perception data; After obtaining the three types of embeddings, the spatio-temporal encoder layer uses a variant of the Transformer encoder to extract spatio-temporal information in the spatio-temporal fusion embedding features for behavior classification; the Transformer encoder consists of a multi-head self-attention module, a normalization residual connection module, and a feed-forward network module; Before performing feature fusion, the spatio-temporal encoder maps the content embedding to a high-dimensional space using a linear transformation, as shown in the following formula: In formula (7), the feature dimension of the content embedding C emb is mapped from M to H dim ; to ensure the fusion of spatio-temporal features, after transformation, all embedded feature dimensions are H dim × T′; Next, the Transformer encoder is used to fuse temporal and spatial features and extract spatio-temporal fusion features; First, after performing dimensionality transformation using (7), the content is embedded into C emb and fused with the time embedding T emb as shown in the following formula: Secondly, after fusing the Emb features using the multi-head self-attention layer MultiAttn(·) in the time dimension, the spatial feature S is further fused. emb Meanwhile, a residual mechanism is added to ensure the stable convergence of the deep network, as shown in the following formula: where λ is an adjustable parameter that adjusts the relative contribution of the embedding vector in space; Then, layer normalization is used to normalize the fusion features corresponding to the same measurement time; Finally, the FFN(·) is used to further transform the feature Emb′, and the residual connection and layer normalization mechanism are used to ensure the stability of training; finally, a spatio-temporal fusion feature representation O is output.
7. A human behavior recognition method based on a fine-grained data spatio-temporal graph neural network according to claim 6, characterized in that: The spatio-temporal graph Transformer model uses two fully connected layers to form a classifier for human behavior classification; the output of the spatio-temporal encoder is flattened into a one-dimensional vector and sent into the fully connected layer, and the number of neurons in the output layer of the classifier corresponds to the specific classification in the human behavior data.