A Human Action Recognition Method Based on Self-Attention Mechanism and Bi-GRU
By combining the self-attention mechanism and Bi-GRU's Encoder-Decoder model, the problems of low accuracy and high computational complexity in the prior art are solved, and efficient human movement recognition and feature extraction are achieved.
Patent Information
- Application Number
- CN202211304941.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-10-24
AI Technical Summary
The existing human body movement recognition technology based on convolutional neural networks has problems such as insufficient spatial feature extraction, high computational complexity, large number of parameters, and difficulty in extracting time features with long time intervals, resulting in low accuracy of human body movement recognition.
The Encoder-Decoder model combined with Bi-GRU is adopted to extract global time-related features through the self-attention mechanism and splice them with the original input data to ensure that Bi-GRU can extract local temporal features, thereby achieving complete extraction of time-domain features.
It improves the accuracy of human body movement recognition, reduces the complexity of the algorithm and the amount of parameters, simplifies the model structure, reduces the consumption of computing resources, and ensures the integrity of time domain features.
Smart Images

Figure CN115690906B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of human action recognition, and particularly relates to a human action recognition method based on self-attention mechanism and Bi-GRU. Background Art
[0002] Human action recognition refers to classifying motions into predefined human action categories according to the data obtained by sensors. It has played a very important role in fields such as health monitoring systems, telemedicine, and motion detection. Human action recognition based on inertial sensors has the advantages of being free from external interference, not being restricted by scenes, and having strong anti-interference ability, and is more suitable for daily exercises and military applications.
[0003] The proposal of deep learning has brought breakthrough progress to machine learning and also brought a new development direction to human action recognition. Deep learning can automatically learn deep features from raw data, solving the problem that the feature extraction of traditional machine learning depends on the prior knowledge of researchers, resulting in poor algorithm generalization ability.
[0004] The human action recognition technology based on convolutional neural network and recurrent neural network is one of the most widely used technologies in the current human action recognition technology based on deep learning. Convolutional neural network can extract spatial features, and recurrent neural network can extract temporal features. However, there are still the following problems: 1. For a task with strong temporal correlation such as human action recognition, the spatial features extracted by the convolutional network are not effective enough, resulting in low recognition accuracy for complex actions. 2. The convolutional network has too high computational complexity and too many parameters. 3. The recurrent neural network is difficult to extract the temporal features between data with a long time interval, resulting in insufficient accuracy of human action recognition. Therefore, a new feature extraction and recognition method needs to be proposed to improve the accuracy of human action recognition and reduce the algorithm complexity.
[0005] The present invention has an essential difference from the patent CN114639169A. The data source of the present invention is inertial sensing, while CN114639169A uses WiFi, and the present invention does not use complex convolutional algorithms.
[0006] The present invention extracts global temporal correlation features through the self-attention mechanism. To ensure that Bi-GRU can extract the local temporal order features of the original data, the output of the self-attention mechanism is concatenated with the original input data. Then, the local temporal order features are extracted through Bi-GRU, realizing the complete extraction of time-domain features. At the same time, the combination of the self-attention mechanism and Bi-GRU has a simple structure and low number of parameters, solving the problems of large number of parameters and complex structure of the convolutional network. Summary of the Invention
[0007] The present invention aims to solve the problems of the above prior art. A human action recognition method based on self-attention mechanism and Bi-GRU is proposed. The technical solution of the present invention is as follows:
[0008] A human action recognition method based on self-attention mechanism and Bi-GRU, comprising the following steps:
[0009] S1: Record the inertial sensor data of human actions, and intercept the data and the corresponding action category labels through a sliding window.
[0010] S2: Construct an Encoder-Decoder model; the Encoder-Decoder model includes an Encoder and a Decoder. Input the data into the Encoder for encoding. Extract the temporal correlation features between the input data through the multi-head self-attention layer in the Encoder, and then splice them with the original input data.
[0011] S3: Decoder decoding: The Decoder includes a bidirectional gated recurrent unit Bi-GRU, a fully connected layer, and a Softmax layer. Input the output data of the Encoder into the bidirectional gated recurrent unit Bi-GRU for further temporal order feature extraction; the fully connected layer integrates the features into a vector, and the Softmax layer converts the output of the fully connected layer into a probability distribution.
[0012] S4: Input the output features of Bi-GRU into the fully connected layer to obtain an output vector. The dimension of the output vector is the total number of classification labels, and the value of the Nth dimension of the vector is the possibility that the action corresponding to the input inertial sensor data is the Nth action.
[0013] S5: Train the model according to the sample data, and then input the inertial sensor data with unknown classification labels into the trained model to obtain its human action category.
[0014] Further, the S1 specifically includes:
[0015] Use the inertial sensor located on the torso to record the inertial sensor time series data of human actions, and set a sliding window of a certain length to intercept the corresponding length of data and the human action category corresponding to each sliding window.
[0016] Further, the multi-head self-attention layer in step S2 includes three fully connected layers: query, key, and value. The input data passes through these three fully connected layers to obtain the Q, K, and V matrices respectively, and then the Attention-Score attention score matrix is obtained through further calculation. To ensure that Bi-GRU can learn the time-domain features of the original data, the Attention-Score matrix is concatenated with the original data on the last dimension to obtain the output of the Encoder.
[0017] Further, the calculation formula of the Attention-Score matrix is as follows:
[0018]
[0019] where Head_size represents the dimension size of each head in Multi-Head, and Softmax represents the Softmax function, which is calculated for each row of the matrix. The Softmax formula is as follows:
[0020]
[0021] where y a represents the value of the a-th column in the a-th row of the Attention-Score matrix, y b represents the value of the b-th column in the a-th row of the matrix, and w represents the number of columns of the matrix.
[0022] Further, a Softmax layer is connected after the fully connected layer. The Softmax layer uses the Softmax formula to classify the probability Q(i|x) of the sensor time series data x currently input to the Encoder-Decoder model as the class label i according to the vector output by the fully connected layer. The Softmax formula is as follows:
[0023]
[0024] where z i represents the output of the i-th neuron in the last fully connected layer corresponding to the input sequence x. Among them, z c represents the output of the c-th neuron in the fully connected layer. The N-th dimension value is the probability that the action corresponding to the inertial sensor data within the input sliding window is the N-th type of action. Among them, Softmax(z i ) = Q(i|x);
[0025] Select the action i corresponding to the maximum Q(i|x) as the human action recognition result.
[0026] If Softmax(z i) is the maximum value of the Softmax function result, then the action recognition result corresponding to the input data x is the action of the i-th class label.
[0027] Furthermore, the loss function adopts the balanced cross-entropy function:
[0028]
[0029] where the first half of the right side of the equation is the balanced cross-entropy loss function, and α i represents the loss weight of the i-th action, N represents the number of action categories, P represents the probability distribution after converting the true label into a one-hot encoding, and Q represents regarding the vector output by the model as the action probability distribution; P(x ji ) represents the probability of the i-th action in the true label corresponding to the j-th input sequence x, and Q(x ji ) represents the probability of the i-th action in the model output corresponding to the j-th input sequence x; by assigning different loss weights, the problem of unbalanced sample sizes in the dataset can be solved; the second half is the L2 regularization term; where λ is the regularization coefficient, θ represents the set of learnable parameters in the algorithm, and m is the number of learnable parameters in the algorithm.
[0030] The advantages and beneficial effects of the present invention are as follows:
[0031] The Encoder-Decoder model of the present invention is a neural network model with a simple and lightweight network structure. Different from ordinary human action recognition methods based on deep recurrent neural network learning, the present invention first encodes through the self-attention mechanism in the Encoder, extracts the global time correlation features between data regardless of time intervals, and solves the disadvantage that it is difficult for recurrent neural networks to extract the time correlation features between data with long time intervals. Secondly, to ensure that the Bi-GRU can learn the time domain features of the original data, the Attention-Score matrix is concatenated with the original data on the last dimension to obtain the output of the Encoder. The output of the Encoder is then passed through the gated recurrent unit of the Decoder to extract the time sequence features of the data, improving the accuracy of human action recognition. The present invention can efficiently process inertial sensor data and can automatically learn complete and effective time sequence features from sensor data. At the same time, the present invention only uses recurrent neural networks and does not use convolutional neural networks, has a relatively simple structure and a small number of parameters, so the calculation is simpler and consumes less computer resources. Concatenating the Attention-Score matrix with the original data on the last dimension ensures the integrity of the time domain features of the original data and improves the recognition accuracy. The present invention can provide a new perspective and new thinking for human action recognition and contribute to the development of human action recognition. Description of the Drawings
[0032] Figure 1 It is a schematic structural diagram of the Encoder-Decoder model provided by the preferred embodiment of the present invention.
[0033] Figure 2 It is a flowchart of the method implementation of the embodiment of the present invention.
[0034] Figure 3 It is a flowchart of the generation process of the Encoder-Decoder model.
[0035] Figure 4 It is a method for calculating the Attention-Score matrix of the embodiment of the present invention. Specific embodiments
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0037] The technical solution of the present invention to solve the above technical problems is:
[0038] First, the present invention provides a human action recognition method based on self-attention mechanism and Bi-GRU, as Figure 2 shown, including collecting and processing data, and inputting the data into the model to obtain the human action recognition result. The specific steps of model training are as Figure 3 shown, including the following steps 1, 2, and 3:
[0039] Step 1: Use the inertial sensor located on the torso to record the inertial sensor time series data of the human action, and set a sliding window of a certain length to intercept the corresponding length of data and the human action category corresponding to each sliding window;
[0040] Step 2: Construct an Encoder-Decoder model;
[0041] As Figure 1 shown, the Encoder-Decoder model includes an Encoder and a Decoder. Among them, the Encoder: includes a Multi-Head-Self-Attention layer; the Decoder: includes a bidirectional gated recurrent unit network, a fully connected layer, and a Softmax layer;
[0042] In the present invention, the Multi-Head-Self-Attention layer includes three fully connected layers: query, key, and value. The input data passes through these three fully connected layers to obtain the Q, K, and V matrices respectively. Then, through further calculation, the Attention-Score matrix is obtained. To ensure that the Bi-GRU can learn the time-domain features of the original data, the Attention-Score matrix is concatenated with the original data on the last dimension to obtain the output of the Encoder. The Decoder includes a bidirectional gated recurrent unit to further extract the chronological features; a fully connected layer to integrate the output of the bidirectional gated recurrent unit and output a vector with the dimension equal to the total number of classification labels; a Softmax layer. The output of the fully connected layer passes through the Softmax function to obtain a vector with the dimension equal to the total number of classification labels. The value of the Nth dimension of the vector is the probability that the action corresponding to the inertial sensor data within the input sliding window is the Nth type of action.
[0043] The first module of the Encoder-Decoder model is the Encoder, which includes the Multi-Head-Self-Attention layer. The Multi-Head-Self-Attention layer includes three fully connected layers: query, key, and value. The input data passes through these three fully connected layers to obtain the Q, K, and V matrices respectively, and then through further calculation, the Attention-Score matrix is obtained. The calculation formula of the Attention-Score matrix is:
[0044]
[0045] where Head_size represents the dimension size of each head in Multi-Head, and Softmax represents the Softmax function, which is calculated for each row of the matrix. The Softmax formula is as follows:
[0046]
[0047] where y a represents the value of the a-th column in the m-th row of the Attention-Score matrix, y b represents the value of the b-th column in the m-th row of the matrix, and w represents the number of columns of the matrix.
[0048] Figure 4 shows the calculation process of the Attention-Score matrix with an input time series length of 2, a data dimension of 1×3 for each time step, 4 heads, and a head dimension of 3. The Attention-score matrix calculated by each head ( Figure 3 where S 0Concatenating them by column (etc.) gives the output of the Multi-Head-Self-Attention layer ( Figure 3 in S).
[0049] To ensure that the Bi-GRU can learn the time-domain features of the original data, the Attention-Score matrix is concatenated with the original data on the last dimension to obtain the output of the Encoder. The second module is the Decoder, including a bidirectional gated recurrent unit: used to extract the high-dimensional time-series features of the data; a fully connected layer: integrating and outputting the high-dimensional time-series features obtained by the bidirectional gated recurrent unit into a vector, denoted as (z 1 , z 2 ,..., z N ). The dimension of this vector is the number of action types; the Softmax function: calculating the output of the fully connected layer through the Softmax formula to obtain a vector, and the dimension of this vector is the total number of classification labels. The value of the Nth dimension of the vector is the probability that the action corresponding to the inertial sensor data in the input sliding window is the Nth type of action. The Softmax formula is as follows:
[0050]
[0051] where z i represents the output of the i-th neuron in the last fully connected layer corresponding to the input sequence x, where z c represents the output of the c-th neuron in the fully connected layer, and the value of the Nth dimension is the probability that the action corresponding to the inertial sensor data in the input sliding window is the Nth type of action. Among them, Softmax(z i ) = Q(i|x);
[0052] If Softmax(z i ) is the maximum value of the result of the Softmax function, then the action recognition result corresponding to the input data x is the i-th type of label action, and N is the number of action types.
[0053] Step 3: Train the Encoder-Decoder model according to the sensor time-series data samples intercepted in Step 1 and their corresponding human action category labels, that is, stop training when the loss function value is lower than the set threshold.
[0054] To enable the neural network to learn more discriminative features, the loss function is used to judge the closeness between the actual output and the expected output of the Encoder-Decoder model. The loss function of the present invention adopts the following balanced cross-entropy loss function:
[0055]
[0056] where the first half of the right side of the equation is the balanced cross-entropy loss function, αi denotes the loss weight of the i-th action, n denotes the number of samples in one training, N denotes the number of action types, P denotes the probability distribution after converting the true label into a one-hot encoding, and Q denotes regarding the vector output by the model as the action probability distribution. P(x ji ) represents the probability of the i-th action in the true label corresponding to the j-th input sequence x, and Q(x ji ) represents the probability of the i-th action in the model output corresponding to the j-th input sequence x. By assigning different loss weights, the problem of unbalanced sample sizes in the dataset can be solved. The latter part is the L2 regularization term, which helps to alleviate the overfitting of the algorithm. Among them, λ is the regularization coefficient, θ represents the set of learnable parameters (weights and biases) in the algorithm, and m is the number of learnable parameters in the algorithm.
[0057] Step 4: Use the trained Encoder-Decoder model to identify and classify human actions.
[0058] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0059] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0060] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising said element.
[0061] The above embodiments should be understood as being only for the purpose of illustrating the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A human action recognition method based on self-attention mechanism and Bi-GRU, characterized in that, it includes the following steps: S1: Record the inertial sensor data of human actions, and intercept the data and the corresponding action category labels of the data through a sliding window; S2: Construct an Encoder-Decoder model; the Encoder-Decoder model includes an Encoder and a Decoder. Input the data into the Encoder encoder for encoding. Extract the temporal correlation features between the input data through the multi-head self-attention layer in the Encoder encoder, and then splice them with the original input data; S3: Decoder decoding: The Decoder decoder includes a bidirectional gated recurrent unit Bi-GRU, a fully connected layer, and a Softmax layer. Input the output data of the Encoder into the bidirectional gated recurrent unit Bi-GRU for further temporal order feature extraction; the fully connected layer integrates the features into a vector, and the Softmax layer converts the output of the fully connected layer into a probability distribution; S4: Input the output features of Bi-GRU into the fully connected layer to obtain an output vector. The dimension of this output vector is the total number of classification labels. The value of the Nth dimension of the vector is the possibility that the action corresponding to the input inertial sensor data is the Nth action; S5: Train the model according to the sample data, and then input the inertial sensor data with unknown classification labels into the trained model to obtain its human action category; The multi-head self-attention layer in S2 includes three fully connected layers: query, key, and value. The input data respectively obtains Q, K, and V matrices through these three fully connected layers, and then obtains the Attention-Score attention score matrix through further calculation. To ensure that Bi-GRU can learn the temporal domain features of the original data, splice the Attention-Score matrix with the original data on the last dimension to obtain the output of the Encoder; The calculation formula of the Attention-Score matrix is: where Head_size represents the dimension size of each head of Multi-Head, and Softmax represents the Softmax function, which is calculated for each row of the matrix. The Softmax formula is as follows: where y a represents the value of the a-th column in a certain row of the Attention-Score matrix, and y b represents the value of the b-th column in a certain row of the matrix, and w represents the number of columns of the matrix.
2. The human action recognition method based on self-attention mechanism and Bi-GRU according to claim 1, characterized in that, the S1 specifically includes: Use the inertial sensor located on the torso to record the inertial sensor time series data of human actions, and set a sliding window of a certain length to intercept the corresponding length of data and the human action category corresponding to each sliding window.
3. The human action recognition method based on self-attention mechanism and Bi-GRU according to claim 1, characterized in that, The fully connected layer is followed by a Softmax layer. The Softmax layer uses the Softmax formula to calculate the probability Q(i|x) of classifying the current input sensor time series data x of the Encoder-Decoder model into label i based on the vector output by the fully connected layer. The Softmax formula is as follows: where z i represents the output of the i-th neuron in the last fully connected layer corresponding to the input sequence x, where z c represents the output of the c-th neuron in the fully connected layer, N is the number of action types, and the value of the N-th dimension is the probability that the action corresponding to the inertial sensor data within the input sliding window is the N-th type of action. Among them, Softmax(z i ) = Q(i|x); Select the action i corresponding to the maximum Q(i|x) as the human action recognition result. If Softmax(z i ) is the maximum value of the result of the Softmax function, then the action recognition result corresponding to the input data x is the action of the i-th class label.
4. A human action recognition method based on self-attention mechanism and Bi-GRU according to claim 3, characterized in that, The loss function uses a balanced cross-entropy function: Among them, the first half of the right side of the equation is the balanced cross-entropy loss function, and α i represents the loss weight of the i-th action, N represents the number of action types, P represents the probability distribution after converting the true label into a one-hot encoding, and Q represents regarding the vector output by the model as the action probability distribution; P(x ji ) represents the probability of the i-th action in the true label corresponding to the j-th input sequence x, and Q(x ji ) represents the probability of the i-th action in the model output corresponding to the j-th input sequence x; by assigning different loss weights, the problem of unbalanced sample sizes in the dataset can be solved; the second half is the L2 regularization term; where λ is the regularization coefficient, θ represents the set of learnable parameters in the algorithm, and m is the number of learnable parameters in the algorithm.