Complex human body activity identification method based on MSRC-BiGRU-SA

Through the MSRC-BiGRU-SA model, the problem of insufficient multi-scale feature capture in complex human activity recognition in the prior art is solved, and the recognition accuracy and robustness are improved, especially in the recognition ability when dealing with multi-stage and action interactive activities.

CN120296483APending Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510170636.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture multi-scale spatial and temporal features in the recognition of complex human activities, and ignores the fusion of multi-scale features and the capture of key features, resulting in insufficient recognition accuracy and robustness.

Method used

The MSRC module is used to extract multi-scale spatial and temporal features, combine the BiGRU module to capture the front and back dependencies of the time series, and enhance the capture ability of key features through the SA module to build the MSRC-BiGRU-SA model.

Benefits of technology

It significantly improves the classification performance of complex activities, improves the recognition accuracy and robustness of the model, especially when dealing with multi-stage, action interactive activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296483A_ABST
    Figure CN120296483A_ABST
Patent Text Reader

Abstract

The invention discloses a complex human body activity identification method based on MSRC-BiGRU-SA, and the method comprises the following steps: S1, collecting daily activity data including typewriting, writing, coffee drinking, speech, smoking and eating through a wearable sensor, and taking the daily activity data as a data sample; s2, performing data preprocessing on the collected data samples, and dividing the preprocessed data into a training set and a test set according to a certain proportion; s3, an MSRC module, a BiGRU module and a self-attention mechanism module are connected in sequence, and an MSRC-BiGRU-SA model is constructed; dropout layers are respectively added behind the MSRC module and the BiGRU module, and a global average pooling layer is adopted between the self-attention mechanism module and the output layer; and S4, training the MSRC-BiGRU-SA model by adopting the training set, and storing the model with the optimal performance. The method can effectively improve the classification performance of complex human body activities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a complex human activity recognition method based on the combination of multi-scale residual convolution (MSRC), bidirectional gated recurrent unit (BiGRU) and self-attention mechanism (SA). Background Art

[0002] With the rapid development of intelligent wearable devices and Internet of Things technology, people's demand for understanding and analyzing human behavior models is increasing day by day. Human Activity Recognition (HAR) aims to automatically identify and classify human activities by analyzing human actions and behaviors in different scenarios, and it is widely used in health monitoring (such as exercise tracking and sleep monitoring), smartphone applications (such as step counting and navigation), virtual reality (such as hand movement tracking), industrial safety (such as detection of dangerous activities), and other fields.

[0003] In order to improve the recognition accuracy, in recent years, deep learning algorithms have become the mainstream technology in the field of HAR, especially convolutional neural network (CNN) and recurrent neural network (RNN). Many studies have tried to combine the structures of CNN and RNN, such as models like CNN-LSTM, CNN-GRU, CNN-BiLSTM, CNN-BiGRU, etc. to capture the spatial and temporal features of activities. These algorithms have achieved certain success in recognizing simple activities (such as walking, running, sitting, etc.), which are coarse-grained activities. However, when facing complex activities, which are fine-grained activities, especially activities involving multiple stages, actions and interactions such as eating and talking, most of the existing algorithms only focus on the extraction of single-scale spatial and temporal features, ignoring the fusion of multi-scale features and the capture of key features, resulting in insufficient accuracy and robustness in complex activity tasks. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a complex human activity recognition method based on MSRC-BiGRU-SA, which fully extracts the multi-scale spatial and temporal features of sensor data through the MSRC module; fully captures the forward and backward dependencies in the time series through the BiGRU module; and enhances the model's ability to capture key features of complex activities through the SA module, so as to improve the classification performance of complex activities.

[0005] Technical Solution: A complex human activity recognition method based on MSRC-BiGRU-SA includes the following steps:

[0006] S1, collecting daily activity data covering typing, writing, drinking coffee, giving a speech, smoking and eating through wearable sensors as data samples;

[0007] S2. Preprocess the collected data samples, and divide the preprocessed data into a training set and a test set according to a certain ratio;

[0008] S3. Connect the MSRC module, the BiGRU module, and the self-attention mechanism module in sequence to construct the MSRC-BiGRU-SA model; and add Dropout layers after the MSRC module and the BiGRU module respectively, and use a global average pooling layer between the self-attention mechanism module and the output layer;

[0009] S4. Use the training set to train the MSRC-BiGRU-SA model and save the model with the best performance.

[0010] Furthermore, the format of the data samples is a CSV file, where each sample contains the feature values of the time step and the corresponding activity label.

[0011] Furthermore, the implementation steps of preprocessing the collected data samples are as follows:

[0012] S21. Normalize the data samples and scale the data to the range [0, 1];

[0013] S22. Smooth the data using the moving average algorithm, apply a moving window of size 5 time steps to the time series of each feature for average calculation, and keep the original values for the first 4 time steps;

[0014] S23. Concatenate the smoothed data with the original data, and then divide the concatenated data through a moving window to generate a sample dataset. Among them, the sample dataset is divided into a training set and a test set according to a ratio of 7:3.

[0015] Furthermore, the MSRC module extracts multi-scale features from the input time series data through convolutional kernels of multiple different sizes, and the implementation process is as follows:

[0016] First, parallel process the input data through multiple convolutional kernels of different sizes, and continuously integrate the feature information of the previous adjacent channel between convolutional layers to gradually increase the receptive field of each layer; at the same time, there is an original path without a convolutional kernel, concatenate the data after parallel processing with the original path data, and then perform feature fusion through a 1×1 convolution; finally, use a residual connection to directly process the input features through a 1×1 convolutional kernel and then add them to the output of the 1×1 convolutional layer.

[0017] Furthermore, the working method of the BiGRU module is as follows: in the current state, the output of the BiGRU is obtained by combining the input data processed forward and the input data processed backward.

[0018] Furthermore, the self-attention mechanism module performs a weighting process on the feature sequence output by the BiGRU. By calculating the correlation between each feature and other features, the self-attention mechanism can assign higher weights to key features and lower weights to irrelevant features.

[0019] Compared with the prior art, the present invention has the following remarkable effects:

[0020] 1. Aiming at the problem that the prior art is difficult to capture multi-scale spatial and temporal features in complex activities, the present invention adopts the MSRC module structure. Among them, the multi-scale parallel structure can effectively extract multi-scale spatial features in complex human activities and effectively fuse the feature information of the original data; the convolutional layers can capture multi-scale temporal features in complex human activity recognition through a layer-by-layer expansion method, and the 1×1 convolution combined with the residual convolution can effectively fuse the captured multi-scale spatial and temporal features;

[0021] 2. Aiming at the problems of insufficient extraction of key features and inadequate modeling of temporal dependencies in the prior art for complex activities, the present invention introduces the SA module after the BiGRU module, so that the MSRC-BiGRU-SA model can further optimize the feature weight allocation on the basis of capturing the forward and backward dependencies of the time series, enhance the model's attention to key features, and thus effectively improve the classification performance of complex human activities. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the algorithm flowchart of the present invention;

[0023] Figure 2 is the structural diagram of the MSRC module of the present invention;

[0024] Figure 3 is the working diagram of the BiGRU module of the present invention;

[0025] Figure 4 is the overall structural diagram of the model of the present invention;

[0026] Figure 5 is the confusion matrix of the model of the present invention on the complex human activity dataset. DETAILED DESCRIPTION OF THE INVENTION

[0027] The present invention will be further described in detail below with reference to the accompanying drawings of the specification and specific embodiments.

[0028] As Figure 1 shown, the specific embodiment of the present invention includes the following steps:

[0029] Step 1, constructing a dataset;

[0030] In this embodiment, the initial dataset contains time series data generated by various complex human activities. The dataset is collected by two wearable sensors, including an accelerometer and a gyroscope. The collected data samples cover six daily activities such as typing, writing, drinking coffee, giving a speech, smoking, and eating. The format of the data samples is a CSV file, where each sample contains the feature values of the time steps and the corresponding activity labels.

[0031] Step 2, data preprocessing;

[0032] In the data preprocessing stage, first, the data samples are normalized to scale the data to the range [0, 1] to eliminate the impact of data dimensionality differences on the model in advance, ensure the consistency of the scales of different features, and promote the rapid convergence of the model. The normalization formula is:

[0033]

[0034] where x max is the maximum value of the data, x min is the minimum value of the data, X is the dataset matrix after normalization processing, and x is the input dataset sample matrix.

[0035] After the normalization processing, then the moving average algorithm is used to smooth the data to reduce noise interference. Specifically, a moving window of size 5 time steps is applied to the time series of each feature for average calculation, and the original values are maintained for the first 4 time steps. Subsequently, the smoothed data is concatenated with the original data to enhance data diversity. Finally, the concatenated data is segmented through a moving window to generate a sample dataset for model training and testing. Among them, the sample dataset is divided into a training set and a testing set in a ratio of 7:3 to ensure that the model can be effectively trained and evaluated.

[0036] Step 3, build the MSRC-BiGRU-SA model;

[0037] The MSRC (Multi-Scale Residual Convolution) module built in the present invention is as Figure 2As shown, its core function is to perform multi-scale feature extraction on the input time series data through convolutional kernels of various different sizes. This module first processes the input data in parallel through multiple convolutional kernels of different sizes (such as 1×3, 1×5, 1×7, etc.), so as to extract spatial features of different scales. To further enhance the model's feature extraction ability, the convolutional layer can gradually increase the receptive field of each layer by continuously integrating the feature information of adjacent channels. Through this layer-by-layer expansion method, the MSRC module can capture time features of different scales. In addition, it also includes an original path without a convolutional kernel, which helps to retain the feature information in the original data and avoid losing important detailed features during the convolution process. On the basis of feature extraction, to further improve the performance of the MSRC-BiGRU-SA model, the MSRC module performs feature fusion through 1×1 convolution. 1×1 convolution can effectively compress and fuse the features from the paths, map the multi-scale features into a more efficient and compact space, thereby reducing the computational complexity and enhancing the model's expressive ability. Through the feature fusion of 1×1 convolution, the MSRC module can combine multi-scale information while retaining the key features of each scale. In addition, the MSRC module also adopts the design of residual connection. Residual connection can alleviate the gradient vanishing problem in deep networks and accelerate the training process. By directly adding the input features to the output of the convolutional layer, the residual connection helps the flow of information, enabling the network to better learn complex non-linear mappings and effectively improving the training stability and convergence speed of the MSRC-BiGRU-SA model.

[0038] The working mode of the BiGRU (Bidirectional Gated Recurrent Unit) module built in the present invention is as Figure 3 shown, where x and y respectively represent the inputs and outputs at different times, and the subscripts represent different times. A and A ′ respectively represent the forward and backward hidden state parameters. As can be seen from Figure 3 , in the current state, the output of the BiGRU module is obtained by combining the input data processed forward and the input data processed backward. This enables the MSRC-BiGRU-SA model to comprehensively consider the input data in the past and future during learning and evaluation. In the HAR task, using the sensor data at the moments before and after the current action posture captured as the recognition basis can significantly improve the accuracy of action recognition.

[0039] The main function of the SA module (self-attention mechanism module) built in the present invention is to perform weighted processing on the feature sequence output by the BiGRU, so that the MSRC-BiGRU-SA model can automatically focus on the important features related to the classification task. By calculating the correlation between each feature and other features, the self-attention mechanism can assign higher weights to key features and lower weights to irrelevant features. In terms of weight assignment, the self-attention mechanism measures the correlation between features by calculating the dot product similarity between the query vector Q and the key vector K. The higher the dot product similarity, the higher the degree of alignment of the features in the feature space and the stronger their correlation. Subsequently, the similarity is normalized by the Softmax function so that the sum of the attention weights of all features is 1. The specific values of the weights depend on the correlation between features: if a certain feature is more important for the final classification result in the context, its dot product similarity with other features is larger, and after the Softmax transformation, the attention weight corresponding to this feature is relatively higher; otherwise, it is lower. The calculation formula is as follows:

[0040] Q = XW q (2)

[0041] K = XW k (3)

[0042] V = XW v (4)

[0043]

[0044] In the formula, X represents the input feature sequence, Q, K, and V represent the query vector, key vector, and value vector respectively, W q , W k , W v represent the linear transformation matrices, T represents the transpose, d k represents the dimension of the query vector and the key vector, and softmax is the activation function, which is used here to calculate the self-attention weights.

[0045] Specifically, introducing the SA module enables the MSRC-BiGRU-SA model to enhance its ability to focus on key action segments, thereby improving the MSRC-BiGRU-SA model's understanding and recognition ability of complex activities, which also makes up for the deficiency that traditional models are difficult to capture the subtle differences between different complex activities.

[0046] The complete MSRC-BiGRU-SA model built in the present invention is as Figure 4As shown in the figure, in order to reduce the overfitting risk of the MSRC-BiGRU-SA model, the present invention adds Dropout layers after the MSRC module and the BiGRU module respectively. Among them, adding a Dropout layer after the MSRC module can reduce feature redundancy, prevent the MSRC-BiGRU-SA model from over-relying on local features, and improve the robustness of features; adding a Dropout layer after the BiGRU module can prevent the bidirectional RNN structure from overfitting and improve the robustness of time series modeling. Considering the parameter complexity of the MSRC-BiGRU-SA model, the present invention uses a global average pooling layer to replace the fully connected layer, reducing the number of parameters while retaining important feature information.

[0047] Step 4, train the MSRC-BiGRU-SA model;

[0048] Use the training set as input to train the MSRC-BiGRU-SA model. After the data in the training set is input into the network, calculate the cross-entropy loss between the predicted result and the true label. The MSRC-BiGRU-SA model learns by continuously inputting the data in the training set and updates the network parameters to reduce the value of the loss function. When the loss function of the MSRC-BiGRU-SA model converges on the training set or reaches the preset number of training rounds, the training process ends.

[0049] During the training process, to ensure the convergence effect of the MSRC-BiGRU-SA model and effectively measure the difference between the predicted value and the true value, a cross-entropy loss function is used to evaluate the difference between the predicted probability distribution and the true distribution. To improve the stability and efficiency during the training process, the Adam optimizer is used for iteration. The learning rate is set to 0.001, the number of training rounds is set to 300, the batch size is set to 32, and the main parameters of the MSRC-BiGRU-SA model are shown in Table 1. To ensure the reliability and repeatability of the results, these parameters are determined through grid search and cross-validation to obtain the best performance on the dataset.

[0050] Table 1 Layers and parameters of the MSRC-BiGRU-SA model

[0051] Parameter Name Parameter Setting One-dimensional Convolutional Layer filters = 128, kernel size = 3, stride = 1 One-dimensional Convolutional Layer filters = 128, kernel size = 5, stride = 1 One-dimensional Convolutional Layer filters = 128, kernel size = 7, stride = 1 One-dimensional Convolutional Layer filters = 32, kernel size = 1, stride = 1 One-dimensional Convolutional Layer filters = 32, kernel size = 1, stride = 1 Dropout Layer Dropout Rate 0.5 BiGRU unit = 64 Dropout Layer Dropout Rate 0.5 Self-attention Mechanism - Global Average Pooling - Output Layer unit = 6

[0052] Step 5, evaluate the MSRC-BiGRU-SA model;

[0053] In the present invention, the performance of the MSRC-BiGRU-SA model was comprehensively evaluated through a variety of evaluation metrics: recall rate R, accuracy Acc, precision Pre, and F1 score. The evaluation results of the classification metrics of the MSRC-BiGRU-SA model for various complex activities are shown in Table 2. As can be seen from Table 2, the MSRC-BiGRU-SA model of the present invention has obtained relatively high values in all evaluation metrics, indicating that the MSRC-BiGRU-SA model has strong classification ability in the task of identifying complex activities. To further understand the classification effect of the MSRC-BiGRU-SA model, the confusion matrix was also calculated in this embodiment, and the confusion matrix is as Figure 5 shown, and it can be seen from Figure 5 that the MSRC-BiGRU-SA model can accurately identify complex human activities.

[0054] Table 2 Evaluation Table of Model Classification Metrics Unit: %

[0055] Model R Acc Pre F1 MSRC-BiGRU-SA 0.9749 0.9750 0.9757 0.9750

[0056] To further verify the superiority of the present invention, the MSRC-BiGRU-SA model was compared and analyzed with other existing mainstream models. The experimental results are shown in Table 3. Through comparison, it can be observed that the performance of the MSRC-BiGRU-SA model of the present invention is significantly better than those of these classical algorithms such as KNN, DT, and 1D CNN. Compared with the currently mainstream and advanced CNN-LSTM, CNN-GRU, CNN-BiLSTM, CNN-BiGRU, TS-DyConv, and Deep-ConvLSTM, the accuracy has increased by 6.70%, 6.95%, 6.35%, 4.78%, 5.77%, and 2.37% respectively, and the F1 score has increased by 6.65%, 7.06%, 6.47%, 4.49%, 5.79%, and 2.19% respectively. The experimental results show that the MSRC-BiGRU-SA model has significant advantages in processing complex activity data compared with other mainstream models.

[0057] Table 3 Comparison of Experimental Results of Different Algorithms Unit: %

[0058]

[0059]

[0060] In addition, to further analyze the contributions of the proposed MSRC module and SA module in the MSRC-BiGRU-SA model of the present invention to the model performance, an ablation experiment was also designed in this embodiment, and the results of the ablation experiment are shown in Table 4. As can be seen from Table 4, after replacing the traditional CNN with the MSRC module on the basis of the CNN-BiGRU model, the accuracy and F1 score of the CNN-BiGRU model were improved by 2.99% and 2.70% respectively; after introducing the self-attention mechanism on the basis of the CNN-BiGRU model, the accuracy and F1 score of the CNN-BiGRU-SA model were improved by 1.32% and 0.83% respectively; after improving the CNN in the CNN-BiGRU-SA model to the MSRC module, the accuracy and F1 score of the MSRC-BiGRU-SA model were improved by 3.46% and 3.66% respectively; after improving the CNN in the CNN-BiGRU model to the MSRC module and introducing the self-attention mechanism, the accuracy and F1 score of the MSRC-BiGRU-SA model were improved by 4.78% and 4.49% respectively. The experiment shows that the MSRC module can effectively capture the multi-scale time and space features of complex human activities, and the self-attention mechanism can enhance the model's ability to capture the important features of human activities and improve the model performance.

[0061] Table 4 Results of Ablation Experiment Unit: %

[0062] Model Acc F1 CNN-BiGRU 92.72 93.01 MSRC-BiGRU 95.71 95.71 CNN-BiGRU-SA 94.04 93.84 MSRC-BiGRU-SA 97.50 97.50

[0063] The technical means disclosed in the present invention are not limited to the specific technical solutions disclosed in the above embodiments, but also include other technical solutions formed by any combination of the above technical features. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, various forms of improvements and adjustments can be made to the technical details, and these improvements and adjustments are also included in the protection scope of the present invention.

Claims

1. A complex human activity recognition method based on MSRC - BiGRU - SA, characterized in that, The steps are as follows: S1. Collect daily activity data covering typing, writing, drinking coffee, giving a speech, smoking, and eating through wearable sensors as data samples; S2. Perform data preprocessing on the collected data samples, and divide the preprocessed data into a training set and a test set according to a certain ratio; S3. Connect the MSRC module, the BiGRU module, and the self-attention mechanism module in sequence to construct the MSRC-BiGRU-SA model; and add Dropout layers after the MSRC module and the BiGRU module respectively, and use a global average pooling layer between the self-attention mechanism module and the output layer; S4. Use the training set to train the MSRC-BiGRU-SA model and save the model with the best performance.

2. The complex human activity recognition method based on MSRC-BiGRU-SA according to claim 1, wherein The format of the data samples is a CSV file, where each sample contains the feature values of the time steps and the corresponding activity labels.

3. The complex human activity recognition method based on MSRC-BiGRU-SA according to claim 1, characterized in that, The implementation steps for performing data preprocessing on the collected data samples are as follows: S21. Normalize the data samples and scale the data to the interval [0, 1]; S22. Smooth the data using the moving average algorithm, apply a moving window of size 5 time steps to the time series of each feature for average calculation, and keep the original values for the first 4 time steps; S23. Concatenate the smoothed data with the original data, and then divide the concatenated data through a sliding window to generate a sample data set. Among them, the sample data set is divided into a training set and a test set according to a ratio of 7:

3.

4. The complex human activity recognition method based on MSRC - BiGRU - SA according to claim 1, characterized in that, The MSRC module performs multi-scale feature extraction on the input time series data through convolutional kernels of various different sizes. The implementation process is as follows: First, parallel process the input data through multiple convolutional kernels of different sizes, and continuously incorporate the feature information of the previous adjacent channel between convolutional layers to gradually increase the receptive field of each layer; at the same time, there is an original path without a convolutional kernel, concatenate the data after parallel processing with the data on the original path, and then perform feature fusion through a 1×1 convolution; finally, use a residual connection to directly process the input features through a 1×1 convolutional kernel and then add them to the output of the 1×1 convolutional layer.

5. The complex human activity recognition method based on MSRC-BiGRU-SA according to claim 1, wherein The working method of the BiGRU module is as follows: In the current state, the output of the BiGRU is obtained by combining the input data processed forward with the input data processed backward.

6. The complex human activity recognition method based on MSRC - BiGRU - SA according to claim 1, wherein, The self-attention mechanism module performs weighted processing on the feature sequence output by the BiGRU. By calculating the correlation between each feature and other features, the self-attention mechanism can assign higher weights to key features and lower weights to irrelevant features.

Citation Information

Patent Citations

  • Human body activity identification method based on residual shrinkage network

    CN117523672A