Human body activity recognition improvement method and system based on wearable sensor

By constructing a wearable sensor human activity recognition model based on residual network and multi-head attention mechanism, the problems of insufficient local feature capture, gradient disappearance and generalization in the prior art are solved, and high-precision and robust human activity recognition are achieved, which is suitable for medical monitoring and motion analysis.

CN120277535AActive Publication Date: 2025-07-08BEIJING INFORMATION SCI & TECH UNIV
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202510421940.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing human activity recognition technology based on wearable sensors is insufficient in fine local feature capture, gradient vanishing problems, symmetric and asymmetric activity recognition, insufficient generalization, noise sensitivity and hyperparameter dependence, and it is difficult to perform well in complex activity classification tasks.

Method used

Combining the residual network, multi-head attention mechanism and improved Pelican optimization algorithm, a fusion model of local spatiotemporal and global spatiotemporal dependency characteristics is constructed. Through data preprocessing, convolutional neural network, bidirectional gated recurrent neural network, multi-head attention mechanism and improved Pelican optimization algorithm, model parameters are optimized to improve recognition accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of human activity recognition, can maintain high-precision recognition performance in complex scenarios, and is suitable for medical monitoring and motion analysis fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277535A_ABST
    Figure CN120277535A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human body activity recognition, in particular to an improved human body activity recognition method and system based on a wearable sensor, and the method comprises the following steps: collecting the time sequence original data of human body activities through the wearable sensor, and obtaining a time sequence feature tensor to construct local space-time features; constructing a first human body activity recognition improved model in combination with residual connection to obtain weighted fusion features of the human body activity; constructing a second human body activity recognition improved model in combination with an improved multi-head attention mechanism so as to obtain global space-time dependency features of the time sequence original data; optimizing the second improved human body activity recognition model according to an improved pelican optimization algorithm to obtain a human body activity recognition optimization model; and optimizing the global space-time dependency feature to obtain a space-time dependency optimization feature, and realizing recognition improvement of the human body activity. The invention provides an improved technical scheme for human body activity recognition with both efficiency and precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human activity recognition, and in particular to an improved method and system for human activity recognition based on wearable sensors. Background Art

[0002] With the intensification of aging, healthy aging has become an important issue, providing new opportunities for the application of human activity recognition technology (HAR). By analyzing human activity signals, human activity recognition technology can identify movement behaviors, showing important application value and research potential in fields such as medical health, smart home, and human-computer interaction. The human activity recognition system based on artificial intelligence is mainly divided into two categories: based on computer vision and wearable sensors. Compared with the former, the latter is more widely used due to its advantages of low cost, no privacy restrictions, and no environmental interference.

[0003] In recent years, the rapid development of deep learning technology has provided new possibilities for automatic feature extraction, improving the recognition accuracy and practicality of human activity recognition systems. However, although many studies have improved classification performance by improving feature engineering and constructing deep learning models, there are still many deficiencies, including: lack of fine capture of local features; over-reliance on stacked network structures leading to gradient vanishing or explosion problems; focusing on the recognition of symmetric and asymmetric activities, being insufficiently comprehensive in dealing with complex activity classification tasks and having insufficient generalization; relying on a large amount of data, being sensitive to noise, and having performance highly dependent on hyperparameter settings, with the accuracy performance being lacking.

[0004] In view of this, the present invention combines a residual network to solve the challenges of high-dimensional and long-term wearable device data in human activity recognition; introduces a multi-head attention mechanism based on sine and cosine periodic functions to generate time encoding, enhancing the extraction of key time step features; at the same time, aiming at the problems of the traditional Pelican Optimization Algorithm (POA) in terms of insufficient global search ability, poor adaptability, and slow convergence speed, an improvement scheme is proposed, aiming to refine the adjustment of key hyperparameters and enhance the collaborative effect between modules; strengthens the understanding of global features and the ability to capture important information, enhances its ability to effectively model long-term dependencies, ensures the transmission and retention of important information, and improves the stability and performance of the model. Summary of the Invention

[0005] Aiming at the defects in the prior art, the present invention provides an improved method and system for human activity recognition based on wearable sensors.

[0006] To achieve the above object, in a first aspect, the present invention provides an improved method for human activity recognition based on wearable sensors, and the method includes the following steps: collecting original time series data of human activities by using wearable sensors, obtaining a temporal feature tensor according to the original time series data to construct local spatio-temporal features of the human activities; based on the local spatio-temporal features, constructing a first improved model for human activity recognition in combination with residual connections to obtain weighted fusion features of the human activities; based on the weighted fusion features, constructing a second improved model for human activity recognition in combination with an improved multi-head attention mechanism to obtain global spatio-temporal dependence features of the original time series data; optimizing the second improved model for human activity recognition according to an improved pelican optimization algorithm to obtain an optimized model for human activity recognition; optimizing the global spatio-temporal dependence features through the optimized model for human activity recognition to obtain spatio-temporal dependence optimized features, so as to achieve the improvement of the recognition of the human activities. The present invention significantly improves the accuracy and robustness of human activity recognition by constructing local spatio-temporal features and multi-stage optimization models; the temporal feature tensor effectively fuses the time dynamics and spatial correlation of sensor data, and finely depicts the multi-dimensional motion patterns of complex activities; the weighted fusion features combined with residual connections enhance the model's ability to express deep features and avoid the problem of gradient disappearance; the improved multi-head attention mechanism captures spatio-temporal dependence relationships at different time scales in parallel, improving the ability to extract long-period action features; introducing an improved pelican optimization algorithm to achieve global optimization of model parameters, effectively avoiding local optima while reducing computational complexity; the finally constructed optimized model for human activity recognition maintains high-precision recognition performance in complex scenarios through a multi-level feature enhancement mechanism; it has strong applicability in the fields of medical monitoring and sports analysis, providing a solution that takes into account both computational efficiency and model depth for human activity recognition of wearable devices.

[0007] Optionally, collecting the original time series data of human activities using wearable sensors, and obtaining a time series feature tensor according to the original time series data to construct the local spatio-temporal features of the human activities, including: performing data preprocessing on the original time series data to obtain the time series feature tensor, where the data preprocessing includes Gaussian filtering, sliding window partitioning, and normalization; extracting features from the time series feature tensor based on a convolutional neural network to construct the local spatio-temporal features. The present invention improves the accuracy and adaptability of human activity recognition through data preprocessing and deep feature extraction; Gaussian filtering effectively eliminates high-frequency noise interference in the original sensor data and retains key action signals; sliding window partitioning converts continuous time series data into overlapping segments, which can capture local details of dynamic changes while maintaining the continuity characteristics of actions; normalization processing eliminates the dimensional difference of different sensors and improves the consistency of data distribution; by constructing a time series feature tensor, multi-dimensional sensor signals are deeply fused in the spatio-temporal dimension to form a structured representation depicting the spatio-temporal correlation of human movement; feature extraction is performed based on a convolutional neural network, and more discriminative motion features are abstracted layer by layer through multiple convolutional kernels, thereby effectively distinguishing complex actions with high similarity; providing high-quality and low-redundancy feature inputs for subsequent model construction, taking into account both noise robustness and feature expression ability.

[0008] Optionally, based on the local spatio-temporal features, combining residual connections to construct a first improved human activity recognition model to obtain the weighted fusion features of the human activities, including: parsing the local spatio-temporal features according to a time distribution layer to obtain forward time series information and backward time series information; performing weighted splicing on the forward time series information and the backward time series information based on a weighted splicing layer to obtain a spliced output feature; according to the first improved human activity recognition model, fusing the local spatio-temporal features and the spliced output feature to obtain the weighted fusion features. The present invention enhances the spatio-temporal correlation and discriminative ability of human activity features through a bidirectional time series modeling and residual feature fusion mechanism; the time distribution layer captures the action evolution trend and historical correlation respectively by parsing the forward and backward time series information, forming complementary time dynamic representations; the weighted splicing layer dynamically adjusts the contribution weights of bidirectional features through learnable parameters, balances the sensitivity of bidirectional information in sudden actions, and realizes adaptive feature fusion; the residual network performs weighted fusion of the original local spatio-temporal features and the spliced output feature, retaining both shallow detail information and fusing deep abstract semantics; making up for the deficiency of traditional unidirectional models in modeling the temporal dependence of complex actions.

[0009] Optionally, based on the weighted fusion features, a second improved human activity recognition model is constructed by combining an improved multi-head attention mechanism to obtain the global spatio-temporal dependence features of the original time series data, including: using the weighted fusion features as the input features of the improved multi-head attention mechanism; obtaining the attention scores of the input features through linear transformation, and obtaining the multi-head attention distribution according to the weight matrix in combination with the attention scores; based on the input features, obtaining different time encodings based on the sine periodic function and the cosine periodic function, and obtaining the fused multi-scale time encoding according to the time encodings; based on the second improved human activity recognition model, performing weighted summation on the multi-head attention distribution and the fused multi-scale time encoding to obtain the output features, and the output features are used as the global spatio-temporal dependence features. Through the improved multi-head attention mechanism, the present invention significantly enhances the spatio-temporal dependence modeling ability of complex actions; maps the weighted fusion features to different subspaces through parallel linear transformation, enabling the model to simultaneously focus on multiple key spatio-temporal nodes and effectively capture the global correlation across time steps; obtains multi-scale time encodings based on the sine periodic function and the cosine periodic function, adaptively representing the periodic law and non-stationary characteristics of actions, and by fusing time encodings of different frequencies, both short-term action details are retained and long-term behavior patterns are modeled; through weighted summation, the context correlation features extracted by the attention mechanism and the multi-scale time semantics are deeply fused, breaking through the dependence of traditional models on fixed time windows, and still being able to accurately analyze action intentions in strenuous exercise scenarios, solving the problems of local noise interference and temporal drift in wearable sensor data, and having stronger spatio-temporal feature generalization ability.

[0010] Optionally, the obtaining the attention scores of the input features through linear transformation includes: wherein, is the attention score, is the query, is the key, is the value matrix, is the normalization function, is the dimension of the key. Through the dynamic weight allocation mechanism, the present invention improves the efficiency of spatio-temporal feature correlation modeling; calculates the feature similarity through the dot product operation of the query matrix and the key matrix, and performs gradient stabilization processing through the dimension of the key to avoid the problem of excessive numerical values caused by the product of high-dimensional matrices, enhancing the stability of model training; normalization converts the similarity into a probability distribution, enabling the model to adaptively focus on the significant features of key time nodes; uses the value matrix to achieve the aggregated expression of context information; accurately captures the dependence relationship between different timestamps in sensor data, and demonstrates higher computational efficiency and feature discriminability in the spatio-temporal pattern recognition of complex human actions.

[0011] Optionally, obtaining the multi-head attention distribution by combining the attention scores according to the weight matrix includes: Wherein, is the head of the multi-head attention mechanism, is the attention function, is the query, is the key, is the value matrix, is the weight matrix of the th head, is the multi-head attention distribution, is the concatenation function, is the number of heads, is the concatenated weight matrix. Through the multi-head parallel computing and feature subspace decomposition strategy, the present invention significantly enhances the multi-dimensional modeling ability of spatio-temporal dependence relationships; each attention head maps the input to different subspaces through an independent weight matrix, enabling the model to focus on different features such as motion intensity, direction change, and temporal rhythm respectively; after the concatenation function integrates multi-perspective features, the weight matrix performs adaptive dimensionality reduction and fusion on it, retaining both the specific information of each subspace and eliminating redundant interference; breaking through the representation bottleneck of single-head attention, synchronously analyzing multi-granularity associations across channels and across time in complex actions, and improving the anti-interference ability of the model to local noise and the generalization performance of long sequences through parameterized feature recombination.

[0012] Optionally, optimizing the second human activity recognition improved model according to the improved pelican optimization algorithm to obtain a human activity recognition optimized model, including: improving the pelican optimization algorithm based on chaotic mapping, dynamic non-linear inertia weight, vertical crossover operator and Pareto distribution to construct the improved pelican optimization algorithm; optimizing the key hyperparameters of the second human activity recognition improved model through the improved pelican optimization algorithm to obtain the human activity recognition optimized model. The present invention improves the parameter optimization efficiency and generalization performance of the human activity recognition model through a multi-strategy collaborative optimization intelligent search algorithm; the chaotic mapping initializes the population to enhance the global exploration ability of the algorithm and avoid falling into local optimum due to traditional random initialization; the dynamic non-linear inertia weight adaptively adjusts the proportion of global search and local development according to the iteration progress; the vertical crossover operator enhances the population diversity through cross-dimensional information interaction and accelerates the convergence to the optimal solution domain; the Pareto distribution guides the screening of the non-dominated solution set to achieve multi-objective balance optimization of model accuracy and computational complexity; while reducing the manual parameter tuning cost, the improved pelican optimization algorithm enables the model to adaptively match the feature distribution laws of different activity types. Finally, the constructed human activity optimized model has significant improvements in training efficiency, noise robustness and cross-user generalization ability, providing a reliable algorithm support for real-time activity monitoring of wearable devices in dynamic environments.

[0013] Optionally, improving the pelican optimization algorithm based on chaotic mapping, dynamic non-linear inertia weight, vertical crossover operator and Pareto distribution to construct the improved pelican optimization algorithm, including: performing chaotic interference on the pelican population based on chaotic mapping, satisfying the following relationship: where, is the position of the th pelican individual in the -dimensional space, is the perturbation amplitude control parameter, is the sine function, is the chaotic mapping, and the mathematical model of the chaotic mapping satisfies the following relationship: where, is the next position of the pelican individual, is the cosine function, is the order of the chaotic mapping, is the current position of the pelican individual; obtaining the new position of the pelican individual in the pelican population according to the dynamic non-linear inertia weight, and the dynamic non-linear inertia weight satisfies the following relationship: where, is the dynamic non-linear inertia weight, is the minimum value of the dynamic non - linear inertia weight, is the maximum value of the dynamic non - linear inertia weight, is the current iteration number, is the maximum iteration number; The updated position of the pelican individual is obtained by updating the new position of the pelican individual according to the vertical crossover operator, satisfying the following relationship: where, is the position of the th pelican individual in the -dimensional space, is the control parameter, is the position of the optimal individual in the current pelican population in the -dimensional space, is the position of any random individual in the current pelican population in the -dimensional space; The optimal position of the pelican individual is obtained by iteratively optimizing the updated position of the pelican individual using the Pareto distribution to construct the improved pelican optimization algorithm. The present invention significantly improves the global search ability and convergence efficiency of the pelican algorithm through multi - mechanism collaborative optimization; The chaotic mapping initializes the population, breaking the uniformity limitation of the traditional random distribution, generating a more diverse initial solution set through high - dimensional non - linear mapping, and avoiding premature convergence; The dynamic non - linear inertia weight uses a quadratic decay function to dynamically balance global exploration and local development, giving a larger weight in the initial stage of iteration to enhance the cross - region search ability and narrowing the weight in the later stage to focus on fine - tuning parameters; The vertical crossover operator linearly interpolates the global optimal solution and the random solution, guiding individuals to approach the Pareto front while maintaining population diversity; The Pareto distribution optimization mechanism screens multi - objective optimal solutions through non - dominated sorting, ensuring that the parameter combination achieves balance in indicators such as recognition rate and calculation delay; It solves the problems that traditional optimization algorithms are prone to fall into local optima and have a slow convergence speed in the human activity recognition model, providing an efficient solution for the model adaptive optimization in the dynamic scenario of wearable devices.

[0014] Optionally, the spatio-temporal dependence optimized feature is obtained by optimizing the global spatio-temporal dependence feature through the human activity recognition optimization model to improve the recognition of the human activity, including: based on a fully connected layer, combining the spatio-temporal dependence optimized feature to obtain the global feature of the human activity; using an output layer to map the global feature into a class probability, and classifying the human activity to obtain a human activity classification result. The present invention significantly improves the accuracy and generalization performance of human activity recognition through a hierarchical structure of feature abstraction and probabilistic classification; the fully connected layer performs non-linear mapping and global integration on the spatio-temporal dependence optimized feature, mines the deep association between different sensor channels through an activation function, and eliminates redundant noise to form a highly discriminative global feature vector; the output layer converts the global feature into a class probability distribution, suppresses overfitting, and ensures the robustness of the model to individual differences and dynamic environments; while retaining local action details, it strengthens the global temporal logic and can still maintain a high classification accuracy in complex scenarios.

[0015] In a second aspect, the present invention provides an improved system for human activity recognition based on wearable sensors. The system executes the improved method for human activity recognition based on wearable sensors provided by the present invention. The system includes an input device, an output device, a processor, and a memory. The hardware facilities integrated in the present invention have excellent performance, and the input device, output device, processor, and memory are interconnected with each other. The present invention constructs an efficient information processing system through a high-performance hardware cooperative architecture, improving the real-time performance and reliability of human activity recognition; the input device accurately collects multi-modal sensor signals to ensure the integrity of the original data; the processor is equipped with a parallel computing unit to efficiently execute complex algorithms such as spatio-temporal feature extraction and multi-head attention mechanism; the memory uses a cache mechanism to optimize the access efficiency of feature tensors and model parameters, supporting large-scale temporal data processing; the output device provides real-time feedback on the classification results. Description of the Drawings

[0016] Figure 1 It is a flowchart of an improved method for human activity recognition based on wearable sensors according to an embodiment of the present invention; Figure 2 It is a convolutional neural network structure diagram according to an embodiment of the present invention; Figure 3 It is a structure diagram of a bidirectional gated recurrent neural network unit according to an embodiment of the present invention; Figure 4 It is an improved bidirectional gated recurrent neural network structure diagram according to an embodiment of the present invention; Figure 5 It is an improved multi-head attention mechanism structure diagram according to an embodiment of the present invention; Figure 6 It is a structural flowchart of an improved pelican optimization algorithm according to an embodiment of the present invention; Figure 7 Structural schematic diagram of the human activity recognition optimization model according to an embodiment of the present invention; Figure 8 Accuracy comparison curve graph of the human activity recognition optimization model according to an embodiment of the present invention; Figure 9 Loss comparison curve graph of the human activity recognition optimization model according to an embodiment of the present invention; Figure 10 Confusion matrix graph of the human activity recognition optimization model according to an embodiment of the present invention; Figure 11 Comparison graph of ablation experiment results according to an embodiment of the present invention; Figure 12 Framework diagram of an improved system for human activity recognition based on wearable sensors according to an embodiment of the present invention. Detailed implementation manners

[0017] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described here are only for illustrative purposes and are not used to limit the present invention. In the following description, in order to provide a thorough understanding of the present invention, a large number of specific details are set forth. However, it is obvious to those of ordinary skill in the art that the present invention does not have to employ these specific details. In other instances, well-known circuits, software, or methods have not been described in detail in order to avoid obscuring the present invention.

[0018] Throughout the specification, the reference to "one embodiment", "an embodiment", "one example", or "an example" means that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment", "in an embodiment", "one example", or "an example" appearing throughout the specification do not necessarily all refer to the same embodiment or example. In addition, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0019] Please refer to Figure 1 , an embodiment of the present invention provides an improved method for human activity recognition based on wearable sensors, and the method includes the following steps: S1. Collect the time series raw data of human activities by using wearable sensors, and obtain a time series feature tensor according to the time series raw data to construct the local spatio-temporal features of the human activities.

[0020] In this embodiment, first, the original time-series data of human activities are collected using wearable sensors (plantar pressure sensors, accelerometer sensors, and gyroscopes); subsequently, since the data collection may be affected by factors such as environmental interference and sensor errors, the data is preprocessed, and Gaussian filtering, sliding window partitioning, and normalization methods are used to reduce noise, effectively removing noise and retaining important signal features. Normalization ensures that all sensor data is processed on the same scale, improving the efficiency and stability of model training; the combination of the three methods not only significantly reduces the impact of noise on the data but also enhances the usability of the data, further improving the model performance; finally, the original data after preprocessing is feature-integrated and converted into a time-series feature tensor with a shape of (number of samples, time step, feature dimension).

[0021] Among them, the Gaussian filtering parameter is set to , the time window size is set to 50, and the step size is 10, that is, each window contains 50 data points, and the window moves forward with a step size of 10 data points.

[0022] Furthermore, a convolutional neural network (CNN) is used to automatically extract local spatio-temporal features from the time-series feature tensor.

[0023] Please refer to Figure 2 , which shows the structure diagram of the convolutional neural network; as the core component for extracting the behavior features of sensor data, the basic network structure of the convolutional neural network consists of an input layer, a hidden layer, and an output layer; the hidden layer usually includes modules such as a convolutional layer and a pooling layer; through the combination of these structures, the convolutional neural network can effectively transform the original sensor signals into features with stronger representation capabilities, thereby improving the accuracy and efficiency of human behavior recognition.

[0024] In this embodiment, a one-dimensional convolutional neural network (1D-CNN) is used to process one-dimensional time-series data with high dimensions and long lengths, significantly reducing the computational complexity and memory consumption, avoiding the problems of gradient disappearance or explosion commonly found in long time-series data, and making the model more stable during the training process; to further improve the performance of the combined model and reduce the risk of overfitting, batch normalization and Dropout layers are added after the convolutional layer; the model is composed of the combination of two convolutional layers. Before optimization, each convolutional layer is set with 64 convolutional kernels, and the size of the convolutional kernel is 3; after each convolutional layer, a batch normalization layer, a MaxPooling1D layer, and a Dropout layer at the end are added in sequence; the above additional layers enhance the generalization ability of the model through regularization means.

[0025] S2. According to the local spatio-temporal features, a first improved model for human activity recognition is constructed by combining residual connections to obtain the weighted fusion features of the human activities.

[0026] Please refer toFigure 3 , shown as the structure diagram of a bidirectional gated recurrent neural network unit; the gated recurrent neural network unit (GRU) is an improved recurrent neural network (RNN), designed specifically to handle long-range dependencies in sequential data; by introducing update gates and reset gates to alleviate the vanishing gradient problem, thus learning long-term dependencies, the bidirectional gated recurrent neural network unit (BiGRU) has a bidirectional improvement in structure, enabling it to make more full use of the context information before and after, better utilize the context information, and thus more comprehensively understand the temporal relationship in the data.

[0027] In the prior art, the CNN-BiGRU model achieved efficient parsing of human activity data by combining automatic feature extraction and temporal relationship modeling; the spatial feature extraction of CNN and the temporal dependency modeling of BiGRU worked together, enabling the model to not only focus on the local details of activities but also capture global time patterns.

[0028] In addition, the combination of CNN and BiGRU also demonstrated unique advantages in balancing the expressive power and computational efficiency of the model; CNN was responsible for feature extraction while avoiding the complexity of manually designing features, and BiGRU effectively captured the dynamic changes in temporal data through the comprehensive learning of forward and backward information.

[0029] However, there are also some limitations in the application of CNN-BiGRU: Although deep convolutional neural networks can extract multi-layer features, gradient vanishing or explosion may still occur in complex networks, resulting in unstable training or difficulty in model convergence; BiGRU mainly focuses on local temporal dependencies and fails to make full use of long-term relationships before and after, so the efficiency of the model in capturing global information when dealing with long sequence data and activities with frequent dynamic changes still has limitations, and there is a slight deficiency in accuracy.

[0030] In this embodiment, the local spatio-temporal features are parsed according to the time distribution layer (forward GRU, backward GRU) to obtain forward temporal information and backward temporal information; the forward temporal information and backward temporal information are weighted and concatenated based on the weighted concatenation layer to obtain the concatenated output features; on the basis of the CNN-BiGRU model, a residual connection is added to construct the first improved model for human activity recognition, enhancing its ability to effectively model long-term dependencies, and according to the first improved model for human activity recognition, the local spatio-temporal features and the concatenated output features are fused to obtain the weighted fusion features.

[0031] Please refer to Figure 4 , shown as the structure diagram of the improved bidirectional gated recurrent neural network; in order to enhance the model's ability to model long-term dependencies and promote information flow, the present invention introduces a residual connection between CNN and BiGRU.

[0032] Specifically, the residual connection adds the input of the BiGRU to its output to form a skip connection, effectively alleviating the vanishing gradient problem and promoting information flow. In addition, the residual connection helps the BiGRU better capture long-term dependencies, accelerate the convergence process, and enhance the learning ability for complex time-series data patterns.

[0033] In the human activity recognition task, the contributions of forward temporal information and backward temporal information may not be the same. Forward temporal information usually reflects the evolution of the current state, while backward temporal information provides a review of the previous state. Traditional BiGRU simply concatenates the forward and backward outputs directly to form the final output.

[0034] To make full use of forward temporal information and backward temporal information, a weighted concatenation mechanism is introduced. By assigning different weights to forward temporal information and backward temporal information, the model can handle these two types of information more flexibly. Forward temporal information can be strengthened by weights, and backward temporal information can be adjusted by weights. The forward weight is used as a key hyperparameter, and it is optimized through an optimization algorithm to find the best weight ratio suitable for the model. The given initial forward weight is 0.5.

[0035] S3. Based on the weighted fusion features, a second improved human activity recognition model is constructed by combining an improved multi-head attention mechanism to obtain the global spatio-temporal dependence features of the original time-series data.

[0036] A multi-head attention mechanism (MHA) incorporating fused time encoding is introduced to strengthen the model's understanding of global features and its ability to capture important information.

[0037] Please refer to Figure 5 , which shows the structure diagram of the improved multi-head attention mechanism. The improvement mainly generates different time encodings based on sine periodic functions and cosine periodic functions, and fuses these encodings with the input features. Then, the encodings are input into the multi-head attention mechanism, enhancing the model's understanding of dynamic changes, enabling the model to more flexibly select important features, capture deeper correlations, and perform more powerfully in complex and dynamic tasks. The encoding range is usually between ; The number of "heads" is set to 4, and the initial attention key value is 64.

[0038] In this embodiment, the weighted fusion features are used as the input features of the improved multi-head attention mechanism. The attention scores of the input features are obtained through linear transformation, and the multi-head attention distribution is obtained by combining the attention scores according to the weight matrix. Different time encodings are obtained based on the input features according to sine periodic functions and cosine periodic functions, and the fused multi-scale time encoding is obtained based on the time encodings. Based on the second improved human activity recognition model, the multi-head attention distribution and the fused multi-scale time encoding are weighted and summed to obtain the output features.

[0039] Specifically, the multi-head attention mechanism is a mechanism that enhances the model's attention to different parts of the data and is commonly used in sequence data analysis. By calculating different attention distributions through multiple "heads", the model can learn various features in different dimensions and positions.

[0040] The attention score satisfies the following relationship: where is the attention score, is the query, is the key, is the value matrix, is the normalization function, is the dimension of the key.

[0041] The calculation of each "head" satisfies the following relationship: where is the head of the multi-head attention mechanism, is the attention function, is the query, is the key, is the value matrix, is the weight matrix of the head index.

[0042] The attention distribution satisfies the following relationship: where is the multi-head attention distribution, is the query, is the key, is the value matrix, is the concatenation function, is the head of the multi-head attention mechanism, is the number of heads, is the concatenated weight matrix.

[0043] S4. Optimize the improved human activity recognition model according to the improved pelican optimization algorithm to obtain an optimized human activity recognition model.

[0044] The Pelican Optimization Algorithm (POA) is a nature-inspired algorithm, which is inspired by the hunting behavior of reptiles and mainly divided into two stages: global search and local search; before hunting, the pelican population needs to be initialized. In the first stage, the position of the prey is randomly generated within the search space. If the fitness value of the new position is better than the previous position, the current position will be updated; in the second stage, the pelican individuals determine the flight direction and distance based on their own positions, the global optimal solution and the positions of other individuals.

[0045] The traditional Pelican Optimization Algorithm has the following deficiencies: the global search ability of the algorithm in complex optimization spaces is insufficient and it is easy to fall into local optimal solutions; the fixed inertia weight reduces the adaptability of the algorithm and affects the convergence speed; the lack of solution diversity may lead to the loss of exploration potential.

[0046] In response to the above problems, the Pelican Optimization Algorithm is improved based on chaotic mapping, dynamic non-linear inertia weight, vertical crossover operator and Pareto distribution to construct an improved Pelican Optimization Algorithm.

[0047] Please refer to Figure 6 , the structure flow chart of the improved Pelican Optimization Algorithm is shown in the figure.

[0048] First of all, the initialization stage remains unchanged; in the first stage, chaotic mapping (Chebyshev mapping) is introduced for chaotic interference to increase the randomness and solution diversity of the pelican population, thus avoiding the algorithm from falling into local optimum; the following relationship is satisfied: Among them, is the position of the th pelican individual in the th dimensional space, is the perturbation amplitude control parameter, is the sine function, is the chaotic mapping, and the mathematical model of the chaotic mapping satisfies the following relationship: Among them, is the next position of the pelican individual, is the cosine function, is the order of the chaotic mapping, is the current position of the pelican individual.

[0049] In the second stage, a dynamic non-linear inertia weight is introduced; a higher dynamic non-linear inertia weight is used in the initial stage to strengthen the global search ability, and then the dynamic non-linear inertia weight is gradually reduced to optimize the local search; the new position of the pelican individual of the pelican population is obtained based on the dynamic non-linear inertia weight, and the dynamic non-linear inertia weight satisfies the following relationship: Among them, is the dynamic non-linear inertia weight, is the minimum value of the dynamic non-linear inertia weight, is the maximum value of the dynamic non-linear inertia weight, is the current iteration number, is the maximum iteration number.

[0050] Among them, , .

[0051] Next, the cross operator is used to cross the characteristics of different particles to generate new solutions, further enhancing the exploration ability of the algorithm; according to the vertical cross operator, the new positions of the pelican individuals are updated to obtain the updated positions of the pelican individuals, satisfying the following relationship: Among them, is the position of the th pelican individual in the -dimensional space, is the control parameter, is the position of the optimal individual in the current pelican population in the -dimensional space, is the position of any random individual in the current pelican population in the -dimensional space.

[0052] Among them, .

[0053] Finally, the Pareto distribution (Pareto distribution) is used to iteratively optimize the updated positions of the pelican individuals to obtain the optimal positions of the pelican individuals. The Pareto distribution is used to optimize the solutions to ensure that the optimal solutions are retained in each iteration, accelerating the optimization process.

[0054] In this embodiment, the key hyperparameters of the second human activity recognition improvement model are optimized by improving the pelican optimization algorithm to obtain the human activity recognition optimization model (POA-CNN-ResBiGRU-MHA model). The key hyperparameters are optimized by the improved pelican optimization algorithm (POA), including the learning rate, the number of convolutional kernels, the number of GRU units, the output weight of the forward GRU, and the key dimension of the attention mechanism, so as to improve the stability and performance of the model, improve the accuracy and robustness of the model under different activity categories. The improved optimization algorithm can find the best parameters more efficiently and can effectively avoid falling into local optima, with better effects.

[0055] S5. Optimize the global spatio-temporal dependence features through the human activity recognition optimization model to obtain spatio-temporal dependence optimization features, so as to realize the improvement of the recognition of the human activity.

[0056] In this embodiment, based on the fully connected layer, the global features of the human activity are obtained by combining spatio-temporal dependence to optimize the features; the output layer is used to map the global features into class probabilities, and the human activity is classified to obtain the human activity classification result, including walking, going up stairs, going down stairs, going uphill, and going downhill.

[0057] Specifically, please refer to Figure 7 , which is a schematic structural diagram of the human activity recognition optimization model; it includes five main modules: a data preprocessing module, a feature extraction module (1DCNN), an improved temporal modeling module (ResBiGRU), an improved attention module (MHA), and an improved model optimization algorithm module (POA).

[0058] The sensor data first undergoes feature integration through the preprocessing module and is converted into a temporal feature tensor with a shape of (number of samples, time steps, feature dimension). The convolutional neural network then automatically extracts local spatio-temporal features, and subsequently the time distribution layer operates on the local spatio-temporal features in the time dimension. The BiGRU module captures the temporal information in the forward and backward directions, and performs weighted fusion on the forward and backward outputs in the weighted splicing layer. Through residual connection, the input features are added to the spliced output features to reduce information loss.

[0059] The multi-head attention mechanism that fuses two types of time encoding then calculates the relative importance between the input features, pays attention to different parts of the data in parallel, and comprehensively learns the temporal and spatial dependencies in the data. Then, the fully connected layer activated by ReLU further processes and maps the extracted features. Finally, through the output layer activated by Softmax, the features are mapped into class probabilities to achieve classification. Finally, an improved pelican optimization algorithm (POA) is introduced to optimize the key hyperparameters to ensure that the model reaches the best performance, thereby improving the recognition accuracy and robustness of different categories of human activities.

[0060] S6. Experimental verification and analysis.

[0061] Among them, S6 specifically includes the following steps: S61. Making experimental datasets and public datasets.

[0062] The data acquisition device consists of a self-made intelligent insole, a self-made intelligent bracelet, and a smartphone, integrating various sensors such as flexible thin-film pressure sensors, accelerometers, and gyroscopes, realizing the comprehensive acquisition of human activity data.

[0063] In the experimental preparation stage, in order to evaluate the performance of the improved algorithm, each experimental personnel wears the corresponding data acquisition device. The thin-film pressure sensor is attached to the foot, the inertial sensor is placed at the ankle, the intelligent bracelet is placed on the wrist, and the smartphone is placed in the outer pocket of the upper garment.

[0064] Twelve volunteers, six males and six females, were selected for the experiment. The volunteers needed to perform five actions, namely walking, going up stairs, going down stairs, going uphill, and going downhill, in the selected experimental environment in sequence until the entire journey was completed. The rate of all sensors was set to 20 Hz, and the entire process was completed in an average of 7.3 minutes. A total of 98,696 sets of data were collected from all the experiments, constituting the experimental dataset.

[0065] The publicly available dataset used is a multi-modal gait database created and provided by Honda Research Institute Europe. It mainly records daily walking scenarios in natural urban environments, including the synchronized collection of IMU (Inertial Measurement Unit), FSR (Force Sensitive Resistor), and gaze data. The participants included 20 men and women of different age groups, and the data was collected from three different site routes. The data of all sensors was downsampled to 60 Hz, and a total of 846,715 sets of data were collected.

[0066] S62. Data preprocessing.

[0067] To better classify the time series, the sensor data collected in the experiment needs to be specifically processed to meet the model requirements. First, the required data is exported from the host computer and then preprocessed. Since the data collection may be affected by factors such as environmental interference and sensor errors, Gaussian filtering, sliding window partitioning, and standardization methods are used to reduce noise, and the parameter settings are the same as those in the preprocessing in step S1.

[0068] S63. Training process and result evaluation.

[0069] In terms of the experimental environment, all classification methods were implemented using the TensorFlow framework and related libraries in the Python 3.9 environment. During the model training process, the batch size was set to 64, the learning rate was , and the Adam optimizer was selected. The experimental dataset was divided into a training set, a validation set, and a test set in a ratio of 7:1:2. All experiments were conducted on the experimental dataset, and the final results were all the performances on the test set.

[0070] The confusion matrix is a basic tool for evaluating the classification performance of a model. However, when dealing with a large amount of data, it may be difficult to accurately evaluate the overall performance of the model relying solely on the confusion matrix. Therefore, accuracy, error value, precision, recall, and F1 metric are introduced to extend the analysis of the confusion matrix, and the mathematical expressions of the evaluation metrics are as follows: Among them, is the accuracy, is a true positive example, is a true negative example, is a false positive example, is a false negative example, is precision, is recall, is the F1 measure, During the training process, the model is trained and validated under preset parameters until the test set results are obtained after reaching a sufficient number of iterations. By introducing an early stopping function, it is continuously tried to determine the optimal number of iterations. After more than twenty rounds, the results fluctuate within an acceptable range, so the set number of training rounds is 30.

[0071] Please refer to Figure 8 , which shows the accuracy comparison curve of the optimized model for human activity recognition; including the accuracy curves on the training set and the validation set, describing the changing trend of the accuracy with the number of training rounds.

[0072] Please refer to Figure 9 , which shows the loss comparison curve of the optimized model for human activity recognition; including the loss curves on the training set and the validation set, describing the changing trend of the loss with the number of training rounds.

[0073] From Figure 8 and Figure 9 it can be seen that the fitting of the optimized model for human activity recognition is very good, and the model can achieve an accuracy of 98.52% on the training set and 97.32% on the validation set.

[0074] Please refer to Figure 10 , which shows the confusion matrix diagram of the optimized model for human activity recognition; the confusion matrix of the model on the experimental dataset is plotted, further verifying the effectiveness and stability of the proposed model in the multi-person activity recognition task in complex activities and dynamic scenarios.

[0075] Furthermore, the excellent performance and effectiveness of the optimized model for human activity recognition in the human behavior recognition task are verified. Through experimental comparison with the latest existing models in the field of human activity recognition on the multi-modal gait public dataset, mainly including CNN-ResBiGRU-MHA, Transform-BiLSTM, CNN-BiLSTM, CNN-BiGRU, and DWCNN. The comparison results are shown in Table 1: Table 1: Performance comparison table of each algorithm on the multi-modal gait public dataset As can be seen from Table 1, in terms of the overall accuracy rate, the optimized model for human activity recognition of the present invention outperforms other comparative models in all evaluation metrics, and the average accuracy rate on the selected public dataset can reach 98.94%, further verifying the superiority and practical application potential of the optimized model for human activity recognition in the human activity recognition task.

[0076] S64. Ablation experiment.

[0077] Please refer to Figure 11 , which is a comparative graph of the ablation experiment results; in order to verify the effectiveness and contribution of each component of the optimized model for human activity recognition, by removing or replacing different modules or features of the model one by one, the changes in the model performance are observed. All models are tested on the public dataset, the training batch is set to 30, and the loop test is conducted 20 times, and the accuracy rate on the test set each time is recorded.

[0078] The experimental results show that the accuracy rates of the basic CNN and BiGRU are relatively low, reflecting their insufficient capabilities in long-term modeling and local feature capture. However, after introducing the multi-head attention mechanism, the model performance is significantly improved, verifying the effectiveness of the attention mechanism in extracting key time step features of time series data and enhancing the ability to capture global features and important information.

[0079] Compared with CNN-ResBiGRU-MHA, the accuracy rate of the optimized model for human activity recognition is further improved and more stable, which further indicates that the hyperparameters obtained by the optimization algorithm POA can effectively improve the model performance, thereby enhancing the accuracy and robustness under different activity categories, and also proving that each module and improvement are very necessary and reasonable.

[0080] Please refer to Figure 12 , in an optional embodiment, the present invention provides an improved system for human activity recognition based on wearable sensors. The system includes an input device, an output device, a processor, and a memory, and the hardware facilities are interconnected with each other. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the specific steps of the related embodiments of the improved method for human activity recognition based on wearable sensors provided by the present invention. The improved system for human activity recognition based on wearable sensors provided by the present invention has a complete, objective and stable structure, improving the overall applicability and practical application ability of the present invention.

[0081] In summary, an improved method and system for human activity recognition based on wearable sensors provided by the method of the present invention propose an optimized model for human activity recognition based on wearable sensors, and combine a residual network to address the challenges of high-dimensional and long-term wearable device data in human activity recognition. The model introduces a multi-head attention mechanism that generates time encoding based on sine and cosine periodic functions, enhancing the extraction of key time-step features. At the same time, in view of the problems of the traditional pelican optimization algorithm, such as insufficient global search ability, poor adaptability, and slow convergence speed, an improvement scheme is proposed, aiming to refine the adjustment of key hyperparameters and enhance the collaborative effect between modules. Verification is carried out on the experimental data set created in the real environment and the public data set. The results show that the model algorithm has the best performance and good robustness. In the future, the research will explore the application potential of optimization algorithms in edge computing and few-shot learning techniques, which will enhance the practical applicability of the model and promote the development of the field technology towards diversification and universality.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. An improved method for human activity recognition based on wearable sensors, characterized in that, Including the following steps: Collecting the original time series data of human activities by using wearable sensors, and obtaining a time series feature tensor according to the original time series data to construct the local spatio-temporal features of the human activities; Based on the local spatio-temporal features, constructing a first improved human activity recognition model by combining residual connections to obtain the weighted fusion features of the human activities; Based on the weighted fusion features, constructing a second improved human activity recognition model by combining an improved multi-head attention mechanism to obtain the global spatio-temporal dependence features of the original time series data; Optimizing the second improved human activity recognition model according to the improved pelican optimization algorithm to obtain an optimized human activity recognition model; Optimizing the global spatio-temporal dependence features through the optimized human activity recognition model to obtain spatio-temporal dependence optimized features, so as to realize the improvement of the recognition of the human activities.

2. The improved method for human activity recognition based on wearable sensors according to Claim 1, wherein The step of collecting the original time series data of human activities by using wearable sensors and obtaining a time series feature tensor according to the original time series data to construct the local spatio-temporal features of the human activities includes: Performing data preprocessing on the original time series data to obtain the time series feature tensor, and the data preprocessing includes Gaussian filtering, sliding window division and standardization; Based on a convolutional neural network, performing feature extraction on the time series feature tensor to construct the local spatio-temporal features.

3. The improved method for human activity recognition based on wearable sensors according to claim 1, wherein The step of constructing a first improved human activity recognition model by combining residual connections based on the local spatio-temporal features to obtain the weighted fusion features of the human activities includes: Analyzing the local spatio-temporal features according to a time distribution layer to obtain forward time series information and backward time series information; Based on a weighted splicing layer, performing weighted splicing on the forward time series information and the backward time series information to obtain a spliced output feature; According to the first improved human activity recognition model, fusing the local spatio-temporal features and the spliced output feature to obtain the weighted fusion features.

4. The improved method for human activity recognition based on wearable sensors according to claim 1, wherein, The step of constructing a second improved human activity recognition model by combining an improved multi-head attention mechanism based on the weighted fusion features to obtain the global spatio-temporal dependence features of the original time series data includes: Taking the weighted fusion features as the input features of the improved multi-head attention mechanism; Using a linear transformation to obtain the attention scores of the input features, and obtaining a multi-head attention distribution according to a weight matrix in combination with the attention scores; According to the input features, obtaining different time encodings based on a sine periodic function and a cosine periodic function, and obtaining a fused multi-scale time encoding according to the time encodings; Based on the second improved human activity recognition model, performing weighted summation on the multi-head attention distribution and the fused multi-scale time encoding to obtain an output feature, and taking the output feature as the global spatio-temporal dependence features.

5. The improved method for human activity recognition based on wearable sensors according to claim 4, wherein The step of using a linear transformation to obtain the attention scores of the input features includes: wherein, is the attention score, is the query, is the key, is the value matrix, is the normalization function, is the dimension of the key.

6. The improved method for human activity recognition based on wearable sensors according to claim 4, wherein The step of obtaining a multi-head attention distribution according to a weight matrix in combination with the attention scores includes: Among them, is the head of the multi-head attention mechanism, is the attention function, is the query, is the key, is the value matrix, is the weight matrix of the -th head, is the index of the head, is the multi-head attention distribution, is the concatenation function, is the number of heads, is the concatenated weight matrix.

7. The improved method for human activity recognition based on wearable sensors according to claim 1, characterized in that, The step of optimizing the second improved human activity recognition model according to the improved pelican optimization algorithm to obtain an optimized human activity recognition model includes: Improve the pelican optimization algorithm based on chaotic mapping, dynamic non-linear inertia weight, vertical crossover operator and Pareto distribution to construct the improved pelican optimization algorithm; Optimize the key hyperparameters of the second human activity recognition improved model through the improved pelican optimization algorithm to obtain the human activity recognition optimized model.

8. The improved method for human activity recognition based on wearable sensors according to claim 7, characterized in that, The improvement of the pelican optimization algorithm based on chaotic mapping, dynamic non-linear inertia weight, vertical crossover operator and Pareto distribution to construct the improved pelican optimization algorithm includes: Perform chaotic interference on the pelican population based on chaotic mapping, satisfying the following relationship: in, For the The pelican individuals The position in the dimensional space, is the disturbance amplitude control parameter, is a sine function, is a chaotic mapping, and the mathematical model of the chaotic mapping satisfies the following relationship: Among them, is the next position of the pelican individual, is the cosine function, is the order of the chaotic map, is the current position of the pelican individual; Obtain the new position of the pelican individual in the pelican population according to the dynamic non-linear inertia weight, and the dynamic non-linear inertia weight satisfies the following relationship: Among them, is the dynamic non-linear inertia weight, is the minimum value of the dynamic non-linear inertia weight, is the maximum value of the dynamic non-linear inertia weight, is the current iteration number, is the maximum iteration number; Update the new position of the pelican individual according to the vertical crossover operator to obtain the updated position of the pelican individual, satisfying the following relationship: Among them, is the position of the th pelican individual in the -dimensional space, is the control parameter, is the position of the optimal individual in the current pelican population in the -dimensional space, is the position of any random individual in the current pelican population in the -dimensional space; Use the Pareto distribution to iteratively optimize the updated position of the pelican individual to obtain the optimal position of the pelican individual, so as to construct the improved pelican optimization algorithm.

9. The improved method for human activity recognition based on wearable sensors according to claim 1, characterized in that, The optimization of the global spatio-temporal dependence feature through the human activity recognition optimized model to realize the improvement of the recognition of the human activity includes: Based on the fully connected layer, combine the spatio-temporal dependence optimized feature to obtain the global feature of the human activity; Use the output layer to map the global feature to a class probability, and classify the human activity to obtain the human activity classification result.

10. An improved system for human activity recognition based on wearable sensors, characterized in that, The system includes an input device, an output device, a processor and a memory. The input device, the output device, the processor and the memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the improved method for human activity recognition based on wearable sensors according to any one of claims 1-9.

Citation Information

Patent Citations

  • Unmanned aerial vehicle path optimization method based on chaotic mapping pelican optimization algorithm

    CN116225066A

  • Group activity identification method based on coarse granularity-fine granularity nested learning

    CN116630892A

  • Power line bird identification method based on lightweight target detection model

    CN117237871A

  • Robot path planning method based on improved pelican optimization algorithm

    CN117471919A

  • Human body activity identification method based on residual shrinkage network

    CN117523672A