Human body behavior recognition system and method based on sparse three-channel mixed attention

Through the multi-scale dynamic convolution method of sparse three-channel hybrid attention, the problem of poor interpretability of multi-source sensor position contribution and attention mechanism in the prior art is solved, and the behavior recognition effect of high accuracy and low parameter quantity is achieved, and the model is overfitted is prevented.

CN120180244AActive Publication Date: 2025-06-20SHANDONG UNIV +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510668056.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing behavioral recognition algorithms based on deep learning cannot effectively utilize the location contribution of multi-source sensors at the feature extraction layer, and the attention mechanism has poor interpretability between single channels and channels, resulting in insufficient recognition accuracy and model overfitting.

Method used

The multi-scale dynamic convolution method of sparse three-channel hybrid attention is adopted to enhance feature extraction and model protection by independently parallel learning of spatiotemporal features in the single-channel and inter-channel relationships at different locations of multi-source sensors, combined with improved extrusion-excited three-channel dynamic convolution, three-channel hybrid attention and sparse attention, to enhance the ability of feature extraction and model to prevent overfitting.

Benefits of technology

It improves the accuracy of human behavior recognition, reduces the number of model parameters, enhances feature representation ability, prevents model overfitting, and improves the universality of complex behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180244A_ABST
    Figure CN120180244A_ABST
Patent Text Reader

Abstract

The invention discloses a human body behavior recognition system and method based on sparse three-channel mixed attention, and relates to the technical field of human body behavior recognition of deep learning. Comprising a human body behavior information acquisition module, a human body behavior information transmission module, a human body behavior information preprocessing module, a human body behavior classification and recognition module and a human body behavior information application module which are connected in sequence, the human body behavior information acquisition module is used for acquiring behavior data of a user in real time. The technical problem to be solved by the invention is to provide a human behavior recognition system and method based on sparse three-channel mixed attention, which can independently learn spatio-temporal characteristics in parallel from single channels at different positions of a multi-source sensor and a relationship between the channels. The defects of the existing behavior recognition system feature extraction method and the defect that the contribution rate of the sensor position cannot be fully utilized are overcome, the recognition accuracy is effectively improved, and meanwhile the parameter quantity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human behavior recognition in deep learning, and more specifically, to a human behavior recognition system and method based on sparse three-channel hybrid attention. Background Art

[0002] With the development of wearable devices and deep learning technologies, Human Activity Recognition (HAR) has received extensive attention. The HAR technology refers to the behavior monitoring in a specific field through the data acquired by sensors. In recent years, the HAR technology has been widely applied in many fields such as motion tracking, virtual reality, medical health, work safety, smart home, and abnormal behavior recognition, which can significantly improve production efficiency and quality of life. The general process of a HAR system includes data acquisition, data preprocessing, feature extraction, and behavior classification.

[0003] The algorithm design in the feature extraction stage is the core of the human behavior recognition system and the key to achieving high accuracy and low memory. Feature extraction aims to understand and judge different behaviors by studying the detailed features of various behaviors in the acquired data. Wearable-based HAR uses sensors in portable intelligent devices to collect rich motion information and then accurately recognize human activities. Using a deep learning model to extract features reflecting the differences between different activities from massive motion information is the key to achieving accurate recognition in HAR. The Convolutional Neural Network (CNN) is one of the representative models of deep neural networks and can automatically extract local features in data. The Recurrent Neural Network (RNN) is also one of the widely used deep models at present and can effectively extract global features in sequence data and capture long-term relationships in sequence data. In addition, as a sequence modeling method, the attention mechanism can capture long-range dependencies by learning important time points in sequence data, and more and more attention mechanisms are introduced into neural networks. The fusion of attention mechanisms is of great significance for processing complex sequences and enhancing feature extraction capabilities.

[0004] At present, there are many challenges in the research field of behavior recognition algorithms based on deep learning. For example, most studies usually use RNN and other variants to learn the long-term correlation of latent features, but the utilization of the long-term correlation in the original time series data is insufficient. Secondly, the deep learning models for multi-location sensor data are not very advanced and cannot learn the features from multi-source sensors at different locations in the feature extraction layer, ignoring the location contribution of the sensors. Therefore, analyzing the differences and impacts of multi-location sensor data features in different types of human activities and obtaining the independent contribution of each location to activity recognition are necessary steps. In addition, many networks with attention mechanisms attempt to learn spatio-temporal attention at the image level, ignoring the interpretability of single-channel attention and inter-channel attention in time series analysis. Moreover, increasing the model depth helps to improve the accuracy of the training dataset but reduces the generalization ability. Focusing on learning low-contribution features plays an important role in preventing model overfitting. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a human behavior recognition system and method based on sparse three-channel hybrid attention, which can independently and parallelly learn spatio-temporal features from the single-channel and inter-channel relationships of different locations of multi-source sensors, solve the deficiencies in the feature extraction method of the existing behavior recognition system and the defect of being unable to fully utilize the sensor location contribution rate, effectively improve the recognition accuracy while reducing the number of parameters; the present invention also proposes three attention improvement methods: three-channel dynamic convolution based on improved squeeze-and-excitation (TCDC-ISE), three-channel mixed attention (TCMA), and sparse attention (SA).

[0006] For a human behavior recognition system with sparse three-channel hybrid attention involved in the present invention, first, the user wears a wearable device for data collection, which is built with various sensors. Different types can be selected according to different requirements and worn on different parts of the human body. Then, the collected data is sent to a local server or a cloud server for storage through a data transmission module in a wireless transmission manner, and the transmission method can be selected according to different usage scenarios. Next, preprocessing operations are performed on the data, including filtering, normalization / standardization, sliding window segmentation of data, etc. Filtering is to remove the noise in the original sensor signals and the gravitational acceleration signals in the acceleration data; normalization / standardization can transform data with different dimensions and value ranges into the same range, which is beneficial for distance-based calculation models to eliminate the influence caused by different dimensions and value ranges; the sliding window segmentation method divides continuous long-time data into data segments that conform to the input format of the discrimination model. Furthermore, the preprocessed data is input into the human behavior classification and recognition module. Among them, it is first processed by a multi-scale dynamic convolution with sparse three-dimensional mixed attention (MDC-STDMA) module to independently and parallelly learn spatio-temporal features from the single-channel and inter-channel relationships at different positions of multi-source sensors. This operation preprocesses the input of MDC-STCMA through a scanning expansion module and a multi-scale feature segmentation module, focuses on learning time data with long-distance dependence relationships under the same receptive field and groups them according to the sensor positions; pays attention to the single-channel features themselves through the TCDC-ISE module; pays attention to the feature correlation between channels through the TCMA module; and reduces the overfitting tendency of the model through the SA module. Then it is further processed by a global attention based temporal and spatial fusion (GA-TSF) module to achieve the last step of spatio-temporal global information fusion by using the global attention mechanism that focuses on spatio-temporal features. Finally, by calculating the loss function of each category, the model parameters are adjusted during each round of training.

[0007] The present invention is a general human behavior recognition system based on wearable sensors, which can monitor the behavior categories of users in real time and can be used for the monitoring and management of specific work fields, as well as for the guardianship of the elderly and the rehabilitation management of patients. Aiming at the problems that the existing algorithms mix and process the features of each channel, resulting in the inability to separate the features of sensors at different positions, and the interpretability of the attention extraction of each channel and the attention mixing between channels is poor, the present invention provides a multi-scale dynamic convolution method with sparse three-channel hybrid attention to solve this problem, further enhancing the feature representation and improving the accuracy of behavior recognition. In addition, aiming at the problem that the existing methods are prone to overfitting due to the improvement of accuracy, the sparse attention method is used to retrain the nodes with low contribution rate, effectively preventing the model from overfitting, enhancing the universality of the model for complex behaviors, and being more conducive to the application in actual scenarios.

[0008] The present invention adopts the following technical solutions to achieve the invention purpose: A human behavior recognition system and method based on sparse three-channel hybrid attention, characterized by comprising: A human behavior information acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a human behavior classification and recognition module, and a human behavior information application module that are connected in sequence; The human behavior information acquisition module is used to collect the behavior data of the user in real time, and the behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer; The human behavior information transmission module is used to transmit the collected behavior data of the user to the local server or the cloud server; The human behavior information preprocessing module is used to sequentially store, merge, filter, normalize or standardize, perform sliding window segmentation, and label calibration on the collected behavior data of the user, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels; The human behavior classification and recognition module is used to input the behavior data of the user preprocessed by the human behavior information preprocessing module into the trained human behavior classification and recognition module for behavior type discrimination; The human behavior information application module is used to transmit the obtained human behavior classification and recognition result to various application platforms, so as to realize corresponding functions.

[0009] As a further limitation of this technical solution, the human behavior classification and recognition module includes a scanning expansion module, a multi-scale feature segmentation module, a multi-scale dynamic convolution module with sparse three-channel hybrid attention, a global attention spatio-temporal fusion module, and a human behavior classification result output module that are connected in sequence;

[0010] As a further limitation of the technical solution, the human behavior information acquisition module includes a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer for collecting the user's behavior data.

[0011] As a further limitation of the technical solution, the human behavior information preprocessing module includes a sensor data storage unit, a sensor data merging unit, a sensor data smoothing and filtering unit, a data normalization / standardization unit, a sliding window segmentation unit, and a data label calibration unit connected in sequence.

[0012] As a further limitation of the technical solution, the collected user behavior data is transmitted to a local server or a cloud server through any one of the transmission methods of ultra-wideband, Bluetooth, and 5G.

[0013] A recognition method of a human behavior recognition system based on sparse three-channel hybrid attention, characterized by comprising the following steps: S1: Human behavior data acquisition; The behavior data of the user is collected in real time, that is, the human behavior data. The behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer; S2: Human behavior data transmission; The collected user behavior data is transmitted to a local server or a cloud server; S3: Human behavior data preprocessing; The collected user behavior data is sequentially stored, merged, filtered, normalized or standardized, segmented by a sliding window, and labeled, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels; S4: Construct a recognition model and classify and recognize human behaviors; The behavior data segments preprocessed in S3 are input into the recognition model in batches, and the human behavior recognition output is obtained after training; S5: Loss function calculation, update and feedback; Calculate the loss function for the classification result generated in S4, and adjust the model parameters through backpropagation to improve the classification accuracy and generalization ability. Use the cross-entropy loss function to calculate the error between the model prediction value and the true value. The cross-entropy loss function is defined as shown in Equation (1): (1); Where: represents the batch data volume, represents the total number of categories, represents the th sample belonging to the category true label, Indicates that the th sample belongs to the category with a predicted probability; (2); Where: Indicates the th step model parameter, Indicates the th step model parameter, Indicates the learning rate, is a constant added to increase numerical stability, Indicates the first-order momentum, Indicates the second-order momentum, and Indicates the first-order and second-order average coefficients and to the power; The calculation process of the first-order momentum can be expressed as Equation (3): (3); Where: Indicates the gradient of the loss function , Indicates the first-order average coefficient, Indicates the current first-order momentum value, Indicates the first-order momentum value at the previous moment; The calculation process of the second-order momentum can be expressed as Equation (4): (4); Where: Indicates the second-order average coefficient, Indicates the current second-order momentum value, Indicates the second-order momentum value at the previous moment; S6: Application of human behavior classification output results; Transmit the human behavior classification and recognition results to the corresponding application platform in real time.

[0014] As a further limitation of this technical solution, the specific implementation process of S3 is as follows: S31: Merging of human behavior data; Stitch and merge the collected user behavior data. The merged data is in the format of a two-dimensional array, arranged horizontally in the order of the X, Y, and Z axis data of the three-axis acceleration sensor, the X, Y, and Z axis data of the three-axis angular velocity sensor, and the X, Y, and Z axis data of the three-axis magnetometer, and arranged vertically in chronological order; S32: Filtering of human behavior data; Filtering of human behavior data refers to smoothing filtering of sensor data; The sensor data smoothing filter uses a three - point mean smoothing filter to eliminate the noise signals inside the sensor. Let the sensor signal sequence be , denote the number of data points in the sensor signal sequence; The mean smoothing formulas are shown in Formulas (5) and (6): (5); (6); where: denotes the value after one - time mean smoothing, denotes the value after two - time mean smoothing; S33: Normalization / Standardization of human behavior data; The deviation normalization method is used for normalization. Let the input sequence be , and the output sequence after deviation normalization is , as shown in Formula (7): (7); where: and are respectively the maximum and minimum values of all samples in the input sequence ; The standard score standardization method is used for standardization, as shown in Formula (8): (8); where: is the mean of all samples in the input sequence , is the standard deviation of all samples in the input sequence ; S34: Sliding window segmentation and tagging; Sliding window segmentation uses a window with a fixed length to divide the continuous sensor data into data segments with a fixed length, assigns a corresponding tag to each data segment, that is, the name of the behavior category corresponding to the data segment, and encodes the tag using one - hot encoding.

[0015] As a further limitation of this technical solution, the specific implementation process of S4 is as follows: S41: Scan expansion is to perform data augmentation processing on the input time - series data using shuffled arrangement and flattening operations, expand it into four time - sorting combinations, realize time - feature shuffling, and enhance the feature learning effect. Specifically, it means: The time - series data with the size of each feature dimension being is converted through reshape to Class image data, where ; For each data, flattening operations are performed in 4 orders: horizontally from the upper left to the lower right, vertically from the upper right to the lower left, horizontally from the lower right to the upper left, and vertically from the lower right to the upper left, to obtain and , and the calculation process of their shuffled and flattened arrangement is shown in Equation (9): (9); Merge the four new time series sequences and to obtain the data augmentation sample , as shown in Equation (10): (10); Where: Concat represents the merge connection operation; S42: Multi-scale feature segmentation separates the features of sensors at different positions. In the subsequent feature extraction stage, the grouped data is input into several parallel similar block networks to facilitate the extraction of high-level feature representations of sensor data at different positions in the same activity. Specifically, it means: For the data augmentation sample after the above-mentioned S41, group the multi-source sensor features at different positions according to the body position grouping method of the wearable sensor into groups: IMU1, IMU2,..., IMUn; S43: Sparse three-channel hybrid attention multi-scale dynamic convolution performs in-depth extraction and fusion of features; S44: The global attention spatio-temporal fusion module is used to fully fuse spatio-temporal attention features: The global attention unit combines the gated linear unit and the attention mechanism. Through the mapping process of the gated linear unit, matrices and are generated, and then the self-attention mechanism is used for key feature extraction. The calculation process is shown in Equation (11): (11); Where: is the output of the global attention unit; is the weight bias; is the query matrix, and are the key and value represented by the matrix respectively, is the ReLU activation function; Further, through the one-dimensional convolution calculation with a convolution kernel size of 3 and a stride of 1, data with a feature channel of 1 is generated, and the output of the global attention is further extracted to ensure the non-loss of local information; S45: The human behavior classification result output module is used to reduce the dimension and classify the high-dimensional features obtained by feature extraction and fusion, so as to achieve accurate classification of human behaviors, including a flattening layer, a layer normalization, a fully connected layer, and a Softmax output layer connected in sequence.

[0016] As a further limitation of this technical solution, the specific implementation process of S43 is as follows: S431: The improved squeeze-and-excitation three-channel dynamic convolution unit includes an initial convolution block 0 to obtain potential features of each channel, and three convolution blocks 1, convolution block 2, and convolution block 3 connected in series in sequence; S432: The three-channel hybrid attention unit consists of three parts: long and width channel hybrid attention, long and feature channel hybrid attention, and width and feature channel hybrid attention; The Patch Merging unit is used to shuffle different sensor features, randomize the sensor sources of the features, and achieve the merging of different sensor features. The specific implementation process is as follows: In S432, the finally output attention map passes through a shuffle layer to generate a shuffled feature map with odd-even skip connections; through a splicing layer, data with a feature length 4 times the original and long and width channels 1 / 2 of the original is generated; through a linear layer, the feature dimension is halved to generate data with a feature length 2 times the original and long and width channels 1 / 2 of the original; S434: The sparse attention unit is used to prevent the model from overfitting during the training of low contribution rate nodes, and it consists of 2 mask layers, an encoding layer, and a decoding layer; Each attention vector with a length of passes through a mask layer. The mask module sets the mask rate to 0.75, sets the attention with the top 0.75 weights before sorting to 0, and the rest to 1. Its masking process is shown in formulas (12)-(14): (12); (13); (14); Where: is the reverse sorting function, is the attention sequence with a length of sorted from largest to smallest, will be used as the input of the encoding module; is the masking function; is the attention sequence after masking processing; The encoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 2 of the original; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 4 of the original. The decoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 2 of the original; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel of the original dimension.

[0017] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention proposes a sparse three-channel hybrid attention human behavior recognition system, which realizes accurate recognition of the user's motion state through the processing and analysis of motion signals of wearable sensors at different positions, solves the problems that the video-based human behavior recognition method is vulnerable to the environment, has poor privacy, and the problem that it is difficult to collect data for the radio frequency device-based human behavior recognition method.

[0018] The present invention proposes a sparse three-channel hybrid attention human behavior recognition method, which can independently and parallelly learn the spatio-temporal features of multi-source sensors at different body positions from single channels and inter-channel relationships, thereby improving the accuracy of human behavior recognition.

[0019] The present invention designs a multi-scale dynamic convolution module with sparse three-channel hybrid attention to deeply extract and fuse features. To improve the interpretability of single-channel attention and inter-channel attention in the feature extraction layer, this module uses three-channel dynamic convolution with improved squeeze excitation and three-channel hybrid attention to respond to single-channel and inter-channel features respectively, and uses sparse attention to prevent overfitting.

[0020] The present invention introduces a multi-scale feature segmentation module and a scanning expansion module, which effectively utilizes the time multi-scale correlation of the original time series data and the position contribution of multi-source sensors, and at the same time relaxes the requirement for the optimal convolution kernel size of the CNN.. Description of the Drawings

[0021] Figure 1 It is a schematic diagram of the module composition and connection relationship of the wearable human behavior recognition system of the present invention.

[0022] Figure 2 It is a schematic flow diagram of the wearable human behavior recognition method of the present invention.

[0023] Figure 3This is a schematic diagram of the principle of human behavior classification in the wearable human behavior recognition system of the present invention. Detailed implementation mode

[0024] The following combines the accompanying drawings to describe a specific implementation mode of the present invention in detail, but it should be understood that the protection scope of the present invention is not limited by the specific implementation mode.

[0025] Example 1

[0026] In the application of intelligent rehabilitation, there are many problems. For example, the monitoring of the training actions of patients during rehabilitation training is an important standard for evaluating the rehabilitation training effect. Although cameras are arranged everywhere in the rehabilitation training room, it is difficult to monitor the rehabilitation training actions of patients in daily scenarios and evaluate them in a timely manner. Moreover, the cameras have problems such as being easily blocked, insufficient light, and having dead angles. At this time, the human behavior recognition technology based on wearable devices can make up for the deficiencies of the video. The patient wears a wearable device at the corresponding body position, which is built with a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer, and can be used to collect the patient's behavior data in real time, and transmit it to the local server or cloud server for storage and calculation in real time, and obtain the corresponding recognition result. The intelligent rehabilitation application platform can display this result in real time, and the patient's attending physician and the patient's family members can use this result to observe the patient's rehabilitation situation in real time. In addition, the behavior status of the patient can also be managed in the long term to avoid abnormal situations such as the recurrence of the disease in the patient.

[0027] The present invention, as Figure 1 shown, includes: A human behavior information acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a human behavior classification and recognition module, and a human behavior information application module that are connected in sequence.

[0028] The human behavior information acquisition module is used to collect the user's behavior data in real time. The behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer. The user can wear multiple wearable devices at different positions on the body according to needs, and various sensors are built in. The human behavior information acquisition module includes several different sensor units, which are integrated in the wearable device and worn at the corresponding positions on the user's body respectively. The sensor unit includes a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer, and is used to collect the user's behavior data.

[0029] The human body behavior information transmission module is used to transmit the collected behavior data of the user to the local server or the cloud server. According to the scenario where the user is located, a suitable transmission method can be selected to transmit the collected behavior data of the user to the local server or the cloud server through any one of the transmission methods of Ultra Wide Band (UWB), Bluetooth, and 5G. Different transmission methods are different in terms of power consumption, transmission distance, transmission speed, etc., so they are suitable for different application scenarios.

[0030] The human body behavior information preprocessing module is used to sequentially store, merge, filter, normalize or standardize, perform sliding window segmentation, and label calibration on the collected behavior data of the user, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels.

[0031] The human body behavior classification and recognition module is used to input the behavior data of the user preprocessed by the human body behavior information preprocessing module into the trained human body behavior classification and recognition module for behavior type discrimination.

[0032] The human body behavior information application module is used to transmit the obtained human body behavior classification and recognition results to various application platforms, so as to realize corresponding functions. The human body behavior information application module faces various application platforms, such as various application platforms like the elderly care management platform and the rehabilitation nursing monitoring management platform. The results of model classification and recognition are transmitted to the databases of each application platform in real time for storage and management, and the behaviors of users are analyzed, managed, and visualized in real time.

[0033] This system realizes the real-time monitoring and evaluation of the intelligent rehabilitation training actions of patients through the processing and analysis of sensor data. In the human body behavior information classification module, a new deep learning framework for wearable human body behavior recognition based on a sparse three-channel hybrid attention spatio-temporal parallel multi-scale network is proposed, which can independently and parallelly learn the spatio-temporal features of multi-source sensors at different body positions from single-channel and inter-channel relationships, thus improving the accuracy of human body behavior recognition; at the same time, sparse attention is used to retrain the network nodes with lower weights to prevent model overfitting and further improve the robustness and generalization ability of the model.

[0034] Embodiment 2

[0035] The human body behavior classification and recognition module includes a scanning expansion module, a multi-scale feature segmentation module, a multi-scale dynamic convolution module with sparse three-channel hybrid attention, a global attention spatio-temporal fusion module, and a human body behavior classification result output module connected in sequence.

[0036] The scanning expansion module is used to perform data enhancement processing on the input time series data, expanding it into four time sorting combinations, realizing time feature mixed order, and enhancing feature learning effect; the scanning expansion module first converts the feature dimension into Time series data Convert to The scanning expansion module then performs flattening operations in four orders: horizontally from upper left to lower right, vertically from upper right to lower left, horizontally from lower right to upper left, and vertically from lower right to upper left; the scanning expansion module finally merges the four types of flattened time series data to obtain data enhancement samples, and the dimensions of the new samples are The scanning expansion module expands data of different orders along four different directions, disrupts the feature order, and realizes effective learning of numerical features with long-distance dependencies in the same receptive field. It ensures that each element in the feature integrates information with other positions in different directions, effectively learns the time information with long-distance dependencies in the input features, and thus forms a global receptive field. It is convenient for the subsequent feature extraction module to extract different features from different angles, enriching the diversity of features. The multi-scale feature segmentation module is used to separate the features of sensors at different positions; the multi-scale feature segmentation module can segment the multi-source sensor features at different positions into groups according to the body position grouping method: IMU1, IMU2, ..., IMU n ; The grouping method of the multi-scale feature segmentation module can input the grouped data into several parallel similar block networks in the subsequent feature extraction stage, so as to facilitate the extraction of high-level feature representations of sensor data at different locations in the same activity; The sparse three-channel mixed attention multi-scale dynamic convolution module is used for deep feature extraction; the sparse three-channel mixed attention multi-scale dynamic convolution module includes a three-channel dynamic convolution unit based on improved squeeze excitation, a three-channel mixed attention unit, a Patch Merging unit, and a sparse attention unit connected in sequence; the three-channel dynamic convolution unit based on improved squeeze excitation is used to dynamically and independently learn the attention and weighted convolution layers of each channel (long channel, wide channel and feature channel) to improve the independent learning ability of different channel features; the three-channel mixed attention unit is used to fully mix features across channels and capture the feature relationship between channels; the sparse attention unit is used to prevent model overfitting when training low contribution rate nodes.

[0037] The improved squeeze excitation three-channel dynamic convolution unit includes an initial convolution block 0 to obtain potential features of each channel, and three sequentially connected convolution blocks 1, convolution block 2, and convolution block 3. Convolution block 0 performs two-dimensional convolution calculation on the input data with a convolution kernel size of 4×4 and a stride of 1, followed by batch normalization and ReLU activation function to generate data of 64×H0×W0; convolution block 1 performs two-dimensional convolution calculation on the data of 64×H0×W0 through two groups of convolution kernels with a size of 3×3 and a stride of 1, followed by batch normalization and ReLU activation function to generate data of 32×H1×W1; convolution block 2 performs two-dimensional convolution calculation on the data of 32×H1×W1 through two groups of convolution kernels with a size of 3×3 and a stride of 1, followed by batch normalization and ReLU activation function to generate data of 64×H2×W2; convolution block 3 performs two-dimensional convolution calculation on the data of 64×H2×W2 through two groups of convolution kernels with a size of 3×3 and a stride of 1, followed by batch normalization and ReLU activation function to generate data of 128×H3×W3; in addition, the improved squeeze excitation three-channel dynamic convolution unit implements a convolutional layer that dynamically and independently learns the attention of each channel (long channel, wide channel, and feature channel) and weights it to improve the independent learning ability of different channel features; the potential feature 64×H0×W0 output by convolution block 0 passes through two-dimensional adaptive pooling with a size of 1×1 and two fully connected layers to generate independent attention for each channel (long channel 1×H0×1, wide channel 1×1×W0, and feature channel 64×1×1); then the attention of each channel is summed to obtain attention 64×H0×W0; finally, the attention 64×H0×W0 adaptively adjusts the dimension through a fully connected layer and weights the convolutional layers of each module after convolution block 0.

[0038] The three-channel hybrid attention unit consists of three parts: long and wide channel hybrid attention, long and feature channel hybrid attention, and wide and feature channel hybrid attention; for the long and wide hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on the improved squeeze-and-excitation is subjected to Max-Average self-attention, that is, the feature channel attention is reduced to 2 by combining the average pooling operation and the maximum pooling operation, generating data of 2×H3×W3; through two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the feature channel attention is reduced to 1, fixing the feature channel attention, obtaining the long and wide channel hybrid attention, and generating data of 1×H3×W3; for the long and feature channel hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on the improved squeeze-and-excitation is subjected to Max-Average self-attention, that is, the wide channel attention is reduced to 2 by combining the average pooling operation and the maximum pooling operation, generating data of 128×H3×2; through two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the wide channel attention is reduced to 1, fixing the wide channel attention, obtaining the long and feature channel hybrid attention, and generating data of 128×H3×1; for the wide and feature channel hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on the improved squeeze-and-excitation is subjected to Max-Average self-attention, that is, the long channel attention is reduced to 2 by combining the average pooling operation and the maximum pooling operation, generating data of 128×2×W3; through two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the long channel attention is reduced to 1, fixing the long channel attention, obtaining the wide and feature channel hybrid attention, and generating data of 128×1×W3; the three hybrid attention weights output by the three-channel hybrid attention unit at each position are merged and multiplied to generate feature mixed attention data.

[0039] The Patch Merging unit is used to shuffle different sensor features, randomize the sensor sources of the features, and achieve the merging of different sensor features; the attention map is passed through the shuffle layer to generate a shuffled feature map with odd and even skip connections; through the splicing layer, data with a feature length 4 times the original and long and wide channels 1 / 2 of the original is generated; through a linear layer, the feature dimension is halved, generating data with a feature length 2 times the original and long and wide channels 1 / 2 of the original.

[0040] The sparse attention unit is used to prevent the model from overfitting when training low contribution rate nodes. It consists of two mask layers, an encoding layer and a decoding layer; each attention vector length passes through a mask layer. The mask module sets the mask rate to 0.75, sets the attention with the top 0.75 weights before sorting to 0, and the rest to 1; then the encoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 2 of the original; passing through batch normalization and the ReLU activation function; performing two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 4 of the original; then the decoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel that is 1 / 2 of the original; passing through batch normalization and the ReLU activation function; performing two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1, generating data with a feature channel of the original dimension.

[0041] The global attention spatio-temporal fusion module is used to fully fuse spatio-temporal attention features; the global attention spatio-temporal fusion module includes a global attention unit and a one-dimensional standard convolution unit connected in sequence; the global attention unit combines a gated linear unit and an attention mechanism, and the input data is not directly multiplied by the attention parameters, but first mapped and processed by the gated linear unit; the one-dimensional standard convolution unit further extracts the output of the global attention to ensure that local information is not lost.

[0042] The human behavior classification result output module is used to reduce the dimension and classify the high-dimensional features obtained by feature extraction and fusion to achieve accurate classification of human behaviors, and update the parameters of the recognition model according to the loss function. It includes a flattening layer, a layer normalization, a fully connected layer and a Softmax output layer connected in sequence; the classification result output module calculates and outputs the recognition result through the Softmax layer.

[0043] The human behavior information preprocessing module includes a sensor data storage unit, a sensor data merging unit, a sensor data smoothing and filtering unit, a data normalization / standardization unit, a sliding window segmentation unit and a data label calibration unit connected in sequence.

[0044] The sensor data storage unit is used to store the real-time behavior data of the user.

[0045] The sensor data merging unit is used to splice and merge the behavior data of different sensor units. The merged data is in the format of a two-dimensional array, and is arranged horizontally in the order of the X, Y, and Z axis data of the three-axis acceleration sensor, the X, Y, and Z axis data of the three-axis angular velocity sensor, and the X, Y, and Z axis data of the three-axis magnetometer, and is arranged vertically in chronological order. The frequency of sensor data acquisition can be set to an appropriate value, such as , that is, the size of the data generated per second is 128 rows by 27 columns.

[0046] The sensor data smoothing and filtering unit is used to remove noise in the original sensor data; generally, the sensor data contains a large amount of high-frequency noise, and mean smoothing filtering can be used to remove this noise.

[0047] The data normalization / standardization unit is used to convert data with different dimensions and value ranges to the same value range; to avoid the adverse effects of different dimensions and value ranges on calculations. Normalization can use the Min-Max normalization method, that is, transform the values of all similar sensors to between [0,1]. Standardization can use the Z-Score standardization method, that is, convert the data into data with a mean of 0 and a standard deviation of 1, so that the original incomparable values can be made comparable.

[0048] The sliding window segmentation unit is used to segment continuous long-term sensor data into several data segments; suitable for input to the recognition model. According to the characteristics of human behavior, the length of one window can be set to 128, that is, it can contain seconds of sensor data, which can basically include a complete set of common human behavior data. The sliding step can be set to 64, that is the data coverage rate.

[0049] The data label calibration unit is used to calibrate the labels of the data after sliding window segmentation. The size of one data segment is rows columns, is the number of columns after merging all sensor data. When using triaxial acceleration sensors, triaxial angular velocity sensors, and triaxial magnetometers located on the wrist, waist, and chest, the size of one data segment is 128 rows by 27 columns. Each data segment corresponds to a corresponding label data, and one-hot encoding is performed on the labels.

[0050] The collected user behavior data is transmitted to the local server or cloud server through any one of the transmission methods of ultra-wideband, Bluetooth, and 5G.

[0051] The human behavior information application module faces various application platforms. The classification and recognition results are transmitted to the databases of each application platform in real time for storage and management, and the user's behavior is analyzed, managed, and visualized in real time.

[0052] The recognition method of the human behavior recognition system based on sparse three-channel hybrid attention, as Figure 2 shown, includes the following steps: S1: Human body behavior data collection; Collect the behavior data of users in real time, that is, human body behavior data. The behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer. Wearable devices with several built-in different sensor units are respectively worn at different positions of the user's body. The sensor units mainly include a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer to collect the user's behavior information. S2: Human body behavior data transmission; Transmit the collected behavior data of the user to the local server or the cloud server. According to different scenarios and user requirements, select a suitable data transmission method, such as UWB, Bluetooth, 5G, etc. S3: Preprocessing of human body behavior data; Store, merge, filter, normalize or standardize, segment with a sliding window, and label the collected behavior data of the user in sequence, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels. The specific implementation process of S3 is as follows: S31: Merging of human body behavior data; Stitch and merge the behavior data of the user collected at three positions: the wrist, the waist, and the chest. The merged data is in the format of a two-dimensional array, arranged horizontally in the order of the X, Y, and Z axis data of the three-axis acceleration sensor, the X, Y, and Z axis data of the three-axis angular velocity sensor, and the X, Y, and Z axis data of the three-axis magnetometer, and arranged vertically in chronological order. The frequency of sensor data collection Can be set to an appropriate value, such as Hz, that is, the size of the data generated per second is 128 rows and 27 columns. S32: Filtering of human body behavior data; Filtering of human body behavior data refers to smoothing filtering of sensor data; Smoothing filtering of sensor data uses three-point mean smoothing filtering to eliminate the noise signal inside the sensor. Let the sensor signal sequence be , Indicates the number of data points in the sensor signal sequence; The formula for mean smoothing is shown in formulas (5) and (6): (5); (6); Where: Indicates the value after one-time mean smoothing, Indicates the value after two-time mean smoothing; S33: Normalization / Standardization of human body behavior data; Convert data with different dimensions and value ranges to the same value range to avoid the adverse effects of different dimensions and value ranges on calculations. Use the deviation (Min-Max) normalization method for normalization, that is, transform the values of all sensors of the same type to between [0, 1]; the deviation (Min-Max) normalization method scales all data values to between [0, 1] through a linear transformation of the original data; let the input sequence be , and the output sequence after deviation (Min-Max) normalization is , as shown in Equation (7): (7); Where: and are the maximum and minimum values of all samples in the input sequence respectively; Use the standard score (Z-Score) normalization method for normalization, that is, transform the data into data with a mean of 0 and a standard deviation of 1; the Z-Score normalization method transforms the data into data with a mean of 0 and a standard deviation of 1, as shown in Equation (8): (8); Where: is the mean of all samples in the input sequence , is the standard deviation of all samples in the input sequence ; S34: Sliding window segmentation and labeling; Sliding window segmentation uses a window of fixed length to segment continuous sensor data into data segments of fixed length for input into the recognition model. Set the sensor acquisition frequency to Hz, and according to the characteristics of human behavior, select a window of length , that is, each window contains seconds of sensor data. The sliding step can be set to 64, that is, data coverage rate, so as to segment the data into a size of rows columns, is the number of columns after merging all sensor data. When using three-axis accelerometers, three-axis gyroscopes, and three-axis magnetometers located at the wrist, waist, and chest respectively, the size of a data segment is 128 rows and 27 columns. Assign a corresponding label to each data segment, that is, the name of the behavior category corresponding to the data segment, and use one-hot encoding to encode the label.

[0053] S4: Construct an identification model for human behavior classification and recognition; Input the behavior data segments preprocessed in S3 into the identification model in batches, and obtain the human behavior recognition output through training;

[0054] Such as Figure 3 shown, the specific implementation process of S4 is as follows: S41: Scanning expansion is to perform data augmentation on the input time series data using shuffled arrangement and flattening operations, expand it into four time sorting combinations, realize time feature shuffling, and enhance the feature learning effect. Specifically, it means: The size of each feature dimension is The time series data of is converted into types of image data through reshape, where ; Each data is flattened in four orders: horizontally from top left to bottom right, vertically from top right to bottom left, horizontally from bottom right to top left, and vertically from bottom right to top left, to obtain and . The calculation process of its shuffled arrangement and flattening is shown in Equation (9): (9); Merge the four new time series and to obtain the data augmentation sample , as shown in Equation (10): (10); Among them: Concat represents the merge connection operation; S42: Multi-scale feature segmentation is to separate the features of sensors at different positions. In the subsequent feature extraction stage, the grouped data is input into several parallel similar block networks to facilitate the extraction of high-level feature representations of sensor data at different positions in the same activity. Specifically, it means: Group the data augmentation sample processed in S41 according to the body position grouping method of wearable sensors, and segment the multi-source sensor features at different positions into groups: IMU1, IMU2,..., IMUn; S43: Sparse three-channel hybrid attention multi-scale dynamic convolution is to perform deep extraction and fusion of features; The multi-scale dynamic convolution module with sparsified three-channel hybrid attention includes a three-channel dynamic convolution unit based on improved squeeze-and-excitation, a three-channel hybrid attention unit, and a sparse attention unit, which are connected in sequence. The three-channel dynamic convolution unit based on improved squeeze-and-excitation is a convolution layer that dynamically and independently learns the attention of each channel (long channel, wide channel, and feature channel) and weights it to improve the independent learning ability of different channel features. The three-channel hybrid attention unit is used to fully mix features across channels and capture the feature relationships between channels. The sparse attention unit is used to prevent model overfitting during the training of low contribution rate nodes. The multi-scale dynamic convolution module with sparsified three-channel hybrid attention in S43 includes a three-channel dynamic convolution unit based on improved squeeze-and-excitation, a three-channel hybrid attention unit, a Patch Merging unit, and a sparse attention unit, which are connected in sequence. The specific implementation process is as follows: S431: The three-channel dynamic convolution unit based on improved squeeze-and-excitation includes an initial convolution block 0 to obtain the potential features of each channel, and three convolution blocks 1, 2, and 3 connected in series in sequence. Convolution block 0 realizes the two-dimensional convolution calculation of the input data with a convolution kernel size of 4×4 and a stride of 1, and then passes through batch normalization and the rectified linear unit ReLU (Rectified Linear Unit, ReLU) activation function to generate data of 64×H0×W0. Convolution block 1 realizes the two-dimensional convolution calculation of the 64×H0×W0 data with two convolution kernels of size 3×3 and a stride of 1, and then passes through batch normalization and the ReLU activation function to generate data of 32×H1×W1. Convolution block 2 realizes the two-dimensional convolution calculation of the 32×H1×W1 data with two convolution kernels of size 3×3 and a stride of 1, and then passes through batch normalization and the ReLU activation function to generate data of 64×H2×W2. Convolution block 3 realizes the two-dimensional convolution calculation of the 64×H2×W2 data with two convolution kernels of size 3×3 and a stride of 1, and then passes through batch normalization and the ReLU activation function to generate data of 128×H3×W3. The convolutional layer implemented by the three-channel dynamic convolution unit based on improved squeeze excitation realizes dynamic and independent learning of the attention of each channel (long channel, wide channel, and feature channel) and weights it to improve the independent learning ability of different channel features; the latent feature 64×H0×W0 output by convolutional block 0 is passed through a 1×1 two-dimensional adaptive pooling and two fully connected layers to generate independent attention for each channel (long channel 1×H0×1, wide channel 1×1×W0, and feature channel 64×1×1); then the attention of each channel is summed to obtain the attention 64×H0×W0; finally, the attention 64×H0×W0 adaptively dimensions through a fully connected layer and weights the convolutional layers of each module after convolutional block 0. S432: The three-channel hybrid attention unit consists of three parts: long and wide channel hybrid attention, long and feature channel hybrid attention, and wide and feature channel hybrid attention. For the long and wide hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on improved squeeze excitation uses Max-Average self-attention, that is, the way of combining average pooling operation and maximum pooling operation to reduce the feature channel attention to 2, generating data of 2×H3×W3; through a two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the feature channel attention is reduced to 1, fixing the feature channel attention, obtaining the long and wide channel hybrid attention, and generating data of 1×H3×W3. For the long and feature channel hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on improved squeeze excitation uses Max-Average self-attention, that is, the way of combining average pooling operation and maximum pooling operation to reduce the wide channel attention to 2, generating data of 128×H3×2; through a two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the wide channel attention is reduced to 1, fixing the wide channel attention, obtaining the long and feature channel hybrid attention, and generating data of 128×H3×1. For the wide and feature channel hybrid attention, first, the output 128×H3×W3 of the three-channel dynamic convolution unit based on improved squeeze excitation uses Max-Average self-attention, that is, the way of combining average pooling operation and maximum pooling operation to reduce the long channel attention to 2, generating data of 128×2×W3; through a two-dimensional convolution calculation with a convolution kernel size of 3×3 and a stride of 1, the long channel attention is reduced to 1, fixing the long channel attention, obtaining the wide and feature channel hybrid attention, and generating data of 128×1×W3. Merge the output of the three-channel hybrid attention unit at each position, multiply the three hybrid attention weights, and generate the feature mixed attention data. S433: The Patch Merging unit is used to shuffle different sensor features, randomize the sensor sources of the features, and realize the merging of different sensor features. The specific implementation process is as follows: In S432, the finally output attention map passes through a shuffle layer to generate a shuffled feature map with odd and even skip connections; through a concatenation layer, data with a feature length 4 times the original and long and wide channels 1 / 2 of the original is generated; through a linear layer, the feature dimension is halved to generate data with a feature length 2 times the original and long and wide channels 1 / 2 of the original. S434: The sparse attention unit is used to prevent the model from overfitting during the training of low contribution rate nodes, and it consists of 2 mask layers, an encoding layer and a decoding layer. Each attention vector with a length of passes through a mask layer, and the mask module sets the mask rate to 0.75, sets the attention with the top 0.75 before weight sorting to 0, and the rest to 1. Its masking process is shown in equations (12)-(14) as follows: (12); (13); (14); where: is the reverse sorting function, is the attention sequence sorted from largest to smallest with a length of which will be used as the input of the encoding module to achieve the purpose of retraining the attention nodes with small contributions; is the mask function; is the attention sequence after mask processing; The encoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 2 of the original feature channels; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 4 of the original feature channels. The decoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 2 of the original feature channels; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with the original dimensional feature channels. S44: The global attention spatio-temporal fusion module is used to achieve full fusion of spatio-temporal attention features; the global attention spatio-temporal fusion module includes a global attention unit and a one-dimensional standard convolution unit connected in sequence, and the specific implementation process is as follows:

[0055] S44: The global attention spatio-temporal fusion module is used to achieve full fusion of spatio-temporal attention features; the global attention spatio-temporal fusion module includes a global attention unit and a one-dimensional standard convolution unit connected in sequence, and the specific implementation process is as follows: The global attention unit combines the gated linear unit and the attention mechanism. Through the mapping process of the gated linear unit, matrices and are generated. Then, the self-attention mechanism is used for key feature extraction, and its calculation process is shown in Equation (11): (11); Where: is the output of the global attention unit; is the weight bias; is the query matrix, and are the key and value represented by the matrix respectively, is the ReLU activation function; Further, through the one-dimensional convolution calculation with a convolution kernel size of 3 and a stride of 1, data with a feature channel of 1 is generated, and the output of the global attention is further extracted to ensure the non-loss of local information; S45: The human behavior classification result output module is used to reduce the dimension and classify the high-dimensional features obtained by feature extraction and fusion to achieve accurate classification of human behaviors, including a flattening layer, a layer normalization, a fully connected layer, and a Softmax output layer connected in sequence.

[0056] S5: Loss function calculation, update, and feedback;

[0057] Calculate the loss function for the classification result generated by S4, and adjust the model parameters through backpropagation to improve the classification accuracy and generalization ability. Use the cross-entropy loss function to calculate the error between the model prediction value and the true value. The cross-entropy loss function is defined as shown in Equation (1): (1); Where: represents the batch data volume, represents the total number of categories, represents the true label that the th sample belongs to category , represents the predicted probability that the th sample belongs to category ; (2); Where: represents the model parameters at the th step, represents the model parameters at the th step, represents the learning rate, is a constant added to increase numerical stability, represents the first-order momentum, represents the second-order momentum (exponentially weighted average of the squared gradients), and represents the first-order and second-order average coefficients and to the power; The calculation process of the first-order momentum can be expressed as Equation (3): (3); where: represents the gradient of the loss function , represents the first-order average coefficient, represents the current first-order momentum value, represents the first-order momentum value at the previous moment; The calculation process of the second-order momentum can be expressed as Equation (4): (4); where: represents the second-order average coefficient, represents the current second-order momentum value, represents the second-order momentum value at the previous moment; In each round of training, the data is input into the model in batches for forward propagation and backward propagation in sequence until all the data is trained. During the training process, the performance of the model is evaluated using the validation set data, and the learning rate or other hyperparameters are dynamically adjusted according to the validation set results to avoid overfitting or underfitting problems. When a round of training ends, only the loss function in Equation (2) is used to return to S4 for training until convergence, so as to further adjust the parameters of the recognition model; S6: Application of the human behavior classification output result; The human behavior classification and recognition result is transmitted to the corresponding application platform in real time.

[0058] The result output by the human behavior recognition is transmitted to the intelligent rehabilitation training action monitoring platform for patients. The results are stored, processed and calculated, and displayed in real time, so as to realize the real-time monitoring and evaluation of the intelligent rehabilitation training actions of patients.

[0059] The above-disclosed are only specific embodiments of the present invention. However, the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A human behavior recognition system based on sparse three-channel hybrid attention, characterized in that Including: A human behavior information acquisition module, a human behavior information transmission module, a human behavior information preprocessing module, a human behavior classification and recognition module, and a human behavior information application module that are connected in sequence; The human behavior information acquisition module is used to collect the behavior data of the user in real time. The behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer; The human behavior information transmission module is used to transmit the collected behavior data of the user to the local server or the cloud server; The human behavior information preprocessing module is used to sequentially store, merge, filter, normalize or standardize, perform sliding window segmentation, and label calibration on the collected behavior data of the user, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels; The human behavior classification and recognition module is used to input the behavior data of the user preprocessed by the human behavior information preprocessing module into the trained human behavior classification and recognition module for behavior type discrimination; The human behavior information application module is used to transmit the obtained human behavior classification and recognition result to various application platforms, so as to realize corresponding functions.

2. The human behavior recognition system based on sparse three-channel hybrid attention according to claim 1, characterized in that: The human behavior classification and recognition module includes a scanning expansion module, a multi-scale feature segmentation module, a multi-scale dynamic convolution module with sparse three-channel hybrid attention, a global attention spatio-temporal fusion module, and a human behavior classification result output module that are connected in sequence.

3. The human behavior recognition system based on sparse three-channel hybrid attention according to claim 2, characterized in that: The human behavior information acquisition module includes a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer for collecting the behavior data of the user.

4. The human behavior recognition system based on sparse three-channel hybrid attention according to claim 3, characterized in that: The human behavior information preprocessing module includes a sensor data storage unit, a sensor data merging unit, a sensor data smoothing and filtering unit, a data normalization / standardization unit, a sliding window segmentation unit, and a data label calibration unit that are connected in sequence.

5. The human behavior recognition system based on sparse three-channel hybrid attention according to claim 1, characterized in that: The collected behavior data of the user is transmitted to the local server or the cloud server through any one of the transmission methods of ultra-wideband, Bluetooth, and 5G.

6. A recognition method of the human behavior recognition system based on sparse three-channel hybrid attention according to claim 4, characterized in that Including the following steps: S1: Human behavior data acquisition; Collect the behavior data of the user in real time, that is, human behavior data. The behavior data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer; S2: Human behavior data transmission; Transmit the collected behavior data of the user to the local server or the cloud server; S3: Human behavior data preprocessing; Sequentially store, merge, filter, normalize or standardize, perform sliding window segmentation, and label calibration on the collected behavior data of the user, so as to obtain smooth behavior data segments that are not affected by the dimension and carry labels; S4: Construct an identification model and human behavior classification and recognition; Input the behavior data segments preprocessed in S3 into the identification model in batches, and obtain the human behavior recognition output after training; S5: Loss function calculation, update and feedback; Calculate the loss function for the classification results generated by S4, and adjust the model parameters through backpropagation to improve the classification accuracy and generalization ability. Use the cross-entropy loss function to calculate the error between the model prediction value and the true value. The cross-entropy loss function is defined as shown in Equation (1): (1); Wherein: represents the batch data volume, represents the total number of categories, represents the th sample belongs to the category true label, represents the th sample belongs to the category predicted probability; (2); Wherein: represents the step model parameter, represents the step model parameter, represents the learning rate, is a constant added to increase numerical stability, represents the first-order momentum, represents the second-order momentum, and represent the first-order and second-order averaging coefficients and to the power; The calculation process of the first-order momentum can be expressed as Equation (3): (3); Wherein: represents the gradient of the loss function , represents the first-order average coefficient represents the current first-order momentum value represents the first-order momentum value at the previous moment; The calculation process of the second-order momentum can be expressed as Equation (4): (4); Wherein: represents the second-order average coefficient, represents the current second-order momentum value, represents the second-order momentum value at the previous moment; S6: Application of the human behavior classification output result; Transmit the human behavior classification and recognition results to the corresponding application platform in real time.

7. The recognition method according to claim 6, wherein: The specific implementation process of S3 is as follows: S31: Merging of human behavior data; Stitch and merge the collected behavior data of the user. The merged data is in the format of a two-dimensional array, arranged horizontally in the order of the X, Y, and Z axis data of the triaxial acceleration sensor, the X, Y, and Z axis data of the triaxial angular velocity sensor, and the X, Y, and Z axis data of the triaxial magnetometer, and arranged vertically in chronological order; S32: Filtering of human behavior data; Filtering of human behavior data refers to smoothing filtering of sensor data; The sensor data smoothing filter uses a three-point mean smoothing filter to eliminate the noise signals inside the sensor. Let the sensor signal sequence be , represent the number of data points in the sensor signal sequence; the equations for mean smoothing are shown in Equations (5) and (6) as follows: (5); (6); Wherein: represents the value after one-time mean smoothing, represents the value after two-time mean smoothing; S33: Normalization / standardization of human behavior data; Normalization is performed using the deviation normalization method. Let the input sequence be , and the output sequence after deviation normalization is , as shown in Equation (7): (7); Wherein: and are respectively the maximum and minimum values of all samples in the input sequence; Use the standard score normalization method for standardization processing, as shown in Equation (8): (8); Wherein: is the mean value of all samples in the input sequence, is the standard deviation of all samples in the input sequence; S34: Sliding window segmentation and labeling; Sliding window segmentation uses a window with a fixed length to divide the continuous sensor data into data segments with a fixed length, assign corresponding labels to each data segment, that is, the name of the behavior category corresponding to the data segment, and encode the labels using one-hot encoding.

8. The recognition method according to claim 7, wherein: The specific implementation process of S4 is as follows: S41: Scan expansion is to perform data augmentation processing on the input time series data using scrambled arrangement and flattening operations, expand it into four time sorting combinations, achieve time feature scrambling, and enhance the feature learning effect. Specifically, it means: The size of each feature dimension is to reshape the time series data with size into image-like data through reshape, where ; ; Flatten each data in four orders: horizontally from the upper left to the lower right, vertically from the upper right to the lower left, horizontally from the lower right to the upper left, and vertically from the lower right to the upper left, to obtain and . The calculation process of its shuffled and flattened arrangement is shown in Equation (9): (9); Combine four new time series and to obtain data augmentation samples , as shown in Equation (10): (10); Where: Concat represents the merge connection operation; S42: Multi-scale feature segmentation is to separate the sensor features at different positions. In the subsequent feature extraction stage, input the grouped data into several parallel similar block networks to facilitate the extraction of high-level feature representations of sensor data at different positions in the same activity. Specifically, it means: The data augmentation samples that have passed through the step S41 According to the grouping method of the body positions where the wearable sensors are worn, the multi-source sensor features at different positions are segmented into groups: IMU1, IMU2,..., IMUn; S43: Sparse three-channel hybrid attention multi-scale dynamic convolution is to perform deep extraction and fusion of features; S44: The global attention spatio-temporal fusion module is used to achieve full fusion of spatio-temporal attention features: The global attention unit combines the gated linear unit and the attention mechanism. Through the mapping process of the gated linear unit, matrices and are generated. Then, the self-attention mechanism is used for key feature extraction, and its calculation process is shown in Equation (11): (11); Wherein: is the output of the global attention unit; is the weight bias; is the query matrix, and are the key and value represented by the matrix respectively, is the ReLU activation function; Further perform one-dimensional convolution calculation with a convolution kernel size of 3 and a stride of 1 to generate data with a feature channel of 1, and further extract the output of the global attention to ensure that local information is not lost; S45: The human behavior classification result output module is used to reduce the dimension and classify the high-dimensional features obtained by feature extraction and fusion to achieve accurate classification of human behavior, including a flattening layer, a layer normalization, a fully connected layer, and a Softmax output layer connected in sequence.

9. The recognition method according to claim 8, wherein: The specific implementation process of S43 is as follows: S431: The improved squeeze-and-excitation three-channel dynamic convolution unit includes an initial convolution block 0 to obtain potential features of each channel, and three convolution blocks 1, 2, and 3 connected in series in sequence; S432: The three-channel hybrid attention unit consists of three parts: the long- and width-channel hybrid attention, the long- and feature-channel hybrid attention, and the width- and feature-channel hybrid attention; S433: The Patch Merging unit is used to shuffle the features of different sensors, randomize the sensor sources of the features, and achieve the merging of the features of different sensors. The specific implementation process is as follows: In S432, the finally output attention map passes through the shuffle layer to generate a shuffled feature map with odd-even skip connections; through the concatenation layer, a data with a feature length 4 times the original and long and width channels 1 / 2 of the original is generated; through a linear layer, the feature dimension is halved to generate a data with a feature length 2 times the original and long and width channels 1 / 2 of the original; S434: The sparse attention unit is used to prevent the model from overfitting when training low contribution rate nodes. It consists of two mask layers, an encoding layer, and a decoding layer; Each attention vector with a length of is passed through a masking layer, and the masking module sets the masking rate to 0.

75. The top 0.75 of the weights in the sorted attention are set to 0, and the rest are set to 1. The masking process is shown in equations (12)-(14) as follows: (12); (13); (14); Wherein: is the reverse sorting function, is of length the attention sequence sorted from largest to smallest, will be used as the input of the encoding module; is the masking function; is the attention sequence after masking; The encoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 2 of the original feature channels; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 4 of the original feature channels; The decoding layer first performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with 1 / 2 of the original feature channels; then passes through batch normalization and the ReLU activation function; then performs two-dimensional convolution calculations with a convolution kernel size of 3×1 and a stride of 1 to generate data with the original dimensional feature channels.

Citation Information

Patent Citations

  • Behavior recognition system based on space-time multi-feature extraction and working method thereof

    CN110852382A

  • Open set human body behavior recognition system, method and device based on multi-resolution fusion convolution and storage medium

    CN116432113A

  • Multi-scale space-time interaction skeleton action classification method, system and equipment for action jitter and skeleton noise suppression, and medium

    CN117671353A

  • Lightweight multi-attention-based feature extraction and fusion behavior recognition system and method

    CN117789298A

  • Human body behavior recognition method based on multi-branch sparse large kernel convolutional network

    CN118312876A