A Human Behavior Recognition System and Method Based on Multi-Dimensional and Multi-Scale Feature Extraction
Through the human behavior recognition system of multi-dimensional and multi-scale feature extraction, local and global adaptive spatiotemporal feature extraction, residual multi-parallel multi-attention feature extraction and feature fusion, the existing model's problems of insufficient feature extraction and difficulty in complex behavior recognition are solved, and high-accuracy behavior recognition is achieved.
Patent Information
- Application Number
- CN202510329344.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing human behavior recognition models do not fully utilize the correlation between features in feature extraction, it is difficult to adaptively capture complex dynamic information, and it is difficult to distinguish confusing behaviors from identifying long-term complex behaviors.
The human behavior recognition system that uses multi-dimensional and multi-scale feature extraction, including LGAM simple human behavior recognition network and MPFEFN complex human behavior recognition network, is used to fusion and discriminate features through local and global adaptive spatiotemporal feature extraction, residual multi-parallel multi-attention feature extraction and feature fusion, and uses M-reluGRU and multi-head self-attention mechanism for feature fusion and discrimination.
It improves the accuracy of recognition of confusing behaviors, realizes effective recognition of complex human behaviors, and improves the recognition performance of existing models.
Smart Images

Figure CN119851353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a human behavior recognition system and method based on multi-dimensional and multi-scale feature extraction. Background Art
[0002] Human Activity Recognition (HAR) refers to the intelligent recognition of human behaviors or activity states by collecting and analyzing input personal or group motion data. Currently, human activity recognition has become an important research content in the fields of artificial intelligence, pattern recognition, and human-computer interaction, and is widely used in application scenarios such as smart homes, healthcare, and security monitoring, with great commercial value and broad development prospects, and has received close attention from academia and industry. According to the type of data collected, the recognition methods are mainly divided into two categories: vision-based HAR and sensor-based HAR. The former analyzes image or video data, and the latter studies the time-series data collected by wearable sensors and environmental sensors. Compared with vision-based HAR, sensor-based HAR has the advantages of low cost, good privacy, and strong anti-interference ability.
[0003] With the progress and update iteration of related technologies, HAR algorithms and systems have entered a rapid development stage. Currently, the research on HAR mainly focuses on classification models based on deep neural networks. Deep neural networks can automatically extract behavior features, thereby realizing end-to-end behavior recognition and effectively improving the recognition accuracy. Convolutional Neural Network (CNN) is one of the most widely used deep neural networks currently. CNN belongs to a multi-layer stacked deep feedforward network. By processing the input data layer by layer, it realizes the gradual integration of information, transforms the original data into a higher-level feature representation that is more closely related to the output target, and finally completes the label mapping through a classifier. Compared with CNN, Recurrent Neural Network (RNN) pays more attention to the time-series features of data and can capture the correlation between time-series features. Therefore, RNN and its various variant networks such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are widely used in the construction of HAR models.
[0004] Currently, there are still many challenges in the field of human behavior recognition. First, in terms of feature extraction, existing models do not fully utilize the correlation between features and have limited ability to reshape low-level features into effective high-level representations. Second, in terms of information acquisition, existing frameworks cannot adaptively capture the complex dynamic information contained in human behavior motion patterns and speed changes. Finally, in terms of classification, most existing recognition frameworks are difficult to distinguish between easily confused behaviors, and usually can only classify and recognize simple behaviors in a short time, while it is difficult to recognize long-term complex behaviors composed of multiple simple behaviors. Summary of the Invention
[0005] In view of this, the present invention provides a human behavior recognition system and method based on multi-dimensional and multi-scale feature extraction, which can better utilize the multi-dimensional features contained in human behavior data, improve the recognition accuracy of easily confused behaviors, and achieve effective recognition of complex human behaviors, thereby enhancing the recognition performance of existing models.
[0006] In the first aspect, the present invention provides a human behavior recognition system based on multi-dimensional and multi-scale feature extraction, and the system includes:
[0007] A human behavior data acquisition module, a human behavior data transmission module, a human behavior data storage module, a human behavior data preprocessing module, an LGAM simple human behavior recognition network module, an MPFEFN complex human behavior recognition network module, and a human behavior information application module;
[0008] The LGAM simple human behavior recognition network module includes a feature pre-extraction unit, a local and global adaptive spatio-temporal feature extraction unit, a feature fusion unit, and a simple behavior discrimination output unit connected in sequence;
[0009] The MPFEFN complex behavior recognition network module includes an initialization unit, a residual multi-parallel multi-attention feature extraction unit and a residual multi-parallel channel enhancement and extraction unit connected in parallel, a feature fusion unit, and a complex behavior discrimination output unit.
[0010] Optionally, the human behavior data acquisition module includes several different types of motion data perception units and physiological data perception units. The motion data perception units include an acceleration data perception unit, an angular velocity data perception unit, and a magnetic induction intensity data perception unit; the physiological data perception units include a heart rate data perception unit, a blood pressure data perception unit, a blood oxygen data perception unit, and a surface electromyography data perception unit;
[0011] The transmission methods of the human behavior data transmission module include ultra-wideband UWB, Wi-Fi, long-range radio Lora, ZigBee, Bluetooth, 4G, 5G, and long-range radio information transmission;
[0012] The human behavior data preprocessing module includes a data denoising unit, a multi-modal data merging unit, a missing value processing unit, a data normalization unit, and a data sliding window segmentation unit, which are connected in sequence.
[0013] Optionally, the multi-modal data merging unit merges the motion data and physiological data collected by different sensors by aligning them with longitudinal timestamps and splicing them horizontally; the missing value processing unit uses the mean imputation method to complete the missing information, taking the average value of the column where the missing data is located as the missing value; the data normalization unit performs Z-Score normalization on data with different dimensions and value ranges and converts them to the same value range; the data sliding window segmentation unit divides the continuous time series collected by the sensor into multiple data segments.
[0014] Optionally, the feature pre-extraction unit uses a single one-dimensional convolutional module to preliminarily extract behavior features; the data after feature pre-extraction is sequentially input into the local and global adaptive feature extraction unit and the simple behavior discrimination unit; the features output by the feature fusion unit are the input to the initialization unit, the residual multi-parallel multi-attention feature extraction unit connected in parallel, and the residual multi-parallel channel enhancement and feature extraction unit connected in sequence; the local and global adaptive feature extraction unit extracts local and global features by adaptively adjusting the convolutional kernel size; the residual multi-parallel multi-attention feature extraction unit uses a residual multi-parallel structure and uses multiple attention mechanisms to extract multi-dimensional features contained in human behavior data; the residual multi-parallel channel enhancement and feature pre-extraction unit uses a residual multi-parallel structure and uses channel enhancement technology to extract channel features contained in human behavior data.
[0015] Optionally, the feature fusion unit uses M-reluGRU and the multi-head self-attention mechanism to effectively fuse the long time-series features in complex behaviors; M-reluGRU removes the reset gate in the gated recurrent unit GRU, simplifies the GRU to a single-gate structure, uses the ReLU function as the activation function, and uses batch normalization (BN) as the normalization method.
[0016] Optionally, the human behavior information application module includes a human behavior visualization unit, a human behavior statistics unit, and a human behavior analysis unit; the human behavior data storage module includes a cloud database and a local database.
[0017] Optionally, the simple behavior discrimination output unit includes a fully connected layer and a Softmax function connected in sequence. After multi-dimensional feature fusion, the data is output to the fully connected layer and classified by the Softmax function to output the final simple behavior type.
[0018] In a second aspect, the present invention provides a human behavior recognition method based on multi-dimensional and multi-scale feature extraction, which is implemented based on the human behavior recognition system based on multi-dimensional and multi-scale feature extraction in the first aspect or any possible implementation manner of the first aspect; the method includes:
[0019] S1. Human behavior data collection: The behavior information of the user is collected through the multi-modal sensors in the human behavior data collection module, and the behavior information includes motion data and physiological data;
[0020] S2. Human behavior data transmission: The behavior information collected by the human behavior data collection module is transmitted through the human behavior data transmission module to the databases of the local server and the cloud server of the human behavior data storage module for the human behavior data storage module to store the behavior information;
[0021] S3. Human behavior data preprocessing: The behavior information is sequentially subjected to human behavior data denoising, multi-modal human behavior data merging, human behavior data missing value processing, human behavior data normalization, and human behavior data sliding window segmentation through the human behavior data preprocessing module;
[0022] S4. Construct an LGAM simple human behavior recognition network for simple human behavior recognition and classification; the behavior data after feature preprocessing is input into the LGAM simple human behavior recognition network in batches for feature extraction and fusion, and a simple discrimination result is output through the simple behavior discrimination output unit;
[0023] S5. Construct an MPFEFN complex human behavior recognition network for complex human behavior recognition and classification; the initialized behavior data is input into the MPFEFN complex human behavior recognition network in batches for feature extraction and fusion, and a complex discrimination result is output through the complex behavior discrimination output unit;
[0024] S6. Feature fusion: The output feature data is fused through M-reluGRU and the multi-head self-attention mechanism and input into the Generalized Average Pooling (GAP) layer. The GAP layer converts the feature map of each channel into a feature point, and the feature point is the average value of the entire feature map;
[0025] S7. Output the action recognition result: The feature data obtained after multi-dimensional spatio-temporal feature extraction and fusion is input into the fully connected layer, and the number of hidden units in the fully connected layer is the number of simple behavior categories included in the data; the data after passing through the fully connected layer passes through the Softmax classifier to calculate the probability of the corresponding behavior, and the behavior type with the highest probability is the final judgment result;
[0026] S8. Human behavior information display and application. Through the human behavior information application module, the behavior recognition results are visually displayed, statistically analyzed, and analyzed.
[0027] Optionally, step S3 includes:
[0028] S31. Denoising of human behavior data; using the soft-threshold wavelet denoising method based on Stein's unbiased risk estimate to denoise the data collected by the sensor. , is the original signal, is the noisy signal, is the Gaussian white noise, , is the noise intensity. The denoising process is to remove the noise from the signal to obtain the best approximation of the original signal .
[0029] Discrete sampling is performed to obtain points of discrete signal . The wavelet transform coefficients are shown in Equation (1):
[0030] (1);
[0031] where is the wavelet coefficient, is the scaling function, j is the scale parameter, k is the unit number of the translation of the scaling function. Through the two-scale equations (2) and (3), the recursive implementation method of Equation (1) is obtained:
[0032] (2);
[0033] (3);
[0034] where the symbol * represents convolution, h and g represent the low-pass and high-pass filters respectively, represents the original signal , represents the approximation coefficient at the j scale. The wavelet transform reconstruction formula is shown in Equation (4):
[0035] (4);
[0036] The soft-threshold estimation method based on SURE is used to determine the threshold and obtain the likelihood estimate of the given threshold, as shown in Equation (5):
[0037] (5);
[0038] where represents the selected initial threshold. representing the wavelet coefficients from the sub-band , representing the sum of the numbers of the wavelet coefficients of each sub-band; minimizing the likelihood function to obtain the required threshold as shown in Equation (6):
[0039] (6);
[0040] wherein represents the obtained threshold parameter;
[0041] processing the wavelet transform coefficients of the behavior data using a soft threshold function, comparing the absolute value of the data with the threshold, setting the points less than the threshold to zero, and shrinking the points greater than or equal to the threshold towards zero to become the difference between the point value and the threshold. The soft threshold function is as shown in Equation (7):
[0042] (7);
[0043] Performing wavelet reconstruction of the signal according to Equation (3) to obtain the denoised signal;
[0044] S32. Merging multi-modal human behavior data; longitudinally aligning the behavior data collected by the sensor according to the time stamp and performing horizontal splicing and merging. The merged data is in the format of a two-dimensional array;
[0045] S33. Processing missing values in human behavior data; using the mean imputation method to process the missing values in the behavior data and filling the missing data with the average value of the column where the missing data is located;
[0046] S34. Normalizing human behavior data; using the Z-Score normalization method to normalize the human behavior data to ensure that the data is within the same order of magnitude range;
[0047] Assuming that the input data sample sequence is , the output sequence after Z-Score normalization is , and the calculation method is as shown in Equation (8):
[0048] (8);
[0049] wherein is the mean of the input data sample sequence, is the standard deviation of the input data sample sequence;
[0050] S35. Segmenting human behavior data with a sliding window, using a window with a fixed length to segment the continuous sensor data into data segments with a fixed length. When segmenting, ensure that each data segment contains at least one complete action of a simple behavior, and adopt a 50% window overlap rate during the window sliding process.
[0051] Optionally, step S4 includes:
[0052] S41: Input the human behavior data into a simple human behavior recognition network; transform the preprocessed data into a shape suitable for a one-dimensional convolutional layer, and input it into the feature extraction and fusion unit in batches. The data shape is , where is the batch size, is the number of data channels, is the window length;
[0053] Input the feature data output by the feature pre-extraction unit into the locally and globally adaptive spatio-temporal feature extraction unit and the feature fusion unit connected in sequence;
[0054] S42: Locally and globally spatio-temporal feature extraction; the data input by the input module is processed by the initialization unit. The initialization unit includes a one-dimensional convolutional module, which is composed of a one-dimensional convolutional layer, a batch normalization layer, an activation layer, and a max pooling layer in sequence. Among them, the convolutional kernel size of the one-dimensional convolutional layer is 3, the stride size is 1, the padding method is SAME, and the non-linear activation function selects the ReLU function; the calculation method of the one-dimensional convolutional module is shown in Equation (9):
[0055] ;
[0056] Among them, represents the th column of the feature map, represents the th column of the convolutional kernel, represents the bias term;
[0057] The locally and globally adaptive spatio-temporal feature extraction unit constructs an adaptive temporal kernel composed of a local branch and a global branch. Among them, the local branch adopts a locally sensitive modeling scheme, captures short-term temporal dynamic changes by establishing position-sensitive importance weights, and extracts as shown in Equation (10). The global branch is based on the idea of local invariant modeling, obtains global information and long-term temporal dependence information by constructing a dynamic aggregation convolutional kernel, and extracts as shown in Equation (11). The overall process is shown in Equation (12):
[0058] (10);
[0059] Among them, represents the local branch transformation, represents the multiplication operator, represents the one-dimensional convolutional operation, and represent the Sigmoid activation function and the ReLU activation function respectively;
[0060] (11);
[0061] Wherein represents the global branch transformation, represents the fully connected layer operation equivalent transformation function, respectively represent the weights of two FC layers; in order to realize the interaction between the information obtained by the global branch and the information obtained by the local branch, finally, the outputs of the two branches are convolved;
[0062] (12);
[0063] Wherein, represents the output feature of the time adaptive module, represents the convolution operator.
[0064] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the human behavior recognition method based on multi-dimensional and multi-scale feature extraction in the second aspect or any possible implementation manner of the second aspect.
[0065] In a fourth aspect, an embodiment of the present invention provides an electronic device, including: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, when the instructions are executed by the device, the device is made to execute the human behavior recognition method based on multi-dimensional and multi-scale feature extraction in the second aspect or any possible implementation manner of the second aspect.
[0066] In the technical solution provided by the present invention, the system includes a human behavior data acquisition module, a human behavior data transmission module, a human behavior data storage module, a human behavior data preprocessing module, an LGAM simple human behavior recognition network module, an MPFEFN complex human behavior recognition network module, and a human behavior information application module; the LGAM simple human behavior recognition network module includes a feature pre-extraction unit, a local and global adaptive spatio-temporal feature extraction unit, a feature fusion unit, and a simple behavior discrimination output unit connected in sequence; the MPFEFN complex behavior recognition network module includes an initialization unit, a residual multi-parallel multi-attention feature extraction unit and a residual multi-parallel channel enhancement and extraction unit connected in parallel, a feature fusion unit, and a complex behavior discrimination output unit. This system makes better use of the multi-dimensional features contained in human behavior data, improves the recognition accuracy of easily confused behaviors, and realizes the effective recognition of complex human behaviors, enhancing the recognition performance of existing models. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0068] Figure 1 It is a schematic structural diagram of a human behavior recognition system based on multi-dimensional and multi-scale feature extraction provided by an embodiment of the present invention;
[0069] Figure 2 It is a schematic structural diagram of a local and global adaptive spatio-temporal feature extraction unit provided by an embodiment of the present invention;
[0070] Figure 3 It is a schematic structural diagram of a residual multi-parallel multi-attention feature extraction unit provided by an embodiment of the present invention, where (a) is the detailed structure of FEMM and (b) is the overall framework of the residual multi-parallel multi-attention feature extraction unit;
[0071] Figure 4 It is a schematic structural diagram of a residual multi-parallel channel enhancement and extraction unit provided by an embodiment of the present invention, where (a) is the detailed structure of CBEM and (b) is the overall framework of the residual multi-parallel channel enhancement and extraction unit;
[0072] Figure 5 It is a flowchart of a human behavior recognition method based on multi-dimensional and multi-scale feature extraction provided by an embodiment of the present invention;
[0073] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0075] It should be clear that the described embodiments are only some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0076] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said", and "the" used in the embodiments of the present invention are also intended to include the plural forms unless the context clearly indicates otherwise.
[0077] It should be understood that the term "and / or" used herein is merely a relational description of associated objects, indicating that three relationships may exist. For example, a and / or b may represent: a exists alone, a and b exist simultaneously, and b exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0078] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0079] The human behavior recognition system based on multi-dimensional and multi-scale feature extraction provided by the present invention, as Figures 1 to 4 shown, the system includes:
[0080] a human behavior data acquisition module, a human behavior data transmission module, a human behavior data storage module, a human behavior data preprocessing module, an LGAM simple human behavior recognition network module, an MPFEFN complex human behavior recognition network module, and a human behavior information application module;
[0081] The said LGAM simple human behavior recognition network module includes a feature pre-extraction unit, a local and global adaptive spatio-temporal feature extraction unit, a feature fusion unit, and a simple behavior discrimination output unit connected in sequence;
[0082] The said MPFEFN complex behavior recognition network module includes an initialization unit, a residual multi-parallel multi-attention feature extraction unit and a residual multi-parallel channel enhancement and extraction unit connected in parallel, a feature fusion unit, and a complex behavior discrimination output unit. The initialization unit uses a single one-dimensional convolution and max pooling to preliminarily extract behavior features; the data after initialization is simultaneously input into a parallel multi-dimensional feature extraction unit, a multi-head attention unit, and a feature fusion unit connected in sequence;
[0083] In the embodiments of the present invention, the human behavior data acquisition module includes several motion data sensing units and physiological data sensing units of different types. The motion data sensing units include an acceleration data sensing unit, an angular velocity data sensing unit, and a magnetic induction intensity data sensing unit. The motion data includes the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, and the X, Y, and Z axis data of the magnetometer. The physiological data sensing units include a heart rate data sensing unit, a blood pressure data sensing unit, a blood oxygen data sensing unit, and a surface electromyogram data sensing unit. The physiological data includes heart rate values, blood pressure values, blood oxygen values, and surface electromyogram signal values.
[0084] The motion data sensing units mainly include a three-axis acceleration sensor, a three-axis angular velocity sensor, and a three-axis magnetometer. The physiological data sensing units mainly include a heart rate sensor, a blood pressure sensor, and a skin conductance sensor. The sampling frequencies of different types of sensors are set to the same value according to user requirements.
[0085] The transmission methods of the human behavior data transmission module include ultra-wideband UWB, Wi-Fi, long-range radio Lora, ZigBee, Bluetooth, 4G, 5G, and long-range radio information transmission. A suitable transmission method is selected according to the user's application requirements.
[0086] The human behavior data preprocessing module includes a data denoising unit, a multi-modal data merging unit, a missing value processing unit, a data normalization unit, and a data sliding window segmentation unit connected in sequence. Finally, smooth data segments that are not affected by the dimension and are convenient for model processing are obtained.
[0087] In the embodiments of the present invention, the multi-modal data merging unit merges the motion data and physiological data collected by different sensors in a manner of aligning longitudinally by time stamps and splicing horizontally. The missing value processing unit uses the mean imputation method to complete the missing information, taking the average value of the column where the missing data is located as the missing value. The data normalization unit performs Z-Score normalization processing on data with different dimensions and value ranges, converting them to the same value range, that is, transforming all sensor behavior data to between [-1, 1]. The data sliding window segmentation unit divides the continuous time series collected by the sensor into multiple data segments, ensuring that a complete simple behavior action data falls within one sliding window during the segmentation process.
[0088] In the embodiment of the present invention, the feature pre-extraction unit preliminarily extracts behavioral features by using a single one-dimensional convolution module; the data after feature pre-extraction is sequentially input into the local and global adaptive feature extraction unit and the simple behavior discrimination unit; the features output by the feature fusion unit are the input sequentially connected initialization unit, the parallel-connected residual multi-parallel multi-attention feature extraction unit, and the residual multi-parallel channel enhancement and feature extraction unit; the local and global adaptive feature extraction unit extracts local and global features by adaptively adjusting the convolution kernel size; the residual multi-parallel multi-attention feature extraction unit adopts a residual multi-parallel structure and uses multiple attention mechanisms to extract multi-dimensional features contained in human behavior data; the residual multi-parallel channel enhancement and feature pre-extraction unit adopts a residual multi-parallel structure and uses channel enhancement technology to extract channel features contained in human behavior data.
[0089] In the embodiment of the present invention, the feature fusion unit effectively fuses long temporal features in complex behaviors by using M-reluGRU and the multi-head self-attention mechanism; M-reluGRU removes the reset gate in the gated recurrent unit GRU, simplifies the GRU into a single-gate structure, uses the ReLU function as the activation function, and uses batch normalization BN as the normalization method.
[0090] In the embodiment of the present invention, batch normalization is used to avoid numerical instability caused by the unboundedness of the ReLU activation function; M-reluGRU can obtain the outputs of each hidden layer at different times with less computational cost, extract the behavioral context information of the previous and subsequent times, and is more suitable for temporal data compared to GRU; the multi-head self-attention mechanism enables the network to focus on information from different feature subspaces, and by adding a multi-head self-attention layer after M-reluGRU, higher weights are given to the single-window feature data that contributes the most and has the most obvious features for identifying complex behaviors; finally, through global average pooling (Global Average Pooling, GAP), the channel features of each data are averaged to achieve multi-dimensional feature fusion.
[0091] In the embodiment of the present invention, the human behavior information application module includes a human behavior visualization unit, a human behavior statistics unit, and a human behavior analysis unit; the human behavior data storage module includes a cloud database and a local database.
[0092] In the embodiment of the present invention, the simple behavior discrimination output unit includes a fully connected layer and a Softmax function connected in sequence. After multi-dimensional feature fusion, the data is output to the fully connected layer and classified through the Softmax function to output the final simple behavior type.
[0093] The recognition results of the simple behavior recognition module and the complex behavior recognition module can be transmitted to each application platform in real time for display and statistics, and the behaviors of users can be analyzed and managed.
[0094] The present invention provides a feasible solution for behavior recognition of multiple complexities based on sensors. Aiming at problems such as incomplete feature extraction of existing models, it enhances the ability of feature representation and the adaptability to capture complex information, and improves the recognition accuracy. Aiming at the defects of existing models in complex behavior recognition, a recognition method for complex human behaviors is proposed, which enhances the practical application performance of existing models.
[0095] The present invention can be applied to scenarios such as smart homes; taking the recognition of daily behavior activities in the application scenario of electric power operation as an example, the human behaviors in the electric power scenario have strong logic and conform to the time sequence; among them, complex operation behaviors are usually composed of multiple consecutive simple behaviors. For example, the behavior of a person cleaning the bedroom is composed of multiple simple behaviors such as walking, bending down, and sweeping the floor. The recognition system first analyzes the short-term features of the collected actions to realize the recognition of simple behaviors, and on this basis, comprehensively analyzes the behavior context information to obtain the long-time sequence features of the behavior and realize the recognition of complex behaviors.
[0096] Figure 5 It is a flowchart of the human behavior recognition method based on multi-dimensional and multi-scale feature extraction provided by the embodiment of the present invention. As Figure 5 shown, this method is implemented based on the above-mentioned human behavior recognition system based on multi-dimensional and multi-scale feature extraction; this method includes:
[0097] S1. Human behavior data collection. The behavior information of the user is collected through the multi-modal sensors in the human behavior data collection module, and the behavior information includes motion data and physiological data;
[0098] S2. Human behavior data transmission. The behavior information collected by the human behavior data collection module is transmitted to the databases of the local server and the cloud server of the human behavior data storage module through the human behavior data transmission module for the human behavior data storage module to store the behavior information;
[0099] S3. Human behavior data preprocessing. The behavior information is sequentially subjected to human behavior data denoising, multi-modal human behavior data merging, human behavior data missing value processing, human behavior data normalization, and human behavior data sliding window segmentation through the human behavior data preprocessing module;
[0100] S4. Construct an LGAM simple human behavior recognition network for simple human behavior recognition and classification; input the behavior data after feature preprocessing into the LGAM simple human behavior recognition network in batches for feature extraction and fusion, and output simple discrimination results through the simple behavior discrimination output unit;
[0101] S5. Construct an MPFEFN complex human behavior recognition network for complex human behavior recognition and classification; input the initialized behavior data into the MPFEFN complex human behavior recognition network in batches for feature extraction and fusion, and output complex discrimination results through the complex behavior discrimination output unit;
[0102] S6. Feature fusion, fuse the output feature data through the M-reluGRU and the multi-head self-attention mechanism and input it into the Generalized Average Pooling (GAP) layer. The GAP layer converts the feature map of each channel into a feature point, and the feature point is the average value of the entire feature map;
[0103] S7. Output the action recognition result. After multi-dimensional spatio-temporal feature extraction and fusion, input the obtained feature data into the fully connected layer. The number of hidden units in the fully connected layer is the number of simple behavior categories included in the data; the data after passing through the fully connected layer passes through the Softmax classifier to calculate the probability of the corresponding behavior, and the behavior type with the highest probability is the final judgment result;
[0104] S8. Display and application of human behavior information. Through the human behavior information application module, visually display, statistically analyze the behavior recognition results.
[0105] In the embodiment of the present invention, step S3 includes:
[0106] S31. Denoise human behavior data; use the soft-threshold wavelet denoising method of Stein's Unbiased Risk Estimation (SURE) to denoise the data collected by the sensor; , is the original signal, is the noisy signal, is the Gaussian white noise, , is the noise intensity. The denoising process is to remove the noise from the signal to obtain the best approximation of the original signal ;
[0107] Perform discrete sampling to obtain point discrete signal , and the wavelet transform coefficient is shown in Equation (1):
[0108] (1);
[0109] Wherein, is the wavelet coefficient, is the scaling function, j is the scale parameter, k is the unit number of the translation of the scaling function. Through the two-scale equations (2) and (3), the recursive implementation method of equation (1) is obtained:
[0110] (2);
[0111] (3);
[0112] Wherein, the symbol * represents convolution, h and g represent the low-pass and high-pass filters respectively, represents the original signal , represents the approximation coefficient at scale j. The wavelet transform reconstruction formula is shown in equation (4):
[0113] (4);
[0114] The soft threshold estimation method based on SURE is adopted to determine the threshold, and the likelihood estimation of the given threshold is obtained, as shown in equation (5):
[0115] (5);
[0116] Wherein, represents the selected initial threshold, represents the wavelet coefficient from the sub-band and represents the sum of the number of wavelet coefficients in each sub-band; minimizing the likelihood function, the required threshold is obtained, as shown in equation (6):
[0117] (6);
[0118] Wherein, represents the obtained threshold parameter;
[0119] The wavelet transform coefficients of the behavior data are processed by the soft threshold function. The absolute value of the data is compared with the threshold. The points less than the threshold are set to zero, and the points greater than or equal to the threshold are shrunk towards zero and become the difference between the point value and the threshold. The soft threshold function is shown in equation (7):
[0120] (7);
[0121] According to equation (3), the wavelet reconstruction of the signal is performed to obtain the denoised signal;
[0122] S32. Multimodal human behavior data merging: Vertically align the behavior data collected by sensors according to timestamps and perform horizontal splicing and merging. The merged data is in the format of a two-dimensional array. The order of horizontal splicing is the X, Y, and Z axis data of the acceleration sensor, the X, Y, and Z axis data of the angular velocity sensor, the X, Y, and Z axes of the magnetometer, heart rate data, blood pressure data, blood oxygen data, and surface electromyography data in sequence.
[0123] S33. Missing value processing of human behavior data: Use the mean imputation method to process the missing values in the behavior data, and fill the missing data with the average value of the column where the missing data is located.
[0124] S34. Normalization of human behavior data: Use the Z-Score normalization method to normalize the human behavior data to ensure that the data is within the same order of magnitude range and avoid the adverse effects of different dimensions and value ranges on the calculation.
[0125] Assume the input data sample sequence is , and the output sequence after Z-Score normalization is , and the calculation method is shown in Equation (8):
[0126] (8);
[0127] Where is the mean of the input data sample sequence, and is the standard deviation of the input data sample sequence;
[0128] S35. Sliding window segmentation of human behavior data: Use a window with a fixed length to segment the continuous sensor data into data segments with a fixed length. Ensure that each data segment contains at least one complete action of a simple behavior during segmentation, and adopt a 50% window overlap rate during the window sliding process.
[0129] In the embodiment of the present invention, the step S4 includes:
[0130] S41: Input human behavior data into a simple human behavior recognition network; Transform the preprocessed data into a shape suitable for the one-dimensional convolutional layer, and input it into the feature extraction and fusion unit in batches. The data shape is , where is the batch size, is the number of data channels, and is the window length;
[0131] The feature data output by the feature pre-extraction unit is input into the local and global adaptive spatio-temporal feature extraction unit and the feature fusion unit connected in sequence;
[0132] S42: Local and global spatio-temporal feature extraction; The data of the input module is processed by an initialization unit, which contains a one-dimensional convolutional module composed of a one-dimensional convolutional layer, a batch normalization layer, an activation layer, and a max pooling layer connected in sequence. Among them, the convolutional kernel size of the one-dimensional convolutional layer is 3, the stride size is 1, the padding method is SAME, and the ReLU function is selected as the non-linear activation function; The calculation method of the one-dimensional convolutional module is shown in Equation (9):
[0133] ;
[0134] Among them, represents the th column of the feature map, represents the th column of the convolutional kernel, represents the bias term;
[0135] The local and global adaptive spatio-temporal feature extraction unit constructs an adaptive temporal kernel composed of a local branch and a global branch. Among them, the local branch adopts a local sensitive modeling scheme, captures short-term temporal dynamic changes by establishing position-sensitive importance weights, and extracts as shown in Equation (10). The global branch is based on the local invariant modeling idea, obtains global information and long-term temporal dependence information by constructing a dynamic aggregation convolutional kernel, and extracts as shown in Equation (11). The overall process is shown in Equation (12):
[0136] (10);
[0137] Among them, represents the local branch transformation, represents the multiplication operator, represents the one-dimensional convolutional operation, and represent the Sigmoid activation function and the ReLU activation function respectively;
[0138] (11);
[0139] Among them represents the global branch transformation, represents the equivalent transformation function of the fully connected (FC) layer operation, represent the weights of two FC layers respectively; In order to realize the interaction between the information obtained by the global branch and the information obtained by the local branch, finally, the outputs of the two branches are convolved;
[0140] (12);
[0141] Among them, Denote the output features of the Temporal Adaptive Module (TAM). Denote the convolution operator.
[0142] In the embodiments of the present invention, step S5 includes:
[0143] Extracting the temporal features of complex action data; sequentially inputting the feature data into an initialization unit, a parallel residual multi-parallel multi-attention feature extraction unit, and a residual multi-parallel channel enhancement and feature extraction unit; further obtaining local features related to body parts through an initialization layer, and extracting as shown in Equation (13):
[0144] (13);
[0145] Where, is the feature map, is the two-dimensional convolution calculation, and are respectively the average value and the maximum value of the pooling operation;
[0146] The residual multi-parallel multi-attention feature extraction unit adopts a special parallel structure. Instead of simply parallelizing a single branch, it reduces the output dimension through specific addition and dimensionality reduction operations, while ensuring that the branches do not interfere with each other. The calculation process is as shown in Equation (14); each branch uses a dynamic convolution based on the modality-temporal attention mechanism to help capture the modality and temporal correlations between features, improving the ability to identify key information in multi-modal data; the calculation process of the dynamic convolution based on the modality-temporal attention mechanism (Modal and Temporal Based Dynamic Convolution, MT-Dyconv) is as shown in Equations (15)-(22):
[0147] (14);
[0148] Where, is the concatenated vector, denotes the concatenation operation, denotes the addition and dimensionality reduction operation, that is, compression after summing in the first dimension, and denote the operation processes of the Basic Convolution Block (BCB) and the Feature Extraction with Multi-attention Module (FEMM); the dynamic convolution process can be expressed as:
[0149] (15);
[0150] Among them, and are the weights and weights of the nth nucleus respectively, represents the input feature map, represents the output feature map;
[0151] Specifically, first, pooling kernels of sizes ( ) and ( ) are used to encode each channel along the horizontal and vertical coordinates respectively; Therefore, the output feature maps after pooling are respectively expressed as:
[0152] (16);
[0153] (17);
[0154] Among them, represents the output of the channel with height , represents the output of the channel with width ;
[0155] In order to make full use of modal information and time information, the above two feature maps are concatenated and then transformed using a shared convolution, expressed as:
[0156] (18);
[0157] Among them, is the internal feature map of spatial information in the horizontal and vertical directions, represents the sigmoid activation function; represents the sampling rate of downsampling, used to control the size of the module;
[0158] Next, along the spatial dimension, the feature map is divided into two independent tensors and ; Two convolutions and with a kernel of size ( ) are used to transform the feature map and into the same number of channels as the input; The calculation process is as follows:
[0159] (19);
[0160] (20);
[0161] Among them, and are the output feature maps after channel transformation, is the sigmoid activation function;
[0162] Then, the output is expanded into attention weights, which can be expressed as:
[0163] (21);
[0164] Where, and represent the expanded attention weights, represents the input, represents the final output;
[0165] After calculating the attention weights, two convolutional layers are set to change the module capacity and computational cost, and the calculated attention scalar is:
[0166] (22);
[0167] Where, is a hyperparameter used to control the sparsity of the weights;
[0168] The residual multi-parallel channel enhancement and feature extraction unit adopts the same parallel structure, reduces the output dimension through specific addition and dimensionality reduction operations, and ensures that the branches do not interfere with each other. The overall operation is as follows:
[0169] (23);
[0170] Where, is the concatenated vector, represents the concatenation operation, represents the addition and dimensionality reduction operation, that is, compression is performed after summing in the first dimension, and represent the operation processes of the BCB and the Channel Boost and Extraction Module (CBEM);
[0171] The Inception structure and the SE module are used to extract channel features. The Inception structure inputs the input vector into four groups of parallel convolutional blocks respectively, and then performs concatenation in the channel dimension, which is expressed as:
[0172] (24);
[0173] Where, The feature represents the enhanced feature, represents the concatenation of the outputs of multiple two-dimensional convolutional layers, , and respectively represent that the convolutional kernel sizes are , and for the two-dimensional convolution operation;
[0174] The SE calculation method is as follows:
[0175] In the compression stage, for the enhanced features of each branch, global average pooling is used to reduce the context and capture the channel dependencies in the global information, expressed as:
[0176] (25);
[0177] Wherein, is the input, is the element capturing the feature;
[0178] The excitation is to make full use of the information captured in the squeezing process. This process uses two FC layers connected by an activation function as a gating mechanism to learn the non-linear interaction between multiple channels. The excitation stage is defined as:
[0179] (26);
[0180] Wherein, is the attention weight, and are the activation functions ReLU and sigmoid respectively, and represents the weight for the transformation scale;
[0181] Then, the scaling operation distributes the previously obtained attention weights to the features of each channel, expressed as:
[0182] (27);
[0183] Wherein, represents the th element of the output feature map.
[0184] In the embodiment of the present invention, step S6 includes:
[0185] The feature fusion unit can assign different weights to the features at different times, and adopts M-reluGRU and the multi-head self-attention mechanism; the feature data of a single window is first input into M-reluGRU for long-time sequence feature fusion, and the calculation process of M-reluGRU is shown in formulas (27)-(29):
[0186] (28);
[0187] (29);
[0188] (30);
[0189] Among them, represents the input data, represents the output of the update gate, represents the output of the M-reluGRU at the previous moment, represents the candidate hidden state, represents the output of the M-reluGRU at the current moment;
[0190] The data output by the M-reluGRU is input into the multi-head attention layer, which assigns different weights to the single-window simple behavior feature data input at different times; the feature matrix composed of several single-window feature data is , where, represents the single-window feature data output by the simple behavior feature acquisition unit at the th moment, represents the total number of windows; multiplying by the corresponding weight matrix respectively obtains the query matrix , the key matrix and the value matrix , and then repeating multiple times to perform different linear mappings on , , to calculate the outputs of different attention heads, and finally concatenating the outputs of multiple attentions; the calculation method of the multi-head attention layer is shown in Equations (30)-(32):
[0191] (31);
[0192] (32);
[0193] (33);
[0194] where, represents the square root of the dimensions of matrix and matrix , , and respectively represent the weight matrices for the th time to perform linear mappings on , , , represents the calculation result of the th head in the multi-head attention mechanism, It means concatenating the outputs of multiple heads; finally, the preliminarily fused features are input into the GAP layer, which converts the feature map of each channel into a feature point, and this feature point is the average value of the entire feature map.
[0195] Through the acquisition and analysis of sensor data, the present invention realizes human behavior classification and recognition based on sensors, making up for the deficiencies of high cost, susceptibility to interference, and poor privacy of vision-based behavior recognition; through a variety of attention mechanisms and residual parallel structures, the effectiveness, comprehensiveness, and complex dynamic capture ability of feature extraction are effectively improved. Compared with mainstream and latest models, it has obvious advantages in terms of adaptability, reliability, and practicality; the feature extraction and fusion module proposed by the present invention makes up for the deficiencies of existing models in complex human behavior recognition and can effectively solve the problem that existing models can only achieve simple behavior recognition.
[0196] In the technical solution provided by the present invention, the system includes a human behavior data acquisition module, a human behavior data transmission module, a human behavior data storage module, a human behavior data preprocessing module, an LGAM simple human behavior recognition network module, an MPFEFN complex human behavior recognition network module, and a human behavior information application module; the LGAM simple human behavior recognition network module includes a feature pre-extraction unit, a local and global adaptive spatio-temporal feature extraction unit, a feature fusion unit, and a simple behavior discrimination output unit connected in sequence; the MPFEFN complex behavior recognition network module includes an initialization unit, a residual multi-parallel multi-attention feature extraction unit and a residual multi-parallel channel enhancement and extraction unit connected in parallel, a feature fusion unit, and a complex behavior discrimination output unit. The system makes better use of the multi-dimensional features contained in human behavior data, improves the recognition accuracy of easily confused behaviors, and realizes the effective recognition of complex human behaviors, enhancing the recognition performance of existing models.
[0197] Each step of the embodiments of the present invention can be executed by an electronic device. Among them, the electronic device includes, but is not limited to, mobile phones, tablet computers, portable PCs, desktop computers, etc.
[0198] The embodiments of the present invention provide a computer-readable storage medium, which includes a stored program. When the program runs, it controls the electronic device where the computer-readable storage medium is located to execute the embodiments of the above-mentioned human behavior recognition method based on multi-dimensional and multi-scale feature extraction.
[0199] Figure 6 It is a schematic diagram of an electronic device provided by the embodiments of the present invention, as Figure 6As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, it implements the human behavior recognition method based on multi-dimensional and multi-scale feature extraction in the embodiments. To avoid repetition, details are not elaborated here.
[0200] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art can understand that Figure 6 These are merely examples of the electronic device 21 and do not constitute a limitation on the electronic device 21. It may include more or fewer components than shown in the figure, or combine certain components, or have different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0201] The so-called processor 211 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0202] The memory 212 may be an internal storage unit of the electronic device 21, such as the hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 21. Further, the memory 212 may also include both the internal storage unit and the external storage device of the electronic device 21. The memory 212 is used to store the computer program and other programs and data required by the network device. The memory 212 may also be used to temporarily store data that has been output or is to be output.
[0203] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0204] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A human behavior recognition system based on multi-dimensional and multi-scale feature extraction, characterized in that The system includes: a human behavior data acquisition module, a human behavior data transmission module, a human behavior data storage module, a human behavior data preprocessing module, an LGAM simple human behavior recognition network module, an MPFEFN complex human behavior recognition network module, and a human behavior information application module; The LGAM simple human behavior recognition network module includes a feature pre-extraction unit, a local and global adaptive spatio-temporal feature extraction unit, a feature fusion unit, and a simple behavior discrimination output unit connected in sequence; The MPFEFN complex behavior recognition network module includes an initialization unit, a residual multi-parallel multi-attention feature extraction unit and a residual multi-parallel channel enhancement and extraction unit connected in parallel, a feature fusion unit, and a complex behavior discrimination output unit; The feature pre-extraction unit uses a single one-dimensional convolution module to preliminarily extract behavior features; the data after feature pre-extraction is sequentially input into the local and global adaptive feature extraction unit and the simple behavior discrimination unit; the features output by the feature fusion unit are the input connected in sequence to the initialization unit, the residual multi-parallel multi-attention feature extraction unit and the residual multi-parallel channel enhancement and feature extraction unit connected in parallel; the local and global adaptive feature extraction unit extracts local and global features by adaptively adjusting the convolution kernel size; the residual multi-parallel multi-attention feature extraction unit adopts a residual multi-parallel structure and uses multiple attention mechanisms to extract multi-dimensional features contained in human behavior data; the residual multi-parallel channel enhancement and feature pre-extraction unit adopts a residual multi-parallel structure and uses channel enhancement technology to extract channel features contained in human behavior data; The feature fusion unit uses M-reluGRU and a multi-head self-attention mechanism to effectively fuse long temporal features in complex behaviors; M-reluGRU removes the reset gate in the gated recurrent unit GRU, simplifies the GRU to a single-gate structure, uses the ReLU function as the activation function, and uses BN as the normalization method; Complex action data temporal feature extraction; the feature data is sequentially input into the initialization unit, the parallel residual multi-parallel multi-attention feature extraction unit and the residual multi-parallel channel enhancement and feature extraction unit; local features related to body parts are further obtained through the initialization layer, and the extraction is as shown in Equation (13): (13); Among them, is the feature map, is the two-dimensional convolution calculation, and are the average value and the maximum value of the pooling operation respectively; The residual multi-parallel multi-attention feature extraction unit adopts a parallel structure. Instead of simply parallelizing a single branch, it reduces the output dimension through addition and dimensionality reduction operations while ensuring that the branches do not interfere with each other. The calculation process is as shown in Equation (14); each branch uses dynamic convolution based on the modality-time attention mechanism to help capture the modality and time correlations between features, improving the ability to identify key information in multi-modal data; the calculation process of the dynamic convolution based on the modality-time attention mechanism (Modal and Temporal Based Dynamic Convolution, MT-Dyconv) is as shown in Equations (15)-(22): (14); Among them, is the splicing vector, represents the splicing operation, represents the summation and dimensionality reduction operation, that is, compression is performed after summing in the first dimension, and represent the operation processes of the Basic Convolution Block (BCB) and the Feature Extraction with Multi-attention Module (FEMM); the dynamic convolution process can be expressed as: (15); Among them, and are the weights and weights of the kernel respectively, represents the input feature map, represents the output feature map; Specifically, first, pooling kernels of size ( ) and ( ) are used to encode each channel along the horizontal and vertical coordinates, respectively; thus, the output feature maps after pooling are represented as follows: (16); (17); Among them, represents the channel output with a height of , represents the channel output with a width of . To make full use of modal information and temporal information, the above two feature mappings are cascaded and then transformed using shared convolutions, which is expressed as: (18); Among them, is the internal feature map of the spatial information in the horizontal and vertical directions, represents the sigmoid activation function; represents the sampling rate of downsampling, which is used to control the size of the module; Next, along the spatial dimension, the feature map is divided into two independent tensors and ; using two convolutions and with a kernel of size ( ), the feature maps are transformed to the same number of channels as the input; the calculation process is as follows: (19); (20); Among them, and are the output feature maps after channel transformation, is the sigmoid activation function; Then, the output is expanded into attention weights, which can be expressed as: (21); Among them, and represent the extended attention weights, represents the input, represents the final output; After calculating the attention weights, two convolutional layers are set to change the module capacity and computational cost, and the calculated attention scalar is: (22); Among them, is a hyperparameter for controlling the sparsity of the weights; The residual multi-parallel channel enhancement and feature extraction unit adopts the same parallel structure, reduces the output dimension through addition and dimensionality reduction operations, and at the same time ensures that the branches do not interfere with each other. The overall operation is as follows: (23); Among them, is a splicing vector, represents a splicing operation, represents a summation and dimensionality reduction operation, that is, compression is performed after summation in the first dimension, and represent the operation process of the BCB and the Channel Boost and Extraction Module (CBEM); The Inception structure and SE module are used to extract channel features. The Inception structure inputs the input vector into four groups of parallel convolutional blocks respectively, and then concatenates them in the channel dimension, which is expressed as: (24); Among them, The feature representation represents the enhanced feature, indicating the concatenation of the outputs of multiple two-dimensional convolutional layers, , and respectively represent two-dimensional convolutional operations with convolutional kernel sizes of , and ; The SE calculation method is as follows: In the compression stage, for the enhanced features of each branch, global average pooling is used to reduce the context and capture the channel dependencies in the global information, which is expressed as: (25); Among them, is the input, is the element that captures the feature; The excitation is to make full use of the information captured during the squeezing process. This process uses two FC layers connected by activation functions as a gating mechanism to learn the non-linear interactions between multiple channels. The excitation stage is defined as: (26); Among them, is the attention weight, and are the activation functions ReLU and sigmoid respectively, and represent the weights for transforming the scale; Then, the scaling operation distributes the previously obtained attention weights to the features of each channel, which is expressed as: (27); Among them, represents the -th element of the output feature map.
2. The system according to claim 1, wherein The human behavior data acquisition module includes several motion data sensing units and physiological data sensing units of different types. The motion data sensing units include an acceleration data sensing unit, an angular velocity data sensing unit, and a magnetic induction intensity data sensing unit; the physiological data sensing units include a heart rate data sensing unit, a blood pressure data sensing unit, a blood oxygen data sensing unit, and a surface electromyography data sensing unit; The transmission methods of the human behavior data transmission module include ultra-wideband UWB, Wi-Fi, long-range radio Lora, ZigBee, Bluetooth, 4G, 5G, and long-range radio information transmission; The human behavior data preprocessing module includes a data denoising unit, a multi-modal data merging unit, a missing value processing unit, a data normalization unit, and a data sliding window segmentation unit connected in sequence.
3. The system according to claim 2, wherein The multi-modal data merging unit merges the motion data and physiological data collected by different sensors in a way of aligning the longitudinal timestamps and horizontally arranging and splicing; the missing value processing unit uses the mean imputation method to complete the missing information, taking the average value of the column where the missing data is located as the missing value; the data normalization unit performs Z-Score normalization processing on data with different dimensions and value ranges to convert them to the same value range; the data sliding window segmentation unit divides the continuous time series collected by the sensor into multiple data segments.
4. The system according to claim 1, wherein The human behavior information application module includes a human behavior visualization unit, a human behavior statistics unit, and a human behavior analysis unit; the human behavior data storage module includes a cloud database and a local database.
5. The system according to claim 1, characterized in that, The simple behavior discrimination output unit includes a fully connected layer and a Softmax function connected in sequence. After multi-dimensional feature fusion, the data is output to the fully connected layer and classified through the Softmax function to output the final simple behavior type.
6. A human behavior recognition method based on multi-dimensional and multi-scale feature extraction, characterized in that, The method is implemented based on the human behavior recognition system for multi-dimensional and multi-scale feature extraction according to any one of claims 1 to 5; the method includes: S1. Human behavior data collection: Collect the behavior information of the user through the multi-modal sensors in the human behavior data collection module, where the behavior information includes motion data and physiological data; S2. Human behavior data transmission: Transmit the behavior information collected by the human behavior data collection module to the databases of the local server and the cloud server in the human behavior data storage module through the human behavior data transmission module for the human behavior data storage module to store the behavior information; S3. Human behavior data preprocessing: Sequentially perform human behavior data denoising, multi-modal human behavior data merging, human behavior data missing value processing, human behavior data normalization, and human behavior data sliding window segmentation on the behavior information through the human behavior data preprocessing module; S4. Construct an LGAM simple human behavior recognition network for simple human behavior recognition and classification; Batch the behavior data after feature preprocessing and input it into the LGAM simple human behavior recognition network for feature extraction and fusion, and output a simple discrimination result through the simple behavior discrimination output unit; Perform discrete sampling to obtain a point discrete signal, and the wavelet transform coefficient is shown in Equation (1): S5. Construct an MPFEFN complex human behavior recognition network for complex human behavior recognition and classification; Batch the initialized behavior data and input it into the MPFEFN complex human behavior recognition network for feature extraction and fusion, and output a complex discrimination result through the complex behavior discrimination output unit; S6. Feature fusion: Fuse the output feature data through the M-reluGRU and the multi-head self-attention mechanism and input it into the Generalized Average Pooling (GAP) layer. The GAP layer converts the feature map of each channel into a feature point, and the feature point is the average value of the entire feature map; S7. Output the action recognition result: Input the obtained feature data into the fully connected layer after multi-dimensional spatio-temporal feature extraction and fusion. The number of hidden units in the fully connected layer is the number of simple behavior categories included in the data; The data after passing through the fully connected layer passes through the Softmax classifier to calculate the probability of the corresponding behavior, and the behavior type with the highest probability is the final judgment result; S8. Display and application of human behavior information: Visualize, statistically analyze the behavior recognition result through the human behavior information application module.
7. The method according to claim 6, characterized in that The step S3 includes: S31. Denoising of human body behavior data: The soft-threshold wavelet denoising method based on Stein's unbiased risk estimator is used to denoise the data collected by the sensor; , is the original signal, is the noisy signal, is the Gaussian white noise, , is the noise intensity. The denoising process is to remove the noise from the signal to obtain the best approximation of the original signal ; Discrete sampling is performed to obtain a discrete signal of points , and the wavelet transform coefficients are as shown in Equation (1): (1); Among them, is the wavelet coefficient, is the scaling function, j is the scaling parameter, k is the number of units of translation of the scaling function. Through the two-scale equations (2) and (3), the recursive implementation method of equation (1) is obtained: (2); (3); where the symbol * represents convolution, h and g represent the low-pass and high-pass filters respectively, represents the original signal , represents the approximation coefficient at the j-th scale, and the wavelet transform reconstruction formula is shown in Equation (4): (4); Use the soft threshold estimation method based on Stein's Unbiased Risk Estimate (SURE) to determine the threshold and obtain the likelihood estimate of the given threshold, as shown in Equation (5): (5); Among them, represents the selected initial threshold value, represents the wavelet coefficients from the sub-band , represents the sum of the numbers of wavelet coefficients in each sub-band; minimizing the likelihood function gives the required threshold value, as shown in Equation (6): (6); Among them, The expressed threshold parameter; Process the wavelet transform coefficient of the behavior data using the soft threshold function. Compare the absolute value of the data with the threshold. Points less than the threshold are set to zero, and points greater than or equal to the threshold are shrunk towards zero to become the difference between the point value and the threshold. The soft threshold function is shown in Equation (7): (7); Perform wavelet reconstruction of the signal according to Equation (3) to obtain the denoised signal; S32. Merge multi-modal human behavior data: Vertically align the behavior data collected by the sensor according to the timestamp and perform horizontal splicing and merging. The merged data is in the format of a two-dimensional array; S33. Process missing values in human behavior data: Use the mean imputation method to process the missing values in the behavior data, and fill the missing data with the average value of the column where the missing data is located; S34. Normalize human behavior data: Use the Z-Score normalization method to normalize the human behavior data to ensure that the data is within the same order of magnitude range; Suppose the input data sample sequence is , and the output sequence after Z-Score normalization is , and the calculation method is shown in Equation (8): (8); Among them, is the mean of the input data sample sequence, is the standard deviation of the input data sample sequence; S35. Segment human behavior data with a sliding window: Use a window with a fixed length to segment the continuous sensor data into data segments with a fixed length. When segmenting, ensure that each data segment contains at least one complete action of a simple behavior, and adopt a 50% window overlap rate during the window sliding process.
8. The method according to claim 7, wherein The step S4 includes: S41: Input human behavior data into a simple human behavior recognition network; transform the preprocessed data into a shape suitable for the one-dimensional convolutional layer, and input it into the feature extraction and fusion unit in batches. The data shape is , where is the batch size, is the number of data channels, is the window length; The feature data output by the feature pre-extraction unit is input into the locally and globally adaptive spatio-temporal feature extraction unit and the feature fusion unit connected in sequence; S42: Locally and globally spatio-temporal feature extraction: The data input by the input module is processed by the initialization unit. The initialization unit includes a one-dimensional convolution module, which is composed of a one-dimensional convolutional layer, a batch normalization layer, an activation layer, and a max pooling layer connected in sequence. Among them, the convolutional kernel size of the one-dimensional convolutional layer is 3, the stride size is 1, the padding method is SAME, and the ReLU function is selected as the non-linear activation function; the calculation method of the one-dimensional convolution module is shown in Equation (9): (9); Among them, represents the column of the feature map, represents the column of the convolutional kernel, represents the bias term; The locally and globally adaptive spatio-temporal feature extraction unit constructs an adaptive temporal kernel composed of a local branch and a global branch. The local branch adopts a locally sensitive modeling scheme, captures short-term temporal dynamic changes by establishing position-sensitive importance weights, and extracts as shown in Equation (10). The global branch is based on the idea of local invariance modeling, obtains global information and long-term temporal dependence information by constructing a dynamic aggregation convolutional kernel, and extracts as shown in Equation (11). The overall process is shown in Equation (12): (10); Among them, represents a local branch transformation, represents a multiplication operator, represents a one-dimensional convolution operation, and represent the Sigmoid activation function and the ReLU activation function respectively; (11); Among them represents a global branch transformation represents a fully connected layer operation equivalence transformation function respectively represent the weights of two FC layers; in order to realize the interaction between the information obtained by the global branch and the information obtained by the local branch, finally, the outputs of the two branches are convolved (12); Among them, represents the output features of the time adaptive module, represents the convolution operator.
Citation Information
Patent Citations
Pedestrian re-identification method based on residual multi-channel attention multi-feature fusion
CN115830531A
Human body behavior recognition system based on multi-dimensional feature fusion and working method thereof
CN116092119A