Human body emotion recognition system and method based on double-flow residual space-time fusion model

Through the human emotion recognition system of the dual-stream residual spatiotemporal fusion model, wearable devices are used to collect and process peripheral physiological signals, realizing real-time monitoring and accurate recognition of human emotions, solving the problems of insufficient feature extraction and insufficient multimodal signal fusion in the existing technology, and improving the accuracy of emotion recognition.

CN120744600APending Publication Date: 2025-10-03SHANDONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510768671.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing human emotion recognition technology has problems such as cumbersome steps, insufficient accuracy and complex equipment, especially insufficient extraction of peripheral physiological signal feature information based on wearable devices and defects in multimodal signal fusion, and existing public data sets are not collected for specific fields.

Method used

A human emotion recognition system based on a dual-stream residual spatiotemporal fusion model is adopted. Peripheral physiological signals are collected through wearable devices. After data preprocessing, feature extraction and fusion are performed using a multi-scale separable residual module, a time domain convolution feature module and a dual-stream cross-attention fusion module. Finally, the emotional state is obtained at the classification and recognition layer.

Benefits of technology

It realizes real-time monitoring of human emotions, improves the accuracy of emotion recognition, solves the problems of multi-scale feature extraction and multi-dimensional feature fusion, and enhances feature representation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744600A_ABST
    Figure CN120744600A_ABST
Patent Text Reader

Abstract

The invention discloses a human body emotion recognition system and method based on a double-flow residual space-time fusion model, and relates to the technical field of artificial intelligence. Comprising a human body peripheral physiological information acquisition module, a human body peripheral physiological information transmission module, a human body peripheral physiological information preprocessing module, a human body emotion classification and recognition module and a human body emotion information application module which are connected in sequence, the human body peripheral physiological information acquisition module is used for acquiring peripheral physiological data of a user in real time, wherein the peripheral physiological data comprises photoelectric volume pulse waves, skin electrical response and temperature; the human body peripheral physiological information transmission module is used for transmitting the collected peripheral physiological data of the user to a local server or a cloud server; the human body peripheral physiological information preprocessing module is used for sequentially carrying out storage, extreme value removal, filtering, interpolation, normalization, sliding window segmentation and label calibration on the collected peripheral physiological data of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a human emotion recognition system and method based on a dual-stream residual spatiotemporal fusion model. Background Art

[0002] Human emotion recognition uses information from questionnaires, physical signals, and physiological signals to accurately infer human emotional states from a variety of sources and patterns, enabling the classification and identification of emotions. With the continuous development of artificial intelligence technology, research in the field of emotion recognition using peripheral physiological signals collected by wearable devices has received increasing attention. Human emotion recognition has real-world applications, such as the development of human-computer interaction, smart classrooms, and healthcare systems.

[0003] Research on human emotion recognition has made some progress. Current mainstream methods include questionnaires, physical signal monitoring (e.g., audio and video), and physiological signal monitoring. However, researchers both domestically and internationally face numerous challenges. First, questionnaire-based emotion recognition requires specialized psychologists to issue questions, which are then collected and completed by individuals. The responses are then statistically scored and analyzed by specialized psychologists. Therefore, due to the cumbersome process, it is not suitable for real-time monitoring. Second, physical signal monitoring (e.g., audio and video) involves real-time video or audio monitoring of subjects. However, people can disguise their internal emotional changes, significantly affecting recognition accuracy. Furthermore, current devices for emotion recognition based on EEG signals are complex and not portable. Therefore, research on emotion recognition using peripheral physiological signals collected by wearable devices holds great potential.

[0004] In recent years, with the continuous advancement of artificial intelligence, deep learning technologies have been widely used in emotion recognition. These methods can extract characteristic information from peripheral physiological signals, opening new doors in the field of emotion recognition. However, existing deep learning-based network models are not sufficient for extracting the complex features of peripheral physiological signals, and have certain limitations in fusing signals from different modalities. In addition, existing public datasets are not collected for specific fields. Summary of the Invention

[0005] The technical problem addressed by this invention is to provide a human emotion recognition system and method based on a dual-stream residual spatiotemporal fusion model. First, a user wears a wearable device for data collection, which contains various built-in sensors and is placed on a specific part of the body. The collected data is then wirelessly transmitted to a local server or cloud server for storage via a data transmission module. The transmission method can be selected based on different usage scenarios. The data is then preprocessed, including extreme value removal, filtering, interpolation, normalization, and sliding window segmentation. Extreme value removal mitigates the impact of extreme values ​​on the data; filtering removes noise and motion artifacts from the raw sensor signals; interpolation addresses the issue of inconsistent sampling rates in peripheral physiological signals; normalization transforms data with different dimensions and value ranges to the same range, facilitating distance-based computational models to eliminate the effects of these differences; and the sliding window segmentation method segments long, continuous data into data segments that conform to the input format of the discriminant model. The preprocessed data is fed into the human emotion classification and recognition module. After processing by the multi-scale separable residual module, the temporal convolutional feature module, and the two-stream cross-attention mechanism module, fully extracted and fused feature information is obtained. Finally, this information is fed into the classification and recognition layer to obtain human emotion state information. This system can monitor the user's emotional state in real time and can be used for emotion monitoring and management of students, as well as for monitoring and emotion management of the elderly or patients. This system addresses the inability of existing algorithms to effectively extract multi-scale features and fuse multi-dimensional features. It proposes a two-stream residual spatiotemporal fusion feature extraction and fusion method to address this issue, further enhancing feature representation and improving emotion recognition accuracy.

[0006] The present invention adopts the following technical solutions to achieve the invention objectives: The human emotion recognition system and method based on the dual-stream residual spatiotemporal fusion model is characterized by comprising: a human peripheral physiological information acquisition module, a human peripheral physiological information transmission module, a human peripheral physiological information preprocessing module, a human emotion classification and recognition module, and a human emotion information application module connected in sequence; The human peripheral physiological information acquisition module is used to: collect the user's peripheral physiological data in real time, the peripheral physiological data including photoplethysmography, skin galvanic response and temperature; The human peripheral physiological information transmission module is used to: transmit the collected user's peripheral physiological data to a local server or a cloud server; The human peripheral physiological information preprocessing module is used to sequentially store, remove extreme values, filter, interpolate, normalize, perform sliding window segmentation, and label calibration on the collected user's peripheral physiological data, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments; The human emotion classification and recognition module is used to: input the user's physiological data preprocessed by the human peripheral physiological information preprocessing module into the trained human emotion classification and recognition module to perform emotion type discrimination; The human emotion information application module is used to transmit the obtained human emotion classification and recognition results to various application platforms, thereby realizing corresponding functions.

[0007] As a further limitation of the present technical solution, the human emotion classification and recognition module includes a parallel-connected multi-scale separable residual module and a time-domain convolution feature module, a dual-stream cross-attention fusion module that fuses parallel connection features, and an emotion category output module; The multi-scale separable residual module includes a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module connected in parallel; The dual-stream cross-attention fusion module includes a feature injector and a feature extractor; The emotion category output module is used to output the classification recognition result. The emotion category output module is composed of a fully connected layer. The fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.

[0008] As a further limitation of the present technical solution, the human peripheral physiological information acquisition module includes several different sensor units, including a photoplethysmography sensor, a skin galvanic response sensor, and a temperature sensor, for acquiring the user's peripheral physiological data.

[0009] As a further limitation of the present technical solution, the human peripheral physiological information preprocessing module includes a sensor data storage unit, an extreme value removal data unit, a photoplethysmography filtering unit, a sensor data interpolation unit, a data normalization unit, a sliding window segmentation unit, and a data label calibration unit connected in sequence.

[0010] As a further limitation of the present technical solution, the human peripheral physiological information transmission module transmits the collected peripheral physiological data of the user to a local server or a cloud server via any transmission method including ultra-wideband, Bluetooth, and 5G.

[0011] The recognition method of the human emotion recognition system based on the dual-stream residual spatiotemporal fusion model includes the following steps: S1: Human peripheral physiological data collection; Real-time collection of user's peripheral physiological data, i.e. human peripheral physiological data, including photoplethysmography, galvanic skin response, and temperature; S2: Human peripheral physiological data transmission; Transmitting the collected user's peripheral physiological data to a local server or a cloud server; S3: Preprocessing of human peripheral physiological data; The collected user's peripheral physiological data is sequentially stored, de-extremely analyzed, filtered, interpolated, normalized, segmented using a sliding window, and labeled, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments. S4: Building recognition model and sentiment classification and recognition; The peripheral physiological data segments pre-processed by S3 are input into the recognition model in batches, and the emotion recognition output is obtained after training; S5: Loss function calculation update and feedback; The feature vectors output by the previous Dense fully connected layer in the fully connected layer of the sentiment classification and recognition output module in S4 are processed according to the principle of the same category, and the central mean and variance of the feature vectors belonging to the same category and finally classified correctly are calculated, as shown in Formula (1) and Formula (2) respectively. In addition, the variance is normalized, as shown in Formula (3): (1); (2); (3); in: Representation category The central mean vector of Representation category The variance of Represents the normalized category The variance of Representation category The number of correctly classified feature vectors in Representation category Middle The feature vector of the data output in the second to last layer of the fully connected layer, Indicates the The square of the Euclidean distance between a sample and the center of its class, K Indicates the total number of classes; is the total number of all known behavioral categories; The loss function is set as formula (4): (4); in: It represents the maximum distance between the center means of two categories, calculated using the 1-norm; S6: Application of discriminant output results; The results of emotion recognition output are transmitted to the corresponding application platform in real time.

[0012] As a further limitation of this technical solution, the specific implementation process of S3 is as follows: S31: Remove extreme values ​​from human peripheral physiological data; The collected peripheral physiological data of the users were de-extremely analyzed. The Windsor method was used to replace the upper and lower extreme values ​​of the three peripheral physiological signals with the 3% and 97% quantiles. S32: filtering of human peripheral physiological data; Human peripheral physiological data filtering refers to the filtering of photoelectric volumetric pulse wave signals with a sampling rate greater than 10 Hz; a 3rd-order Butterworth low-pass filter with a cutoff frequency of 10 Hz is used to remove high-frequency noise and motion artifacts and extract pure human photoelectric volumetric pulse wave signals. The 3rd-order Butterworth filter is expressed as formula (5): (5); in: is the cutoff frequency, is the passband edge frequency, is the angular frequency, is the passband ripple factor, yes The value at the edge of the passband; S33: Interpolation of human peripheral physiological data; Since the collected peripheral physiological signals have inconsistent sampling rates, the data cannot be directly input into the model for training. Therefore, the data is interpolated using a combination of linear interpolation and nearest neighbor interpolation. The processes of the two interpolation methods are shown in Equations (6) and (7), respectively: (6); (7); in: represents a known point in time, Indicates that it corresponds to The known value of represents the time point of target interpolation, Represents the result after using linear interpolation, Represents the result of using nearest neighbor interpolation; S34: Normalization of human peripheral physiological data; Use the minimum-maximum normalization method for normalization. Suppose the input sequence is , the output sequence after minimum-maximum normalization is , as shown in formula (8): (8); in: The first samples, and are the minimum and maximum values ​​of the input sequence, respectively. is the normalized data; S35: Sliding window segmentation and labeling; Sliding window segmentation uses a fixed-length window to segment continuous sensor data into fixed-length data segments, and assigns a corresponding label to each data segment, that is, the name of the emotion category corresponding to the data segment.

[0013] As a further limitation of this technical solution, the S4 includes the following steps: S41: The multi-scale separable residual module uses a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module to output feature data of different scales respectively, specifically: A combination of a convolution kernel size of 3 and a number of multi-scale separable residual blocks of [1, 2, 1] is used to extract fine-grained multi-scale separable residual features; A convolution kernel size of 5 and the number of multi-scale separable residual blocks of [2, 1, 2] are used to perform coarse-grained multi-scale separable residual feature extraction; In addition, the core component of the multi-scale separable residual module is the multi-scale separable residual block, which is composed of a grouped separable convolution module and a residual connection. Its calculation process is shown in formula (9): (9); in: is the output of the grouped separable convolution module, BN is the batch normalization operation, Conv2D represents the two-dimensional convolution operation, Act represents the LeakyReLU activation function used, and Sum represents the The output of Y represents the final output of the multi-scale separable residual block; The group separable convolution module is an adaptive module that can perform adaptive group convolution according to different data inputs. Its calculation process is shown in formula (10): (10); Where: G1, G2, ..., G n Representative n Group convolution, each group is composed of a depth-wise separable convolution and Conv2D, with BN and LeakyReLU added between the two convolutions. Exp represents the channel expansion operation, and Concat represents the connection operation. Then, the feature addition operation is used to further fuse the features at different scales from the outputs of the fine-grained multi-scale separable residual module and the coarse-grained multi-scale separable residual module; S42: The time domain convolution feature module extracts deep time domain features from the pre-processed user's peripheral physiological data. Specifically, In the time-domain convolution feature module, the query vector and the key vector are regarded as one branch, namely the branch containing the time-domain convolution multilayer perceptron module, and the other branch is regarded as the key vector. Similarly, the attention weights and values ​​in the attention mechanism are calculated respectively, and then the feature enhancement is achieved through the dot product operation. The calculation process is shown in Equations (11), (12) and (13): (11); (12); (13); in: x represents the original input feature sequence, Y 1 is the output of TCM, Y 2 means Y The output after the dot product operation of 1 and another branch, is the output of the time-domain convolution feature module, LN is layer normalization, Conv1D is a 1D convolution operation, Act is the GELU activation function, and TCM is a time-domain convolution multilayer perceptron module. The calculation process is shown in Equations (14), (15), and (16): (14); (15); (16); Among them: Drop is the DropOut operation with a drop rate of 0.5, DSConv1D is a one-dimensional depth-separable convolution, is the output of TCM; S43: The dual-stream cross-attention fusion module deeply fuses feature data of different dimensions, specifically: The two-stream cross-attention fusion module includes a feature injector and a feature extractor; The feature injector and feature extractor are two complementary feature processing operations. The feature injector converts the feature information output by the time domain convolution feature module into Injected into the output of the multi-scale separable residual module The calculation process is shown in formula (17): (17); in: is the injection intensity, attn is the efficient cross-attention mechanism, is the output of the feature injector; The feature extractor is used to extract the feature information output by the multi-scale separable residual module , to enhance the feature information output by the time domain convolution feature module , and its calculation process is shown in formula (18): (18); in: is the output of the feature extractor. FFN is a multi-layer perceptron using the GELU activation function. S44: Classification and recognition is to output the emotion classification and recognition results, specifically: The emotion category output module is composed of a fully connected layer, and the fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention proposes a human emotion recognition system based on a dual-stream residual spatiotemporal fusion model. By processing and analyzing peripheral physiological signals, it realizes real-time classification of human emotions, making up for the deficiency that physical signals such as audio and video are easily affected by personal subjective influences.

[0015] In the human emotion information classification module, this paper proposes a new deep learning algorithm based on a dual-stream residual spatiotemporal fusion model to learn the complex deep features of peripheral physiological data at different scales and dimensions, more accurately extract the features of peripheral physiological data, and improve the accuracy of emotion recognition.

[0016] The present invention proposes a human emotion recognition method, which can solve the problem of poor recognition performance of existing classification models by setting a suitable loss function and adjusting model parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the module composition and connection relationship of the present invention.

[0018] Figure 2 Schematic diagram of the method of the present invention.

[0019] Figure 3 Schematic diagram of the principle of classification of human peripheral physiological information according to the present invention. DETAILED DESCRIPTION

[0020] A specific embodiment of the present invention is described in detail below with reference to the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.

[0021] Example 1

[0022] Many challenges arise in classroom management, such as students becoming bored with the material or feeling frustrated by a lack of understanding. While most classrooms now have cameras and teachers present to impart knowledge and monitor student well-being, it can still be difficult for teachers to detect emotional changes in a crowded classroom. Furthermore, video footage captured by cameras can be easily camouflaged by students. Therefore, human emotion recognition technology based on peripheral physiological signals can overcome the limitations of audio and video. Students wear a wristband on their left wrist, equipped with a photoplethysmography sensor, a galvanic skin response sensor, and a temperature sensor. This wristband collects students' peripheral physiological data in real time and transmits it to a local or cloud server for storage and computation, generating corresponding recognition results. The student management application platform displays these results in real time and provides timely guidance to students exhibiting negative emotions. Furthermore, it can also manage students' mental health over the long term, preventing mental health issues.

[0023] The present invention is applied to student management application platform, such as Figure 1 As shown, it includes: a human peripheral physiological information acquisition module, a human peripheral physiological information transmission module, a human peripheral physiological information preprocessing module, a human emotion classification and recognition module, and a human emotion information application module connected in sequence.

[0024] The human peripheral physiological information acquisition module is used to collect the user's peripheral physiological data in real time. The peripheral physiological data includes photoplethysmographic (PPG), galvanic skin response (GSR) and temperature (TEMP).

[0025] The human peripheral physiological information acquisition module includes several different sensor units, which are integrated into a wearable device and worn on the corresponding position of the user's body. The sensor units include a photoplethysmography sensor, a skin galvanic response sensor, and a temperature sensor to collect the user's peripheral physiological data.

[0026] The human peripheral physiological information transmission module is configured to select an appropriate transmission method based on the user's scenario to transmit the collected user's peripheral physiological data to a local server or cloud server. The collected user's peripheral physiological data is transmitted to the local server or cloud server via any of the following transmission methods: Ultra Wide Band (UWB), Bluetooth, or 5G. Different transmission methods vary in power consumption, transmission distance, and transmission speed, and are therefore suitable for different application scenarios.

[0027] The human peripheral physiological information preprocessing module is used to sequentially store, remove extreme values, filter, interpolate, normalize, perform sliding window segmentation, and label calibration on the collected user's peripheral physiological data, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments.

[0028] The human emotion classification and recognition module is used to input the user's physiological data preprocessed by the human peripheral physiological information preprocessing module into the trained human emotion classification and recognition module to perform emotion type discrimination.

[0029] The human emotion information application module is used to transmit the obtained human emotion classification and recognition results to various application platforms, thereby realizing corresponding functions.

[0030] Example 2 The human emotion classification and recognition module includes a parallel-connected multi-scale separable residual module and a time-domain convolution feature module, a dual-stream cross-attention fusion module that fuses parallel connection features, and an emotion category output module.

[0031] The multi-scale separable residual module is used to extract deep spatial features from the preprocessed user's peripheral physiological data. The multi-scale separable residual module includes a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module connected in parallel. The fine-grained multi-scale separable residual module is a branch using a convolution kernel size of 3 and a combination of multi-scale separable residual blocks of [1, 2, 1]. This branch focuses on extracting fine features from peripheral physiological signals to capture short-term emotional changes. The coarse-grained multi-scale separable residual module is a branch using a convolution kernel size of 5 and a combination of multi-scale separable residual blocks of [2, 1, 2]. This branch focuses on extracting features with a larger receptive field to capture emotional states over a longer timeframe. Finally, the outputs of the fine-grained multi-scale separable residual module and the coarse-grained multi-scale separable residual module are fused through feature addition to achieve multi-scale spatial feature extraction of peripheral physiological signals.

[0032] The time domain convolution feature module is used to perform deep time domain feature extraction on the pre-processed peripheral physiological data of the user; the time domain convolution feature module is specially designed to extract the time dimension features in the peripheral physiological signals, and realizes the idea of ​​the attention mechanism by using convolution. The original attention mechanism is divided into three branches: query vector, key vector and value vector. In the time domain convolution feature module, the query vector and the key vector are regarded as one branch, that is, the branch containing the data convolution multi-layer perceptron module, and the other branch is regarded as the key vector, that is, the attention weights and values ​​in the attention mechanism are calculated similarly, and then feature enhancement is achieved through dot product operation; the time domain convolution feature module can capture the short-term and long-term dependencies of time series in emotion recognition, while also avoiding the problem of large computational complexity of the original attention mechanism.

[0033] The dual-stream cross-attention fusion module is used to deeply fuse feature data of different dimensions; the dual-stream cross-attention fusion module includes a feature injector and a feature extractor; the feature injector and the feature extractor are two complementary feature processing operations, and the feature injector injects the feature information output by the time domain convolution feature module into the output of the multi-scale separable residual module; the feature extractor is used to extract the feature information output by the multi-scale separable residual module to enhance the feature information output by the time domain convolution feature module; the dual-stream cross-attention fusion module uses an efficient cross-attention mechanism to achieve efficient fusion of spatiotemporal feature information.

[0034] The emotion category output module is used to output the classification recognition result. The emotion category output module is composed of a fully connected layer. The fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.

[0035] The classification and recognition module includes a loss function calculation and update module and an output module connected in sequence; The loss function calculation and update module is used to calculate the mean and variance of known category sentiment data during the training process, thereby generating a loss function to update the parameters of the recognition model; the loss function calculation and update module includes a similar data mean and variance calculation unit and a loss function construction unit connected in sequence. In the similar data mean and variance calculation unit, the vectors of the second to last layer in the fully connected layer are arranged according to the same category, and the mean and variance of each category are calculated respectively. The loss function is constructed by the mean and variance; during the training process, first follow the same batch of data. The amount of data for training, Set to any positive integer less than the total amount of data, for example After each round of training, the constructed loss function is used to adjust the parameters of the recognition model, and this process is repeated until convergence.

[0036] The output module is used to output the final emotion recognition results. It receives the optimized recognition model parameters from the loss function calculation update module. This module performs the final classification and recognition task by applying these updated parameters to the input emotion data. Specifically, the output module inputs the emotion data to be recognized into the trained model. The model analyzes and processes the data based on its internally learned features and finally outputs the emotion recognition results. These results are presented in the form of category labels, such as "happy", "sad", etc. The specific labels depend on the category range defined in the training data. Ultimately, the user or the system can make corresponding decisions or responses based on these output emotion recognition results.

[0037] The human peripheral physiological information preprocessing module includes a sensor data storage unit, an extreme value removal data unit, a photoplethysmography filtering unit, a sensor data interpolation unit, a data normalization unit, a sliding window segmentation unit, and a data label calibration unit, which are connected in sequence.

[0038] The sensor data storage unit is used to store real-time peripheral physiological data of the user.

[0039] The extreme value removal data unit is used to replace the upper and lower extreme values ​​in the data with the 3% and 97% quantiles to reduce the impact of extreme values ​​on the data.

[0040] The photoplethysmography filtering unit is used to remove high-frequency noise and motion artifacts in the photoplethysmography signal; specifically, it uses a third-order Butterworth low-pass filter with a cutoff frequency of 10 Hz to remove high-frequency noise and motion artifacts.

[0041] The sensor data interpolation unit is used to process the problem of inconsistent sampling rates of peripheral physiological signals; specifically, it uses a method combining linear interpolation and nearest neighbor interpolation to interpolate different peripheral physiological signals to the same sampling rate.

[0042] The data normalization unit is used to eliminate scale differences between different signals and convert data with different dimensions and value ranges to the same value range. Normalization uses the Min-Max normalization method, which converts the values ​​of all similar sensors to the range [0, 1].

[0043] The sliding window segmentation unit is used to segment the continuous long-term sensor data into a plurality of data segments.

[0044] The sliding window segmentation unit is used to segment the continuous long-term sensor data into several data segments; it is suitable for recognition model input. According to the characteristics of human peripheral physiological signals, if the sampling rate is 64Hz, the length of a window can be set to 320, which can include Seconds of sensor data. Sliding step size Can be set to 160, that is data coverage.

[0045] The data label calibration unit is used to perform label calibration on the data after the sliding window segmentation, and formulate a corresponding label for each data segment, that is, the name of the emotion category corresponding to the data segment.

[0046] The human emotion information application module is oriented to various application platforms, and the classification and recognition results are transmitted to the database of each application platform in real time for storage and management, and the user's emotions are analyzed, managed and visualized in real time.

[0047] Example 3

[0048] The recognition method of human emotion recognition system based on dual-stream residual spatiotemporal fusion model, such as Figure 2 As shown, taking the student management application platform as an example, the following steps are included: S1: Human peripheral physiological data collection; Real-time collection of user's peripheral physiological data, i.e. human peripheral physiological data, including photoplethysmography, galvanic skin response, and temperature; S2: Human peripheral physiological data transmission; The collected user's peripheral physiological data is transmitted to a local server or cloud server; according to different scenarios and user needs, appropriate data transmission methods are selected, such as ultra-wideband, Bluetooth, 5G, etc.

[0049] S3: Preprocessing of human peripheral physiological data; The collected user's peripheral physiological data is sequentially stored, de-extremely analyzed, filtered, interpolated, normalized, segmented using a sliding window, and labeled, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments. S4: Building recognition model and sentiment classification and recognition; The peripheral physiological data segments pre-processed by S3 are input into the recognition model in batches, and the emotion recognition output is obtained after training; S5: Loss function calculation update and feedback; The feature vectors output by the previous Dense fully connected layer in the fully connected layer of the sentiment classification and recognition output module in S4 are processed according to the principle of the same category, and the central mean and variance of the feature vectors belonging to the same category and finally classified correctly are calculated, as shown in Formula (1) and Formula (2) respectively. In addition, the variance is normalized, as shown in Formula (3): (1); (2); (3); in: Representation category The central mean vector of Representation category The variance of Represents the normalized category The variance of Representation category The number of correctly classified feature vectors in Representation category Middle The feature vector of the data output in the second to last layer of the fully connected layer, Indicates the The square of the Euclidean distance between a sample and the center of its class, K Indicates the total number of classes; is the total number of all known behavioral categories; The loss function is set as formula (4): (4); in: It represents the maximum distance between the center means of two categories, calculated using the 1-norm; In each round of training, the data of known categories are first divided into batches according to each batch. The amount of data is input into the emotion recognition model for parameter update. Set to any positive integer less than the total amount of data. At this time, the cross entropy loss function commonly used in training is used for training and output. After one round of training is completed, only the loss function in formula (4) is used. Return to S4 for training until convergence, so as to further adjust the recognition model parameters and make the distance between similar data closer in the multidimensional space and the distance between the mean values ​​of heterogeneous data centers farther; S6: Application of discriminant output results; The results of emotion recognition output are transmitted to the corresponding application platform in real time.

[0050] The specific implementation process of S3 is as follows: S31: Remove extreme values ​​from human peripheral physiological data; The collected user's peripheral physiological data was de-extremeized using the Winsorization method to replace the upper and lower extreme values ​​of the three peripheral physiological signals with the 3% and 97% quantiles; S32: filtering of human peripheral physiological data; Human peripheral physiological data filtering refers to the filtering of photoelectric volumetric pulse wave signals with a sampling rate greater than 10 Hz; a 3rd-order Butterworth low-pass filter with a cutoff frequency of 10 Hz is used to remove high-frequency noise and motion artifacts and extract the pure human photoelectric volumetric pulse wave signal. The 3rd-order Butterworth filter is expressed as formula (5): (5); in: is the cutoff frequency, is the passband edge frequency, is the angular frequency, is the passband ripple factor, yes The value at the edge of the passband; S33: Interpolation of human peripheral physiological data; Since the collected peripheral physiological signals have inconsistent sampling rates, the data cannot be directly input into the model for training. Therefore, the data is interpolated using a combination of linear interpolation and nearest neighbor interpolation. For example, if the collected peripheral physiological signals have 1Hz, 4Hz, and 64Hz, then the signal with a sampling frequency of 1Hz is first processed using linear interpolation to obtain a 4Hz signal. Then, the 4Hz signal is interpolated using nearest neighbor interpolation, and finally all the signal sampling rates are 64Hz. The processes of the two interpolation methods are shown in Equations (6) and (7), respectively: (6); (7); in: represents a known point in time, Indicates that it corresponds to The known value of represents the time point of target interpolation, Represents the result after using linear interpolation, Represents the result of using nearest neighbor interpolation; S34: Normalization of human peripheral physiological data; The Min-Max normalization method is used for normalization, that is, the values ​​of all similar sensors are changed to between [0, 1]. The Min-Max normalization method scales all data values ​​to between [0, 1] by linearly transforming the original data. Suppose the input sequence is , the output sequence after minimum-maximum normalization is , as shown in formula (8): (8); in: The first samples, and are the minimum and maximum values ​​of the input sequence, respectively. is the normalized data; S35: Sliding window segmentation and labeling; Sliding window segmentation uses a fixed-length window to segment continuous sensor data into fixed-length data segments, and assigns a corresponding label to each data segment, that is, the name of the emotion category corresponding to the data segment.

[0051] According to the characteristics of human peripheral physiological signals, if the sampling rate is 64Hz, the length of a window can be set to 320, which can include Seconds of sensor data. Sliding step size Can be set to 160, that is data coverage.

[0052] The S4 comprises the following steps: Input the peripheral physiological data segments pre-processed in step S3 into the recognition model in batches, and obtain the emotion recognition output after training; like Figure 3 As shown, the steps include: S41: The multi-scale separable residual module uses a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module to output feature data of different scales respectively, specifically: A combination of a convolution kernel size of 3 and a number of multi-scale separable residual blocks of [1, 2, 1] is used to extract fine-grained multi-scale separable residual features; A convolution kernel size of 5 and the number of multi-scale separable residual blocks of [2, 1, 2] are used to perform coarse-grained multi-scale separable residual feature extraction; In addition, the core component of the multi-scale separable residual module is the multi-scale separable residual block, which is composed of a grouped separable convolution module and a residual connection. Its calculation process is shown in formula (9): (9); in: is the output of the grouped separable convolution module, BN is the batch normalization operation, Conv2D represents the two-dimensional convolution operation, Act represents the LeakyReLU activation function used, and Sum represents the The output of Y represents the final output of the multi-scale separable residual block; The group separable convolution module is an adaptive module that can perform adaptive group convolution according to different data inputs. Its calculation process is shown in formula (10): (10); Among them: G1, G2, ..., G n Representative n Group convolution, each group is composed of a depth-wise separable convolution and Conv2D, with BN and LeakyReLU added between the two convolutions. Exp represents the channel expansion operation, and Concat represents the connection operation. Then, the feature addition operation is used to further fuse the features at different scales from the outputs of the fine-grained multi-scale separable residual module and the coarse-grained multi-scale separable residual module; S42: The time domain convolution feature module extracts deep time domain features from the pre-processed user's peripheral physiological data. Specifically, The attention mechanism is implemented by convolution. The original attention mechanism is divided into three branches: query vector, key vector and value vector. In the time domain convolution feature module, the query vector and key vector are regarded as one branch, that is, the branch containing the time domain convolution multi-layer perceptron module, and the other branch is regarded as the key vector. That is, the attention weight and value in the attention mechanism are calculated similarly, and then the feature enhancement is achieved through the dot product operation. The calculation process is shown in Equations (11), (12) and (13): (11); (12); (13); in: x represents the original input feature sequence, Y 1 is the output of TCM, Y 2 means Y The output after the dot product operation of 1 and another branch, is the output of the time-domain convolution feature module, LN is layer normalization, Conv1D is a 1D convolution operation, Act is the GELU activation function, and TCM is a time-domain convolution multilayer perceptron module. The calculation process is shown in Equations (14), (15), and (16): (14); (15); (16); Among them: Drop is the DropOut operation with a drop rate of 0.5, DSConv1D is a one-dimensional depth-separable convolution, is the output of TCM; S43: The dual-stream cross-attention fusion module deeply fuses feature data of different dimensions, specifically: The two-stream cross-attention fusion module includes a feature injector and a feature extractor; The feature injector and feature extractor are two complementary feature processing operations. The feature injector converts the feature information output by the time domain convolution feature module into Injected into the output of the multi-scale separable residual module The calculation process is shown in formula (17): (17); in: is the injection intensity, FFN is a multi-layer perceptron using the GELU activation function, and attn is an efficient cross-attention mechanism. is the output of the feature injector; The feature extractor is used to extract the feature information output by the multi-scale separable residual module , to enhance the feature information output by the time domain convolution feature module , and its calculation process is shown in formula (18): (18); in: is the output of the feature extractor. FFN is a multi-layer perceptron using the GELU activation function. S44: Classification and recognition is to output the emotion classification and recognition results, specifically: The emotion category output module is composed of a fully connected layer, and the fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.

Claims

1. A human emotion recognition system based on a dual-stream residual spatiotemporal fusion model, characterized by: include: A human peripheral physiological information acquisition module, a human peripheral physiological information transmission module, a human peripheral physiological information preprocessing module, a human emotion classification and recognition module, and a human emotion information application module connected in sequence; The human peripheral physiological information acquisition module is used to: collect the user's peripheral physiological data in real time, the peripheral physiological data including photoplethysmography, skin galvanic response and temperature; The human peripheral physiological information transmission module is used to: transmit the collected user's peripheral physiological data to a local server or a cloud server; The human peripheral physiological information preprocessing module is used to sequentially store, remove extreme values, filter, interpolate, normalize, perform sliding window segmentation, and label calibration on the collected user's peripheral physiological data, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments; The human emotion classification and recognition module is used to: input the user's physiological data preprocessed by the human peripheral physiological information preprocessing module into the trained human emotion classification and recognition module to perform emotion type discrimination; The human emotion information application module is used to transmit the obtained human emotion classification and recognition results to various application platforms, thereby realizing corresponding functions.

2. The human emotion recognition system based on the dual-stream residual spatiotemporal fusion model according to claim 1 is characterized in that: The human emotion classification and recognition module includes a parallel-connected multi-scale separable residual module and a time-domain convolution feature module, a dual-stream cross-attention fusion module that fuses parallel connection features, and an emotion category output module; The multi-scale separable residual module includes a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module connected in parallel; The dual-stream cross-attention fusion module includes a feature injector and a feature extractor; The emotion category output module is used to output the classification recognition result. The emotion category output module is composed of a fully connected layer. The fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.

3. The human emotion recognition system based on the dual-stream residual spatiotemporal fusion model according to claim 2 is characterized in that: The human peripheral physiological information acquisition module includes several different sensor units, including a photoplethysmography sensor, a galvanic skin response sensor, and a temperature sensor, for acquiring the user's peripheral physiological data.

4. The human emotion recognition system based on the dual-stream residual spatiotemporal fusion model according to claim 1, characterized in that: The human peripheral physiological information preprocessing module includes a sensor data storage unit, an extreme value removal data unit, a photoplethysmography filtering unit, a sensor data interpolation unit, a data normalization unit, a sliding window segmentation unit, and a data label calibration unit, which are connected in sequence.

5. The human emotion recognition system based on the dual-stream residual spatiotemporal fusion model according to claim 1 is characterized in that: The human peripheral physiological information transmission module transmits the collected user's peripheral physiological data to a local server or a cloud server via any transmission method including ultra-wideband, Bluetooth, and 5G.

6. A recognition method for a human emotion recognition system based on a dual-stream residual spatiotemporal fusion model according to claim 4, characterized in that: The following steps are involved: S1: Human peripheral physiological data collection; Real-time collection of user's peripheral physiological data, i.e. human peripheral physiological data, including photoplethysmography, galvanic skin response, and temperature; S2: Peripheral physiological data transmission of the human body; Transmitting the collected user's peripheral physiological data to a local server or a cloud server; S3: Preprocessing of human peripheral physiological data; The collected user's peripheral physiological data is sequentially stored, de-extremely analyzed, filtered, interpolated, normalized, segmented using a sliding window, and labeled, thereby obtaining smooth, diverse, dimensionally independent, and labeled emotional data segments. S4: Building recognition model and sentiment classification and recognition; The peripheral physiological data segments pre-processed by S3 are input into the recognition model in batches, and the emotion recognition output is obtained after training; S5: Loss function calculation update and feedback; The feature vectors output by the previous Dense fully connected layer in the fully connected layer of the sentiment classification and recognition output module in S4 are processed according to the principle of the same category, and the central mean and variance of the feature vectors belonging to the same category and finally classified correctly are calculated, as shown in Formula (1) and Formula (2) respectively. In addition, the variance is normalized, as shown in Formula (3): (1); (2); (3); in: Representation category The central mean vector of Representation category The variance of Represents the normalized category The variance of Representation category The number of correctly classified feature vectors in Representation category Middle The feature vector of the data output in the second to last layer of the fully connected layer, Indicates the The square of the Euclidean distance between a sample and the center of its class, K Indicates the total number of classes; is the total number of all known behavior categories; The loss function is set as formula (4): (4); in: It represents the maximum distance between the center means of two categories, calculated using the 1-norm; S6: Application of discriminant output results; The results of emotion recognition output are transmitted to the corresponding application platform in real time.

7. The identification method according to claim 6, characterized in that: The specific implementation process of S3 is as follows: S31: Remove extreme values ​​from human peripheral physiological data; The collected user's peripheral physiological data were de-extremeized by using the Windsor method to replace the upper and lower extreme values ​​of the three peripheral physiological signals with the 3% and 97% quantiles; S32: filtering of human peripheral physiological data; Human peripheral physiological data filtering refers to the filtering of photoelectric volumetric pulse wave signals with a sampling rate greater than 10 Hz; a 3rd-order Butterworth low-pass filter with a cutoff frequency of 10 Hz is used to remove high-frequency noise and motion artifacts and extract pure human photoelectric volumetric pulse wave signals. The 3rd-order Butterworth filter is expressed as formula (5): (5); in: is the cutoff frequency, is the passband edge frequency, is the angular frequency, is the passband ripple factor, yes The value at the edge of the passband; S33: Interpolation of human peripheral physiological data; Since the collected peripheral physiological signals have inconsistent sampling rates, the data cannot be directly input into the model for training. Therefore, the data is interpolated using a combination of linear interpolation and nearest neighbor interpolation. The processes of the two interpolation methods are shown in Equations (6) and (7), respectively: (6); (7); in: represents a known point in time, Indicates that it corresponds to The known value of represents the time point of target interpolation, Represents the result after using linear interpolation, Represents the result of using nearest neighbor interpolation; S34: Normalization of human peripheral physiological data; Use the minimum-maximum normalization method for normalization. Suppose the input sequence is , the output sequence after minimum-maximum normalization is , as shown in formula (8): (8); in: The first samples, and are the minimum and maximum values ​​of the input sequence, respectively. is the normalized data; S35: Sliding window segmentation and labeling; Sliding window segmentation uses a fixed-length window to segment continuous sensor data into fixed-length data segments, and assigns a corresponding label to each data segment, that is, the name of the emotion category corresponding to the data segment.

8. The identification method according to claim 6, wherein: The S4 comprises the following steps: S41: The multi-scale separable residual module uses a fine-grained multi-scale separable residual module and a coarse-grained multi-scale separable residual module to output feature data of different scales respectively, specifically: A combination of a convolution kernel size of 3 and a number of multi-scale separable residual blocks of [1, 2, 1] is used to extract fine-grained multi-scale separable residual features; A convolution kernel size of 5 and the number of multi-scale separable residual blocks of [2, 1, 2] are used to perform coarse-grained multi-scale separable residual feature extraction; In addition, the core component of the multi-scale separable residual module is the multi-scale separable residual block, which is composed of a grouped separable convolution module and a residual connection. Its calculation process is shown in formula (9): (9); in: is the output of the grouped separable convolution module, BN is the batch normalization operation, Conv2D represents the two-dimensional convolution operation, Act represents the LeakyReLU activation function used, and Sum represents the The output of Y represents the final output of the multi-scale separable residual block; The group separable convolution module is an adaptive module that can perform adaptive group convolution according to different data inputs. Its calculation process is shown in formula (10): (10); Among them: G1, G2, ..., G n Representative n Group convolution, each group is composed of a depth-wise separable convolution and Conv2D, with BN and LeakyReLU added between the two convolutions. Exp represents the channel expansion operation, and Concat represents the connection operation. Then, the feature addition operation is used to further fuse the features at different scales from the outputs of the fine-grained multi-scale separable residual module and the coarse-grained multi-scale separable residual module; S42: The time domain convolution feature module extracts deep time domain features from the pre-processed user's peripheral physiological data. Specifically, In the time-domain convolution feature module, the query vector and the key vector are regarded as one branch, namely the branch containing the time-domain convolution multilayer perceptron module, and the other branch is regarded as the key vector. Similarly, the attention weights and values ​​in the attention mechanism are calculated respectively, and then the feature enhancement is achieved through the dot product operation. The calculation process is shown in Equations (11), (12) and (13): (11); (12); (13); in: x represents the original input feature sequence, Y 1 is the output of TCM, Y 2 means Y The output after the dot product operation of 1 and another branch, is the output of the time-domain convolution feature module, LN is layer normalization, Conv1D is a 1D convolution operation, Act is the GELU activation function, and TCM is a time-domain convolution multilayer perceptron module. The calculation process is shown in Equations (14), (15), and (16): (14); (15); (16); Among them: Drop is the DropOut operation with a drop rate of 0.5, DSConv1D is a one-dimensional depth-separable convolution, is the output of TCM; S43: The dual-stream cross-attention fusion module deeply fuses feature data of different dimensions, specifically: The two-stream cross-attention fusion module includes a feature injector and a feature extractor; The feature injector and feature extractor are two complementary feature processing operations. The feature injector converts the feature information output by the time domain convolution feature module into Injected into the output of the multi-scale separable residual module The calculation process is shown in formula (17): (17); in: is the injection intensity, attn is the efficient cross-attention mechanism, is the output of the feature injector; The feature extractor is used to extract the feature information output by the multi-scale separable residual module , to enhance the feature information output by the time domain convolution feature module , and its calculation process is shown in formula (18): (18); in: is the output of the feature extractor. FFN is a multi-layer perceptron using the GELU activation function. S44: Classification and recognition is to output the emotion classification and recognition results, specifically: The emotion category output module is composed of a fully connected layer, and the fused features are input into the emotion category output module, and finally the emotion recognition result can be obtained.