Emotion recognition method and device

CN120337025APending Publication Date: 2025-07-18GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351218.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-18

Smart Images

  • Figure CN120337025A_ABST
    Figure CN120337025A_ABST
Patent Text Reader

Abstract

The invention provides an emotion recognition method and device, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring a multi-modal physiological signal of a target object; inputting the physiological signal into an emotion recognition model to obtain an emotion recognition result of the target object; wherein the emotion recognition model is used for performing dimensionality reduction on a time domain signal and a frequency domain signal corresponding to the physiological signal to obtain a dimensionality-reduced time domain signal and a dimensionality-reduced frequency domain signal; fusing the time domain signal after dimension reduction and the frequency domain signal after dimension reduction to obtain a first scale fusion feature; fusing the first scale fusion feature with the time domain signal before dimension reduction and the frequency domain signal before dimension reduction to obtain a second scale fusion feature and a third scale fusion feature; fusing the second scale fusion feature and the third scale fusion feature to obtain a fourth scale fusion feature; and obtaining an emotion recognition result of the target object based on the fourth scale fusion feature. The method effectively improves the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of Artificial Intelligence (AI). Specifically, it relates to an emotion recognition method and device, and more specifically, to an emotion recognition method, device, and electronic device. Background Art

[0002] Emotions are a series of complex reactions caused by internal or external stimuli and play a crucial role in human cognition and decision-making processes. Emotion recognition has been used in various practical scenarios, such as autonomous driving assistance, healthcare, social communication, etc.

[0003] Most emotion recognition methods are based on the physical or physiological indicators of the human body. Compared with physical signals such as facial expressions and language, the physiological reactions in certain emotional states are involuntary, so they can provide more objective decisions for the recognition system. The physiological signals used for emotion recognition include photoplethysmogram signal (PPG), electrocardiogram signal (ECG), electrodermal activity signal (EDA), and skin temperature signal (SKT), etc. Among them, PPG or ECG reflects the characteristics of cardiac sympathetic nerve activity, EDA reflects the change of skin impedance, and together with SKT indicates parasympathetic nerve activity and stress perception.

[0004] Currently, the shallow time-frequency feature fusion mechanism adopted in related physiological signal-based emotion recognition tasks is difficult to fully exploit the cross-domain correlation, and the deep interaction relationship between the time-domain dynamic features and the frequency-domain energy features has not been effectively modeled, which limits the accuracy of emotion recognition. Summary of the Invention

[0005] This application aims to at least solve one of the technical problems in the related art to some extent. For this purpose, an object of this application is to propose an emotion recognition method and device that can effectively improve the accuracy of emotion recognition.

[0006] Specifically, the technical solution of this application is as follows:

[0007] In the first aspect of the present application, the present application proposes an emotion recognition method. According to an embodiment of the present application, the method includes: obtaining multi-modal physiological signals of a target object; inputting the physiological signals into an emotion recognition model to obtain an emotion recognition result of the target object; wherein, the emotion recognition model is used to: respectively perform dimensionality reduction on the time-domain signal and the frequency-domain signal corresponding to the physiological signal to obtain the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal; fuse the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal to obtain a first-scale fusion feature; fuse the first-scale fusion feature with the time-domain signal and the frequency-domain signal before dimensionality reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; and based on the fourth-scale fusion feature, obtain the emotion recognition result of the target object.

[0008] The present application realizes an effective improvement in emotion recognition accuracy by constructing a unified multi-modal coding architecture and a cross-level time-frequency feature fusion mechanism. Specifically, a dimensionality reduction mapping and parameter sharing strategy is adopted to complete cross-modal feature alignment while reducing the model complexity, with a significant reduction in the number of parameters and an obvious improvement in the training efficiency compared with traditional methods; secondly, a cross-level bidirectional fusion framework is innovatively proposed to solve the feature decoupling problem caused by shallow fusion in the traditional time-frequency domain through heterogeneous feature complementary fusion, residual cross-scale interaction and time-frequency depth coupling.

[0009] In some examples of the present application, the dimensions of the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal are both target dimensions.

[0010] In some examples of the present application, the fusing the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal to obtain a first-scale fusion feature includes: based on an activation function, fusing the full connection layer weights, the frequency-domain features and the time-domain features to obtain a first-scale fusion feature.

[0011] In some examples of the present application, the based on an activation function, fusing the full connection layer weights, the frequency-domain features and the time-domain features to obtain a first-scale fusion feature includes: performing upsampling on the frequency-domain features and the time-domain features in the time dimension to obtain the upsampled time-domain features and frequency-domain features; and using an activation function to fuse the full connection layer weights and the upsampled time-domain features and frequency-domain features to obtain a first-scale fusion feature.

[0012] In some examples of the present application, the multi-modal physiological signals are selected from multiple types including optoelectronic signals, electrocardiogram signals, galvanic skin response signals and skin temperature signals.

[0013] In some examples of the present application, the foregoing emotion recognition model is obtained through training in the following manner: obtaining a training data set, the foregoing training data set including: multimodal physiological signals and corresponding emotion categories; using a data augmentation method to amplify the foregoing multimodal physiological signals to obtain amplified multimodal physiological signals; inputting the amplified foregoing multimodal physiological signals into a deep learning model, using the foregoing data augmentation method as a pseudo label, performing self-supervised learning on the foregoing deep learning model to obtain a pre-trained model; inputting the foregoing training data set into the foregoing pre-trained model, using the foregoing emotion categories as labels, and fine-tuning the foregoing pre-trained model to obtain the foregoing emotion recognition model.

[0014] In some examples of the present application, the foregoing data augmentation method includes at least one of noise processing, signal transformation, and timing adjustment. Among them, the foregoing signal transformation includes at least one of amplitude inversion and segment distortion, and the foregoing timing adjustment includes at least one of time dimension reversal and segment shuffling.

[0015] In some examples of the present application, the foregoing emotion recognition model is obtained through training in the following manner: obtaining a training data set, the foregoing training data set including: multimodal physiological signals and corresponding emotion categories; inputting the foregoing training data set into a deep learning model, using the foregoing emotion categories as labels, and fine-tuning the foregoing deep learning model to obtain the foregoing emotion recognition model.

[0016] In a second aspect of the present application, the present application proposes an emotion recognition device. According to an embodiment of the present application, the device includes: a signal acquisition module for acquiring multimodal physiological signals of a target object; an emotion recognition module for inputting the physiological signals into an emotion recognition model to obtain an emotion recognition result of the target object; wherein, the emotion recognition model is used to: respectively perform dimensionality reduction on the time domain signal and the frequency domain signal corresponding to the physiological signals to obtain a dimensionality-reduced time domain signal and a dimensionality-reduced frequency domain signal; fuse the dimensionality-reduced time domain signal and the dimensionality-reduced frequency domain signal to obtain a first-scale fusion feature; fuse the first-scale fusion feature with the time domain signal and the frequency domain signal before dimensionality reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; and based on the fourth-scale fusion feature, obtain the emotion recognition result of the target object.

[0017] The emotion recognition device of the present application realizes efficient and accurate real-time emotion analysis through a modular design and a hierarchical feature fusion architecture. Specifically, the signal acquisition module supports parallel acquisition and preprocessing of multi-modal physiological signals. Combining with the unified dimensionality reduction coding strategy of the emotion recognition module, it significantly reduces the consumption of computing resources. Through a three-stage processing chain of time-frequency signal dimensionality reduction fusion → cross-scale residual interaction → multi-level feature coupling, it breaks the time-frequency domain information barrier of traditional devices and significantly improves the emotion classification accuracy on the benchmark dataset.

[0018] In the third aspect of the present application, the present application proposes an electronic device. According to an embodiment of the present application, the device includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the emotion recognition method of the first aspect.

[0019] In the fourth aspect of the present application, the present application proposes a computer-readable storage medium. According to an embodiment of the present application, the computer-readable storage medium includes computer instructions, and when the instructions are executed by a computer, the computer realizes the emotion recognition method of the first aspect.

[0020] In the fifth aspect of the present application, the present application proposes a computer program product. According to an embodiment of the present application, the computer program product includes computer instructions, and when part or all of the computer instructions run on a computer, the emotion recognition method as in the first aspect is executed.

[0021] The aforementioned electronic device, computer-readable storage medium, and computer program product realize high-efficiency automation by automatically executing the emotion recognition method through computer instructions. In addition, due to the characteristics of the instructions, they have better stability in different application scenarios.

[0022] Additional aspects and advantages of the present application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic diagram of the system architecture provided by an embodiment of the present application;

[0025] Figure 2 It is a schematic diagram of the process flow of the emotion recognition method provided by an embodiment of the present application;

[0026] Figure 3 It is a schematic diagram of a deep learning network architecture provided by an embodiment of the present application;

[0027] Figure 4 It is a schematic diagram of a ResNet Block network architecture provided by an embodiment of the present application;

[0028] Figure 5 It is a schematic diagram of an emotion recognition device provided by an embodiment of the present application;

[0029] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0032] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0033] In this article, unless otherwise specified, a physiological signal refers to an electrical, mechanical, acoustic, optical, chemical, etc. signal generated by the physiological activities of an organism, which can reflect the physiological state, functional activities, or pathological changes of the body.

[0034] In this text, unless otherwise specified, a "model" is an algorithm that can automatically perform tasks such as prediction, classification, recognition, or decision-making by learning and analyzing input data. The learning process of this model is based on statistical principles and data pattern recognition, and uses a training data set to adjust model parameters and optimize the model to improve its prediction or inference ability. The model can adopt various algorithms and technologies, such as neural networks, support vector machines, decision trees, random forests, deep learning, etc. These models can be trained and optimized through supervised learning, unsupervised learning, or reinforcement learning. In practical applications, the model can be used in various fields, such as natural language processing, image recognition, pattern recognition, data mining, recommendation systems, predictive analysis, etc. It has important application potential in processing large-scale data, automated decision-making, and intelligent systems. However, it should be noted that in the specific application of the model, in-depth research on the features and models used for prediction is required to obtain relatively satisfactory prediction results. Otherwise, various problems will occur, such as overfitting problems, underfitting problems, poor generalization ability, low robustness problems, etc. The inventors of this application have found through a large number of experimental verifications that using multimodal physiological signals to train a deep learning model can obtain a model that can adaptively perform emotion recognition.

[0035] In this text, unless otherwise specified, a "Training Set" refers to the multimodal physiological signals used to train a deep learning model. It contains a series of input samples (such as time-domain signals and frequency-domain signals corresponding to multimodal physiological signals) and corresponding known outputs (such as emotion categories). In the training phase, the deep learning model uses the samples in the training set to learn and adjust the model's parameters to minimize the error between the prediction result and the true label. By continuously trying different parameters and algorithms, the model gradually learns the rules and features between the samples, thereby obtaining more accurate prediction ability.

[0036] In this text, unless otherwise specified, a "Test Set" or validation set is a data set used to evaluate the performance of a deep learning model. It also contains input samples (such as time-domain signals and frequency-domain signals corresponding to multimodal physiological signals) and corresponding true outputs (such as emotion categories), but the samples in the test set are not used in the training phase of the model, which can ensure that the data in the test set is unknown to the model. After training is completed, the test set is used to evaluate the performance of the model, that is, the model makes predictions on the samples in the test set and compares them with the true labels. By comparing the prediction results and the true labels, the performance of the model on unknown data is evaluated, and its generalization ability and prediction accuracy are judged.

[0037] In this application, unless otherwise specified, data augmentation refers to a technique that generates more training samples by transforming, synthesizing, or expanding the original data without changing the data labels, in order to improve the generalization ability of the model, reduce overfitting, and enhance the adaptability of the model to different environments.

[0038] In this application, unless otherwise specified, self-supervised learning (SSL) is a paradigm of machine learning and a subclass of unsupervised learning. Its core idea is to automatically generate supervision signals through the internal structure or information of the data itself, without relying on manually labeled tags, so as to train the model to learn effective representations of the data. Specifically, the model designs specific pretext tasks (such as predicting masked data segments, reconstructing inputs, contrasting similar samples, etc.), converts the original data into "input-pseudo label" pairs, and then obtains a pre-trained model through supervised learning. The learned features can be transferred to downstream tasks (such as classification, detection, etc.), significantly reducing the dependence on labeled data.

[0039] In this application, unless otherwise specified, supervised learning (SL) is a machine learning method that learns the mapping relationship between input data and output labels by using a labeled training data set, so as to predict or classify new data. During the supervised learning process, the model is trained based on known input-output pairs and adjusts the parameters through optimization algorithms to make the prediction results as close as possible to the true labels.

[0040] Currently, the scale of the data sets that can be provided for emotion recognition tasks based on peripheral physiological signals is relatively small. Through the method of predefined feature extraction, supervised learning with emotion recognition as the task, due to the pre-extraction of features, some information is lost, and it may lead to an increase in individual differences among subjects, thus reducing the generalization ability of the model. Due to insufficient data volume and a relatively large model size, directly performing end-to-end supervised learning with emotion recognition as the final task is likely to lead to underfitting of the model.

[0041] Since the sample size of the data sets for emotion computing tasks using peripheral signals is usually limited, great challenges are encountered in model training. Augmenting the data through signal transformation and then using the transformed data for self-supervised learning (SSL) is an important strategy to enhance the model's ability in limited sample scenarios. At the same time, in subject-independent models, compared with multi-modal scenarios, single-modal peripheral signals exhibit lower performance. However, current self-supervised learning multi-modal models usually design encoders separately for each modality and independently train these encoders for each modality, resulting in a larger model size and longer training time.

[0042] In addition, since the shallow time-frequency feature fusion mechanism adopted by the relevant physiological signal-based emotion recognition tasks is difficult to fully exploit cross-domain correlations, the deep interaction relationship between the time-domain dynamic features and the frequency-domain energy features has not been effectively modeled, which limits the accuracy of emotion recognition.

[0043] To address the above deficiencies, the inventors have carried out optimizations in various aspects, such as augmenting multi-modal signal data (such as masking, reverse output in the time dimension, etc.), constructing a unified multi-modal physiological signal encoding architecture, and adopting a cross-level time-frequency feature fusion mechanism.

[0044] In some embodiments, the system architecture of the embodiments of the present application is as Figure 1 shown.

[0045] Figure 1 FIG. is a schematic diagram of a system architecture related to an embodiment of the present application. The system architecture involves a user device 101, a data acquisition device 102, a training device 103, an execution device 104, a database 105, and a content library 106.

[0046] Among them, the data acquisition device 102 is used to read training samples from the content library 106 and store the read training samples in the database 105. The training samples related to the embodiments of the present application can come from a public database.

[0047] The training device 103 trains a deep learning model based on the training samples maintained in the database 105, and the obtained emotion recognition model can effectively perform emotion recognition.

[0048] In addition, referring to Figure 1 , the execution device 104 is configured with an I / O interface 107 to interact with external devices. For example, it receives an input instruction sent by the user device 101 through the I / O interface; wherein, the input instruction includes: obtaining multi-modal physiological signals of a target object. The computing module 108 in the execution device 104 processes the input instruction using the emotion recognition model to obtain an output instruction; wherein, the output instruction includes: the emotion recognition result of the target object, and sends the output instruction to the user device 101 through the I / O interface.

[0049] Among them, the user device 101 may include a mobile phone, a tablet computer, a laptop computer, a handheld computer, a mobile internet device (MID), a desktop computer, or other terminal devices with a browser function installed.

[0050] The execution device 104 may be a server.

[0051] Exemplarily, the server may be a computing device such as a rack server, a blade server, a tower server, or a cabinet server. The server may be an independent test server or a test server cluster composed of multiple test servers.

[0052] In this embodiment, the execution device 104 is connected to the user device 101 through a network. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), the 4th Generation (4G) network, the 5th Generation (5G) network, Bluetooth, wireless fidelity (Wi-Fi), or a call network.

[0053] It should be noted that Figure 1 is only a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitation. In some embodiments, the above data acquisition device 102, the user device 101, the training device 103, and the execution device 104 may be the same device. The above database 105 may be distributed on one server or on multiple servers, and the above content library 106 may be distributed on one server or on multiple servers.

[0054] The technical solution of the present application will be elaborated in detail below:

[0055] In one aspect of the present application, the present application proposes an emotion recognition method, which can be executed by a training device. For example, it can be executed by Figure 1 the training device 103 shown, but not limited thereto. As Figure 2 shown, the method may include:

[0056] S210, obtaining multi-modal physiological signals of a target object through a training device;

[0057] In some examples of the present application, the foregoing multi-modal physiological signals may be selected from multiple types including photoplethysmogram signal (PPG), electrocardiogram signal (ECG), electrodermal activity signal (EDA), and skin temperature signal (SKT). By providing a more comprehensive emotion representation through multi-modal signals (such as the relationship between skin temperature changes and emotion arousal), the recognition robustness is improved.

[0058] It can be understood that the foregoing multi-modal physiological signals can be detected and obtained by a wearable device.

[0059] S210, input the physiological signal into an emotion recognition model to obtain the emotion recognition result of the target object;

[0060] Among them, the emotion recognition model is used to: respectively reduce the dimensions of the time-domain signal and the frequency-domain signal corresponding to the physiological signal to obtain the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal; fuse the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal to obtain a first-scale fusion feature; fuse the first-scale fusion feature with the time-domain signal and the frequency-domain signal before dimension reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; based on the fourth-scale fusion feature, obtain the emotion recognition result of the target object.

[0061] In some examples of the present application, the foregoing emotion recognition results include: calm, pleasant, anxious, angry, sad, and evaluations on the "tendency-arousal" plane, etc.

[0062] Since the signal change speed characteristics of different modal signals are different, by reducing the dimensions of the time-domain signal and the frequency-domain signal to capture the wider range and slower characteristics of the signal, and further through a multi-stage and different-scale time-domain feature and frequency-domain feature fusion mechanism, it is ensured that the convolutional neural network (CNN) obtains a larger receptive field and enhanced feature expression ability, which can effectively improve the accuracy of emotion recognition.

[0063] In the present application, the time-domain signal refers to a signal that is directly measured in the time dimension and reflects the changes in physiological activities.

[0064] In the present application, the frequency-domain signal is a signal obtained by performing frequency analysis on the time-domain signal, and usually uses techniques such as Fourier transform to convert the time-domain signal to the frequency domain. The frequency-domain signal can reveal the frequency composition of the signal, such as the characteristic information of frequency bands such as α wave, β wave, and θ wave in EEG signals.

[0065] In some examples of the present application, the dimensions of the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal are both target dimensions. Exemplarily, the foregoing target dimension is 1.

[0066] Time-domain signals and frequency-domain signals usually have high-dimensional characteristics (such as thousands of sampling points or frequency bands), and direct fusion will lead to an explosive increase in the amount of calculation. In this application, the signals are compressed to a single dimension through downsampling, significantly reducing the feature dimension, reducing the model calculation complexity, and improving the fusion efficiency. In addition, time-domain and frequency-domain signals have different physical meanings and numerical ranges, and direct fusion may lead to mismatched feature distributions. After downsampling to the target dimension, the time-domain and frequency-domain signals are mapped to the same scale, facilitating subsequent feature fusion modules (such as weighted summation, concatenation) to achieve cross-modal alignment.

[0067] In some examples of this application, fusing the downsampled time-domain signal and the downsampled frequency-domain signal to obtain the first-scale fusion feature includes: based on an activation function, fusing the fully connected layer weights, frequency-domain features, and time-domain features to obtain the first-scale fusion feature. Fusing based on the time-domain signal and frequency-domain signal of the target dimension avoids the conflict between the time semantics and frequency semantics of the time domain and the frequency domain, and at the same time ensures the deep fusion of time-frequency domain features.

[0068] In some examples of this application, based on the activation function, fusing the fully connected layer weights with the frequency-domain features and time-domain features to obtain the first-scale fusion feature includes: upsample the frequency-domain features and time-domain features in the time dimension to obtain the upsampled time-domain features and frequency-domain features; use the Gelu activation function to fuse the fully connected layer weights with the upsampled time-domain features and frequency-domain features to obtain the first-scale fusion feature. Different physiological signals contain information in different frequency bands and time periods, and direct fusion is difficult. By downsampling the time-domain signal and the frequency-domain signal to the target dimension for fusion, the conflict between the time semantics and frequency semantics of the time domain and the frequency domain is avoided, and at the same time, the deep fusion of time-frequency domain features is ensured.

[0069] In some preferred examples of this application, the activation function is selected from the Gelu activation function.

[0070] In some examples of this application, the first-scale fusion feature is expressed as:

[0071]

[0072] where F merg represents the first-scale fusion feature; Gelu represents the function; W m is the fully connected layer weight; represents upsampling along the time axis; F f represents the frequency-domain feature; F t represents the time-domain feature.

[0073] Based on the above-mentioned fusion method of the structure, it is ensured that the fusion of the time-domain signal and the frequency-domain signal will not cause semantic conflicts, thus contributing to the effective extraction of the prominent features in both domains.

[0074] For the sake of easy understanding, refer to the deep learning network architecture shown in this application Figure 3 to elaborate in detail on the data processing process of the emotion recognition model.

[0075] Step 1: Multi-modal signal preprocessing and input construction

[0076] The multi-modal physiological signals PPG, EDA, and SKT of the target object are combined into a time-domain data matrix of (batch, time, channel) according to the channel dimension. At the same time, the amplitude-frequency characteristic curves of each modality are extracted through the fast Fourier transform (FFT) to construct a frequency-domain data matrix of (batch, frequency, channel), forming a time-frequency dual-channel input.

[0077] Step 2: Time-domain and frequency-domain multi-scale feature extraction

[0078] The time-domain processing branch contains multiple cascaded residual connection blocks (ResNet Block, RSB). Each residual block contains a convolutional layer and a skip connection structure ( Figure 4 ) and extracts deep features by gradually increasing the feature dimension, and cooperates with the max pooling layer (Maxpooling) to perform downsampling operations to compress the signal dimension.

[0079] Frequency-domain processing branch: The frequency-domain data matrix is processed by the RSB network with the same structure (the number of filters and the pooling strategy are the same as those of the time-domain branch), and deep frequency-domain features are output.

[0080] Step 3: Time-frequency domain deepest scale feature fusion (Concat)

[0081] The deepest scale features in the time domain and the frequency domain (both the time / frequency dimension is 1) are concatenated along the channel dimension, the time dimension is restored through upsampling (UpSample), and the fully connected layer is used for feature compression and cross-dimensional mapping to achieve complementary fusion of time-frequency information and avoid semantic conflicts.

[0082] Step 4: Multi-scale feature joint optimization

[0083] The time dimension of the original time domain and frequency domain features at each scale is restored through upsampling respectively. After being compressed by the fully connected layer, they are fused along the feature dimension, and the dual-head attention mechanism is introduced to strengthen the interaction of key features. After passing through the flattening layer (Flat) and the fully connected layer, the features output by the time-domain encoder (rU-net Encoder) and the features output by the frequency-domain encoder (rU-net Encoder) are fused to generate discriminative multi-scale fusion features.

[0084] Step Five: Classified Output

[0085] After reducing the dimension of the fused features through a fully connected layer, input them into the classification layer to output the emotion recognition results (such as joy, anxiety, anger, sadness, etc.).

[0086] In some examples of this application, the training data for the following emotion recognition model training has been preprocessed. The aforementioned preprocessing includes: resampling multimodal physiological signals (such as PPG, EDA, and SKT) to the same frequency (such as 64 Hz), and performing sliding truncation with a fixed window length (such as using 20 s as the window and 5 s as the overlap for sliding truncation); for the PPG signal, using a band-pass Hamming window FIR filter with a fixed bandwidth (such as 1 - 8 Hz) to remove high-frequency noise and baseline drift; for the EDA and SKT signals, using a fourth-order Butterworth IIR low-pass filter with a fixed cut-off frequency (such as 5 Hz) to remove high-frequency noise. Compared with some methods that use wavelet transform and Hilbert-Huang transform to preprocess signals, the data preprocessing method of this application is computationally simple.

[0087] In some examples of this application, the emotion recognition model is obtained through the following training method: Obtain a training data set, where the training data set includes: multimodal physiological signals and corresponding emotion categories; use data augmentation methods to amplify the multimodal physiological signals to obtain amplified multimodal physiological signals; input the amplified multimodal physiological signals into a deep learning model, and use the data augmentation method as a pseudo-label to perform self-supervised learning on the deep learning model to obtain a pre-trained model; input the training data set into the pre-trained model, and use the emotion category as a label to fine-tune the pre-trained model to obtain the emotion recognition model. Based on the network architecture of this application ( Figure 3 ) for pre-training and fine-tuning can effectively improve the accuracy of emotion recognition. In the pre-training stage, using multimodal signal augmentation data can effectively improve the robustness of the model to noise and deformation; self-supervised learning using the data augmentation method as a label can effectively learn the general features of different data augmentation methods; fine-tuning based on the weights in the pre-training stage for specific emotion recognition tasks can effectively improve the accuracy and stability of the prediction of the emotion recognition model.

[0088] In some examples of the present application, the data augmentation methods include at least one of noise processing, signal transformation, and timing adjustment. Among them, the signal transformation includes at least one of amplitude inversion and segment distortion, and the timing adjustment includes at least one of time dimension reverse order and segment scrambling. Currently, the scale of physiological signal data is small. Through various data augmentation methods, the data volume of physiological signals for pre-training can be effectively increased. In addition, physiological signals are vulnerable to environmental noise, equipment errors, or individual physiological differences, resulting in poor model generalization. The augmented data covers more potential perturbation scenarios (such as noise simulating equipment errors and segment scrambling simulating signal interruption), improving the robustness of the model.

[0089] For ease of understanding, the present application takes public datasets (such as CASE and DEAP) as examples to elaborate on the training process of the emotion recognition model of the present application in detail. Among them, the training process includes: self-supervised pre-training and fine-tuning.

[0090] Step 1: Data preprocessing

[0091] 1.1 Use three physiological signals, PPG, EDA, and SKT, to resample the original signal to obtain a signal sequence with a predetermined sampling frequency, and use a sliding time window for signal segmentation, setting the window length and overlap step.

[0092] 1.2 For the PPG signal, use a band-pass Hamming window FIR filter with a bandwidth of 1 - 8 Hz to remove high-frequency noise and baseline drift.

[0093] 1.3 For the EDA and SKT signals, use a 4th-order Butterworth IIR low-pass filter with a cut-off frequency of 5 Hz to remove high-frequency noise.

[0094] Step 2: Self-supervised pre-training based on pseudo-labels

[0095] 2.1 Perform amplitude and time transformations on the three-modal signals to generate corresponding transformed signals and corresponding transformed category labels as pseudo-labels for the pre-training task. These transformations include:

[0096] 2.1.1 Add Gaussian noise to the original signal, control the signal-to-noise ratio SNR = 50, and the pseudo-label is noise class 1;

[0097] 2.1.2 Invert the original signal around the mean, and the pseudo-label is amplitude inversion class 2;

[0098] 2.1.3 Randomly set 25% of the data of the original signal to 0, and the pseudo-label is mask class 3;

[0099] 2.1.4 Reverse the output of the original signal in the time dimension, and the pseudo-label is time reverse order class 4;

[0100] 2.1.5 The original signal is divided into n non - overlapping segments, and then these segments are shuffled in chronological order and finally recombined together, with the pseudo - label being segment - shuffled class 5;

[0101] 2.1.6 The original signal is divided into n non - overlapping segments. A part of the segments are randomly selected and stretched through a linear interpolation function, and the remaining ones are squeezed through downsampling. The time - warped signal can be concatenated by the transformed segments and finally adjusted to the original length, with the pseudo - label being segment - warped class 6;

[0102] 2.1.7 The original signal is divided into n non - overlapping segments, and one of them is randomly selected and resampled to the original length, with the pseudo - label being segment - selected class 7.

[0103] 2.2 The original signal is given the pseudo - label original - signal class 0. Together with the 7 transformed signals and pseudo - labels described in 2.1 above, it is fed into the time - frequency information fusion rU - Net Encoder for the pre - training task of feature extraction. The specific structure and operation steps of this model are as follows:

[0104] 2.2.1 The three - modality original signals are combined into a data matrix of (batch, time, channel) by increasing the number of channels.

[0105] 2.2.2 The data matrix is fed into a residual connection block (ResNetBlock, RSB) composed of 5 1D convolutional layers (1D CNN). The structures of these 5 1D CNNs are as Figure 4 shown.

[0106] 2.2.3 The max - pooling layer is used to downsample the scale features output by the RSB in 2.2.2.

[0107] 2.2.4 Repeat the process from 2.2.2 to 2.2.3 until the feature extraction of 5 scales is completed. The structures of the 5 - scale RSBs are the same, and the selected numbers of CNN filters are (16, 32, 64, 128, 256) respectively. The pooling factors of the max - pooling layers between every two - scale RSBs are (8, 8, 5, 4) respectively.

[0108] 2.2.5 Parallelly, the three - modality original signals are used to extract the amplitude - frequency characteristic curves through the fast Fourier transform and combined into a frequency - domain data matrix of (batch, frequency, channel) by increasing the number of channels.

[0109] 2.2.6 The frequency - domain data matrix processed in 2.2.5 is fed into a neural network with the same structure as that for processing the signal time series, and repeat steps 2.2.2 to 2.2.4.

[0110] 2.2.7 The data processed in 2.2.4 has a scale of 1 in the second dimension of the deepest scale. The time-domain and frequency-domain features at the deepest scale are fed into the feature fusion network for multi-scale time-frequency feature fusion. The features at the deepest scale (i.e., the first-scale fusion features described above) can be obtained by linearly transforming the upsampled frequency-domain features and time-domain features through a fully connected layer and then processing them through an activation function.

[0111] 2.2.7.1 In the output feature dimension of the deepest scale, the time-domain features and frequency-domain features are concatenated along the last dimension.

[0112] 2.2.7.2 Through the upsampling layer, the output of 2.2.7.1 is upsampled along the second dimension, i.e., the time dimension.

[0113] 2.2.7.3 Through the fully connected layer, the output of 2.2.7.2 is feature-processed and compressed, and the information in the third dimension of the data matrix is transferred to the second dimension, thus completing the information fusion at the deepest scale.

[0114] 2.2.8 The time-domain and frequency-domain features of each scale output by 2.2.7 and its sub-steps are respectively fused. The specific fusion structure is as follows:

[0115] 2.2.8.1 The output features of the 5 scales are respectively upsampled to restore the size of the second dimension of the input data matrix, and the fully connected layer is used to compress the data features after upsampling restoration.

[0116] 2.2.8.2 The features output in 2.2.8.1 are concatenated along the feature dimension, and then the multi-head attention layer is used for multi-scale feature fusion and the output features are strengthened according to the attention scores, and the number of heads of the multi-head attention mechanism is configured.

[0117] 2.2.9 The time-domain and frequency-domain features output by 2.2.8 and its sub-steps are respectively merged using the fully connected layer, and the number of units of the fully connected layer is set.

[0118] 2.2.10 The features output in 2.2.9 are fed into the fully connected layer with the SoftMax activation function to complete the classification task of the pseudo labels. The dimension of the output layer is set according to the classification requirements, and the loss function is the multi-class cross-entropy loss function.

[0119] Step 3: Fine-tuning based on the emotion recognition task

[0120] 3.1 Use the network weights after completing Step 2 as the initial network weights of this step, and input the original signal data of the three modalities into the network model according to the input dimensions in Step 2.2.

[0121] Replace the units of the fully connected layer in step 2.2.10 with the number of target classifications for emotion recognition, and use the cross-entropy loss function to fine-tune the emotion recognition task.

[0122] In some examples of the present application, the emotion recognition model is obtained by training in the following manner: Obtain a training data set, where the training data set includes: multimodal physiological signals and corresponding emotion categories; Input the training data set into a deep learning model, and use the emotion category as a label to fine-tune the deep learning model to obtain the emotion recognition model. Based on the network architecture of the present application ( Figure 3 ) Only fine-tuning can effectively reduce the training time and consumption of computing resources, and quickly adapt to the needs of specific scenarios.

[0123] For ease of understanding, the present application uses public data sets (such as CASE, DEAP) as examples to detail the training process of the emotion recognition model of the present application. Among them, the training process includes: fine-tuning.

[0124] Step 1: Data preprocessing

[0125] 1.1 Use three physiological signals of PPG, EDA, and SKT. Resample the signals to 64 Hz, with a window of 20 s and an overlap of 5 s, and slide and truncate.

[0126] 1.2 The PPG signal uses a band-pass Hamming window FIR filter with a bandwidth of 1 - 8 Hz to remove high-frequency noise and baseline drift.

[0127] 1.3 The EDA and SKT signals use a 4th-order Butterworth IIR low-pass filter with a cut-off frequency of 5 Hz to remove high-frequency noise.

[0128] Step 2: Fine-tuning based on the emotion recognition task

[0129] 2.1 Input the multimodal signal data preprocessed in step 1 into the network model, and use the cross-entropy loss function to fine-tune the emotion recognition task.

[0130] It can be understood that the training set and validation set of the emotion recognition model training method of the present application are selected from the CASE data set and the DEAP data set, and there is no intersection between the training set and the validation set.

[0131] It should be understood that the emotion categories involved in the foregoing two model training methods include: calm, pleasant, anxious, angry, sad, etc.

[0132] The above-mentioned emotion recognition method is based on the fusion of frequency-domain information and time-domain information features, ensuring the comprehensive extraction and feature fusion of multi-domain information of quasi-periodic signals; the size of the second dimension of the deepest scale feature obtained by the per-scale maximum pooling of the rUnet structure is the target dimension, and the design of the time-frequency feature cross-fusion module based on the output features of this layer avoids the conflict between the time semantics and frequency semantics of the time domain and frequency domain dimensions, and at the same time ensures the deep fusion of time-frequency domain features; secondly, the multi-scale feature fusion module ensures that the CNN obtains a larger receptive field, takes into account the characteristics of the signal change speed of different modal signals, performs information fusion at different scales, takes into account local and global features, and improves the model expression ability; the construction and use of the emotion recognition model based on multi-modal physiological signals by increasing the number of channels avoids the model expansion caused by the stacking of feature extraction encoders of different modalities, making the model smaller and the training process simpler.

[0133] In another aspect of the present application, the present application proposes an emotion recognition device 500, which includes: a signal acquisition module 510 and an emotion recognition module 520. Among them,

[0134] The signal acquisition module 510 is used to acquire multi-modal physiological signals of a target object; the emotion recognition module 520 is used to input the physiological signals into an emotion recognition model to obtain an emotion recognition result of the target object; wherein, the emotion recognition model is used to: respectively reduce the dimensions of the time-domain signal and the frequency-domain signal corresponding to the physiological signals to obtain the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal; fuse the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal to obtain a first-scale fusion feature; fuse the first-scale fusion feature with the time-domain signal and the frequency-domain signal before dimension reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; and obtain the emotion recognition result of the target object based on the fourth-scale fusion feature.

[0135] In some examples of the present application, the dimensions of the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal are both the target dimension.

[0136] In some examples of the present application, the step of fusing the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal to obtain a first-scale fusion feature includes: fusing the weights of the fully connected layer, the frequency-domain features, and the time-domain features based on an activation function to obtain a first-scale fusion feature.

[0137] In some examples of the present application, fusing the full-connection layer weights with the frequency-domain features and time-domain features based on an activation function to obtain first-scale fused features includes: upsample the frequency-domain features and time-domain features in the time dimension to obtain the upsampled time-domain features and frequency-domain features; use the activation function to fuse the full-connection layer weights with the upsampled time-domain features and frequency-domain features to obtain the first-scale fused features.

[0138] In some examples of the present application, the multi-modal physiological signals are selected from multiple types including optoelectronic signals, electrocardiogram signals, galvanic skin response signals, and skin temperature signals.

[0139] In some examples of the present application, the emotion recognition model is trained in the following manner: obtain a training data set, where the training data set includes: multi-modal physiological signals and corresponding emotion categories; use data augmentation methods to amplify the multi-modal physiological signals to obtain amplified multi-modal physiological signals; input the amplified multi-modal physiological signals into a deep learning model, and use the data augmentation methods as pseudo-labels to perform self-supervised learning on the deep learning model to obtain a pre-trained model; input the training data set into the pre-trained model, and use the emotion categories as labels to fine-tune the pre-trained model to obtain the emotion recognition model.

[0140] In some examples of the present application, the data augmentation methods include at least one of noise processing, signal transformation, and timing adjustment. Among them, the signal transformation includes at least one of amplitude inversion and segment distortion, and the timing adjustment includes at least one of time dimension reverse order and segment scrambling.

[0141] In some examples of the present application, the emotion recognition model is trained in the following manner: obtain a training data set, where the training data set includes: multi-modal physiological signals and corresponding emotion categories; input the training data set into a deep learning model, and use the emotion categories as labels to fine-tune the deep learning model to obtain the emotion recognition model.

[0142] The aforementioned device is based on the feature fusion of frequency-domain information and time-domain information, ensuring the comprehensive extraction and feature fusion of multi-domain information of quasi-periodic signals. The second dimension size of the deepest scale feature obtained by the progressive scale max pooling of the rUnet structure is the target dimension. Based on the output features of this layer, the design of the time-frequency feature cross-fusion module avoids the conflict between the time semantics and frequency semantics of the time domain and frequency domain dimensions, and at the same time ensures the deep fusion of time-frequency domain features. Secondly, the multi-scale feature fusion module ensures that the CNN obtains a larger receptive field, takes into account the signal change speed characteristics of different modality signals itself, performs information fusion at different scales, takes into account local and global features, and improves the model expression ability. The construction and use of the emotion recognition model based on multi-modal physiological signals by increasing the number of channels avoids the model expansion caused by the stacking of feature extraction encoders of different modalities, making the model smaller in size and the training process simpler.

[0143] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 5 The device 500 shown can execute Figure 2 the corresponding method embodiments, and the foregoing and other operations and / or functions of each module in the device 500 are respectively to implement Figure 2 the corresponding processes in each method in, and for the sake of brevity, they will not be elaborated here.

[0144] In the above, the device 500 of the embodiments of the present application has been described from the perspective of functional modules in conjunction with the drawings. It should be understood that the functional modules can be implemented in hardware form, can also be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0145] In another aspect of the present application, the present application proposes an electronic device.

[0146] The so-called electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computing device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.

[0147] Reference Figure 6 , the electronic device 600 may be an execution device of the above method, but is not limited thereto. As Figure 6 shown, the electronic device 600 may include:

[0148] A memory 610 and a processor 620, the memory 610 is used to store a computer program 630 and transmit the computer program 630 to the processor 620. In other words, the processor 620 can call and run the computer program 630 from the memory 610 to implement the method in the embodiments of the present application.

[0149] For example, the processor 620 can be used to execute the steps in the above method according to the instructions in the computer program 630.

[0150] In some embodiments of the present application, the processor 620 may include but is not limited to:

[0151] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0152] In some embodiments of the present application, the memory 610 includes but is not limited to:

[0153] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0154] In some embodiments of the present application, the computer program 630 can be divided into one or more modules, and the one or more modules are stored in the memory 610 and executed by the processor 620 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 630 in the electronic device.

[0155] As Figure 6 shown, the electronic device 600 may further include:

[0156] A transceiver 660, which can be connected to the processor 620 or the memory 610.

[0157] Among them, the processor 620 can control the transceiver 660 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 660 can include a transmitter and a receiver. The transceiver 660 may further include an antenna, and the number of antennas can be one or more.

[0158] It should be understood that the various components in the electronic device 600 are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.

[0159] According to one aspect of the present application, there is provided a computer-readable storage medium, on which computer instructions or programs are stored. When the computer instructions or programs are executed by a computer, the computer is enabled to execute the methods in the above method embodiments. Or rather, the embodiments of the present application further provide a computer program product containing instructions. When the instructions are executed by a computer, the computer is enabled to execute the methods in the above method embodiments.

[0160] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the methods in the above method embodiments.

[0161] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0162] Those of ordinary skill in the art will realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0163] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be electrical, mechanical, or other forms.

[0164] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0165] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0166] The embodiments of this application will be described in more detail below. The examples of the embodiments are shown in the drawings. The embodiments described below by referring to the drawings are exemplary and are intended to explain this application and should not be construed as limiting this application.

[0167] Embodiment 1: Model Performance Comparison

[0168] Dataset

[0169] This embodiment uses two datasets, the CASE dataset and the DEAP dataset. Among them, the CASE dataset is a real-time method that continuously samples at a frequency of 20 Hz during the entire 8-film stimulation process, and its physiological data is sampled at a frequency of 1 kHz. Since the CASE dataset uses a continuous annotation method, this embodiment evaluates two annotation strategies: (1) overall annotation and (2) fine-grained annotation. The overall label uses the average value of the trajectory as the basis for classification, while the fine-grained label uses the average value of the annotation values in a 20-second segment. The DEAP dataset adopts a classical post-stimulation annotation method dataset and only tests the overall label as the basis for emotion classification. The physiological data in the DEAP dataset was initially sampled at a frequency of 128 Hz. Given the differences in sampling frequency and annotation strategy, the goal of this embodiment is to test the method based on the specific implementation of this application on these two datasets to verify its effectiveness and robustness in different environments.

[0170] Comparison model

[0171] This embodiment compares the emotion recognition model of this application with two classical models. Both of these classical models are pre-trained using self-supervised learning and fine-tuned for emotion classification using the same modality as the method of this application. These two models do not introduce frequency-domain feature extraction and multi-scale information fusion, and both of these methods are built for multi-modal encoder stacking. Their specific descriptions are as follows:

[0172] TCNFomer: A novel self-supervised learning framework that achieves efficient multi-modal fusion through a modality-specific encoder based on temporal convolution and a shared encoder based on transformers.

[0173] SigRep: It adopts a model architecture similar to SimCL, including an encoder composed of four heuristic modules and a unit for each quantity used in the state equation.

[0174] The fine-tuned emotion recognition model (abbreviated as "fine-tuning") and the pre-trained and fine-tuned emotion recognition model (abbreviated as "pre-trained - fine-tuning") involved in the specific implementation part of this application are compared with the above two classical models on the CASE dataset and the DEAP dataset. The results are shown in Tables 1 - 3. On these two datasets, the performance of the model of this application is better than the baseline method, with an average relative improvement of 3%, which proves the effectiveness of including frequency-domain information. Through self-supervised learning pre-training, the accuracy of the model of this application is further improved by about 2% on average, indicating that self-supervised learning can effectively capture emotion-related information in physiological signals.

[0175] As shown in Table 1, the comparison results indicate that the overall label has a higher average accuracy, which may be due to the smoothing effect and the shortened time lag between emotional responses and self-reports. In contrast, as shown in Table 3, the performance of the DEAP dataset is relatively low, mainly for two reasons. First, compared with the CASE dataset, the sampling rate of peripheral signals in the DEAP dataset is lower, which may lead to some information loss. Second, although the subjects' valence-arousal (V-A) ratings for some videos (such as sadness and anger) are similar, the DEAP dataset contains 40 videos, far more than the 8 videos in the CASE dataset. These videos may elicit more diverse physiological responses, ultimately making it more challenging to accurately identify emotions from peripheral signals.

[0176] Table 1 Overall Emotion Recognition Accuracy of the CASE Dataset (%)

[0177]

[0178] Note: Valance and Arousal are two labels for emotion evaluation. Valance represents the positive-negative tendency, and Arousal represents the degree of emotional activation. Among them, -2 indicates binary classification (high and low), and -3 indicates ternary classification (high, medium, and low).

[0179] Acc represents accuracy.

[0180] W-F1: Weighted-F1 score.

[0181] Table 2 Fine-Grained Emotion Recognition Accuracy of the CASE Dataset (%)

[0182]

[0183] Note: Valance and Arousal are two labels for emotion evaluation. Valance represents the positive-negative tendency, and Arousal represents the degree of emotional activation. Among them, -2 indicates binary classification (high and low), and -3 indicates ternary classification (high, medium, and low).

[0184] Acc represents accuracy.

[0185] W-F1: Weighted-F1 score.

[0186] Table 3 Emotion Recognition Accuracy of the DEAP Dataset (%)

[0187]

[0188] Note: Valance and Arousal are two labels for emotion evaluation. Valance represents the positive-negative tendency, and Arousal represents the emotional activation level. Among them, -2 indicates a two-category (high or low) classification, and -3 indicates a three-category (high, medium, or low) classification;

[0189] Acc represents accuracy;

[0190] W-F1: Weighted-F1 score.

[0191] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0192] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application without departing from the principles and purposes of this application.

Claims

1. A method for emotion recognition, characterized in that, Including: Obtain multimodal physiological signals of a target object; Input the physiological signals into an emotion recognition model to obtain an emotion recognition result of the target object; Wherein, the emotion recognition model is used to: Reduce the dimensions of the time-domain signal and the frequency-domain signal corresponding to the physiological signals respectively to obtain the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal; Fuse the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal to obtain a first-scale fusion feature; Fuse the first-scale fusion feature with the time-domain signal and the frequency-domain signal before dimension reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; Fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; Based on the fourth-scale fusion feature, obtain the emotion recognition result of the target object.

2. The method according to claim 1, wherein The dimensions of the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal are both target dimensions.

3. The method according to claim 2, wherein The step of fusing the reduced-dimensional time-domain signal and the reduced-dimensional frequency-domain signal to obtain a first-scale fusion feature includes: Based on an activation function, fuse the fully connected layer weights, the frequency-domain features, and the time-domain features to obtain a first-scale fusion feature.

4. The method according to claim 3, wherein The step of fusing the fully connected layer weights, the frequency-domain features, and the time-domain features based on the activation function to obtain a first-scale fusion feature includes: Upsample the frequency-domain features and the time-domain features in the time dimension to obtain the upsampled time-domain features and frequency-domain features; Use the activation function to fuse the fully connected layer weights and the upsampled time-domain features and frequency-domain features to obtain a first-scale fusion feature.

5. The method according to any one of claims 1-4, characterized in that, The multimodal physiological signals are selected from multiple types including optoelectronic signals, electrocardiogram signals, galvanic skin response signals, and skin temperature signals.

6. The method according to claim 5, wherein The emotion recognition model is obtained through the following training method: Obtain a training data set, where the training data set includes: multimodal physiological signals and corresponding emotion categories; Amplify the multimodal physiological signals by using a data augmentation method to obtain amplified multimodal physiological signals; Input the amplified multimodal physiological signals into a deep learning model, and use the data augmentation method as a pseudo label to perform self-supervised learning on the deep learning model to obtain a pre-trained model; Input the training data set into the pre-trained model, and use the emotion category as a label to fine-tune the pre-trained model to obtain the emotion recognition model.

7. The method according to claim 6, characterized in that, The data augmentation method includes at least one of noise processing, signal transformation, and timing adjustment. Among them, the signal transformation includes at least one of amplitude inversion and segment distortion, and the timing adjustment includes at least one of time dimension reverse order and segment scrambling.

8. The method according to claim 5, characterized in that, The emotion recognition model is obtained through the following training method: Obtain a training data set, where the training data set includes: multimodal physiological signals and corresponding emotion categories; Input the training data set into a deep learning model, and use the emotion category as a label to fine-tune the deep learning model to obtain the emotion recognition model.

9. An emotion recognition device, characterized in that, Including: A signal acquisition module, configured to acquire multi-modal physiological signals of a target object; An emotion recognition module, configured to input the physiological signals into an emotion recognition model to obtain an emotion recognition result of the target object; Wherein, the emotion recognition model is configured to: Reduce the dimensions of the time-domain signal and the frequency-domain signal corresponding to the physiological signals respectively to obtain the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal; Fuse the reduced-dimension time-domain signal and the reduced-dimension frequency-domain signal to obtain a first-scale fusion feature; Fuse the first-scale fusion feature with the time-domain signal and the frequency-domain signal before dimension reduction respectively to obtain a second-scale fusion feature and a third-scale fusion feature; Fuse the second-scale fusion feature and the third-scale fusion feature to obtain a fourth-scale fusion feature; Based on the fourth-scale fusion feature, obtain the emotion recognition result of the target object.

10. An electronic device, characterized in that, It includes: A processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program to implement the emotion recognition method according to any one of claims 1 to 8.