Edge cloud cooperation system and method for emotion recognition
By deploying emotion recognition algorithms on edge computing devices and using dual-channel feature extraction backbone network structures and adaptive cross-domain information fusion modules, the real-time, accuracy and privacy protection problems of emotion recognition in the prior art are solved, and efficient and real-time emotion recognition is achieved.
Patent Information
- Application Number
- CN202510086293.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to achieve efficient, real-time and accurate identification in offline state in emotional recognition, and there are problems with data transmission delay and privacy protection.
By deploying the emotion recognition algorithm on the edge computing device, preprocessing and feature extraction of the collected EEG signals is realized, and dual-channel feature extraction backbone network structure and adaptive cross-domain information fusion module are used for emotion recognition.
Real-time identification of EEG signals is realized offline or online, reducing data transmission delay, improving identification accuracy and efficiency, and enhancing data privacy protection.
Smart Images

Figure CN120105141A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroencephalogram (EEG) information processing, and in particular to an edge-cloud collaborative system and method for emotion recognition. Background Art
[0002] Emotion recognition is a key link in realizing emotional brain-computer interface, and has important application value in human-computer interaction, medical health, traffic safety and other fields. EEG signals can reflect the electrophysiological activities of brain neurons. As a spontaneous physiological signal, it has the advantages of strong objectivity and difficulty in forging. At the same time, the brain plays an important role in the process of individual emotion regulation. Therefore, EEG signals have gradually become an important tool for emotion recognition. With the development of the Internet of Things and information technology, edge computing, as an emerging computing paradigm, provides new possibilities for efficient and real-time emotion recognition. Edge computing can perform real-time data analysis and decision-making at the place where the data is generated, without having to transmit the data to the cloud data center for processing. By performing calculations at a location closer to the data source, the delay of data transmission is reduced, and the efficiency and security of data processing are improved. At the same time, edge computing devices allow certain computing tasks to be performed offline without relying on a continuous network connection. Emotion recognition can be performed even in an environment where the network is unstable or completely offline, and can adapt to the needs of different network environments. In addition, the offline processing capability of edge devices also helps to protect the privacy of subjects. EEG signals belong to the personal privacy of users. Processing EEG signal data locally reduces the risk of data leakage. As artificial intelligence has made remarkable achievements in many fields such as image recognition and natural language processing, emotion recognition technology based on deep learning has gradually become a research hotspot and has achieved a series of innovative results. Combining edge computing with deep learning technology can further improve the accuracy and efficiency of emotion recognition, thereby providing users with a more convenient service experience. Summary of the invention
[0003] In order to solve the problems of the prior art, the present invention provides an edge-cloud collaborative system and method for emotion recognition. The present invention deploys an emotion recognition algorithm on an edge computing device to realize preprocessing and feature extraction of the collected EEG signals of the subject in an offline or online state, and automatically evaluates the emotional state of the subject based on the designed emotion recognition method.
[0004] The present invention adopts the following technical solution:
[0005] An edge-cloud collaborative system for emotion recognition, the system comprising an edge computing device and a cloud server; the edge computing device comprises a data acquisition module, a data preprocessing module, a feature extraction module, an emotion recognition module,
[0006] The edge storage module, the edge interaction module, the device management module and the first data transmission module, the cloud server includes a cloud computing module, a cloud storage module, and a second data transmission module, wherein:
[0007] Feature extraction module: obtains the local user's EEG data after preprocessing from the edge storage module, extracts EEG signal features using signal processing methods, and transmits the extracted EEG signal feature data to the edge storage module for storage;
[0008] Emotion recognition module: loads the extracted user EEG signal feature data and local emotion recognition model from the edge storage module, and outputs the emotion type inferred locally to the edge storage module for storage;
[0009] Edge storage module: used to store and manage data collected, processed and analyzed by other modules of the edge computing device and data transmitted from the cloud server via the first data transmission module;
[0010] The edge interaction module is used to provide a visual graphical operation interface to guide users to input or confirm their basic information during data collection, and to display the recognition results on the graphical interface after the recognition is completed. Users can provide feedback through the edge interaction module; wherein:
[0011] The cloud server initializes the training of the emotion recognition model and transmits the initial model to the edge storage module of each edge computing device via the second data transmission module;
[0012] After the edge computing device adds user data, the data is integrated and uploaded to the cloud server via the first data transmission module;
[0013] Based on the distributed user data uploaded by each edge computing device, the cloud server uses the integrated EEG data to retrain the emotion recognition model to obtain a global model, and transmits the global model to each edge computing device;
[0014] Each edge computing device updates the local emotion recognition model based on the global model and repeats the above process to improve the accuracy of the emotion recognition model.
[0015] Furthermore, the emotion recognition module adopts a dual-path feature extraction backbone network structure, which includes a frequency domain unit, a time domain unit, a first lightweight module, a second lightweight module and an information fusion module; wherein:
[0016] The time domain unit extracts features from the time domain EEG signal according to the following formula to obtain time domain EEG feature data;
[0017] x D=ReLU(BN(DWConv(x)))
[0018] x P =ReLU(BN(PWConv(x D )))
[0019] Where x represents the input EEG signal, DWConv represents the deep convolution operation, BN represents the batch normalization operation, ReLU represents the activation function, and x D is the feature obtained through deep convolution related operations, PWConv represents point-by-point convolution operation, x P It is the feature obtained through point-by-point convolution related operations;
[0020] The frequency domain unit extracts the features of the time domain EEG signal according to the following formula to obtain the frequency domain EEG feature data;
[0021] x fD =ReLU(BN(DWConv(FFT(x))))+x
[0022] x fP =ReLU(BN(PWConv(x fD )))
[0023] Where x represents the input EEG signal, FFT represents the fast Fourier transform processing, and x fD Represents the intermediate features of frequency domain information obtained after deep convolution operation and residual connection, x fP It is the feature obtained by performing point-by-point convolution operations on the intermediate features;
[0024] The first lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the first emotion feature data;
[0025] z o =LSA(DWConv(z))
[0026] z d1 =LSA(Down 1 (DWConv(z)))
[0027] z d2 =LSA(Down 2 (DWConv(z)))
[0028] z m =z o +Up 1 (z d1 )+Up 2 (z d2 )+z
[0029] zl1 =MLP(DWConv(z m ))+z m
[0030] Among them, z represents the input time domain EEG feature data, LSA represents the lightweight self-attention module, and z o Indicates the result obtained by the input data through the original feature path, Down 1 represents a 4x downsampling module, z d1 Indicates the result obtained by downsampling the input data by 4 times. 2 represents an 8-fold downsampling module, z d2 Indicates the result obtained by downsampling the input data by 8 times. 1 Indicates 4x upsampling module, Up 2 Indicates 8x upsampling module, z m It represents the intermediate features after multi-scale feature extraction, and MLP represents multi-layer perceptron;
[0031] The second lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the second emotion feature data;
[0032] z o =LSA(DWConv(z))
[0033] z d1 =LSA(Down 1 (DWConv(z)))
[0034] z d2 =LSA(Down 2 (DWConv(z)))
[0035] z m =z o +Up 1 (z d1 )+Up 2 (z d2 )+z
[0036] z l2 =MLP(DWConv(z m ))+z m
[0037] The information fusion module adopts an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain the emotion type.
[0038] Furthermore, the first lightweight module and the second lightweight module are both composed of a Patch Embedding layer, a PatchMerging layer and a Transformer layer; the lightweight self-attention module is arranged in the Transformer layer;
[0039] The lightweight self-attention module includes a segmentation unit, a self-attention unit and a position encoding unit; wherein:
[0040] The segmentation unit performs feature segmentation according to the following formula;
[0041] X att ,X res =Spi lt(X)
[0042] Where: X represents the input feature of the lightweight self-attention module, Spilt represents the segmentation operation, and X att With X res Respectively represent the features used to input self-attention operation and residual operation;
[0043] The self-attention unit calculates the first emotion feature data and the second emotion feature data according to the following formula to obtain the deep emotion feature:
[0044] X satt =Se l f-Attention(DWConv(X att W Q ,X att W K ,X att W V ))
[0045] Where: W Q , W K , W V Respectively represent the weight matrices used to calculate Q, K, and V, and Self-Attention represents the self-attention operation;
[0046] The position encoding unit processes the deep emotion feature according to the following formula to obtain the global deep emotion feature;
[0048] X pe =X satt +LPE(X att W V )
[0049] X O =X pe +X res
[0050] Where: LPE represents the learnable position encoding operation, X pe Represents the features obtained after the self-attention operation is combined with the learnable position encoding.
[0051] Furthermore, the information fusion module is composed of an embedding layer, a spatial attention layer, a channel attention layer, a gating unit and a multi-layer sensor; the information fusion module adopts an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain an emotion type process, and then fuses the local emotion type process, including:
[0052] The embedding layer performs dimension conversion on the first emotion feature data and the second emotion feature data according to the following formula to obtain the first emotion feature representation and the second emotion feature representation;
[0053] X e1 =Embedding 1 (X 1 )
[0054] X e2 =Embedding 2 (X 2 )
[0055] The spatial attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion space feature and the second emotion space feature;
[0056] X s1 =σ(Conv([AvgPool(X e1 ));MaxPool(X e1 )]))
[0057] X s2 =σ(Conv([AvgPool(X e2 ));MaxPool(X e2 )]))
[0058] The channel attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion channel feature and the second emotion channel feature;
[0059] X c1 =σ(MLP(AvgPool(X e1 ))+MLP(MaxPoo l(X e1 )))
[0060] X c2 =σ(MLP(AvgPool(X e2))+MLP(MaxPoo l(X e2 )))
[0061] The gating unit processes the first emotion space feature and the first emotion channel feature according to the following formula to obtain the first local emotion feature:
[0062] X f1 =σ(W c1 X c1 +W s1 X s1 +b 1 )
[0063] Where: σ represents the activation function, W c1、 W s1 are all learnable weights in the gated unit, b 1 represents the bias term;
[0064] The gate control unit processes the second emotion space feature and the second emotion channel feature according to the following formula to obtain the second local emotion feature:
[0065] X f2 =σ(W c2 X c2 +W s2 X s2 +b 2 )
[0066] Where: σ represents the activation function, W c2、 W s2 are all learnable weights in the gated unit, b 2 represents the bias term;
[0067] The multi-layer perceptron fuses the first local emotion feature and the second local emotion feature according to the following formula to obtain the emotion type:
[0068] X af =MLP(X f1 +X f2 ).
[0069] The present invention can also be implemented by the following technical solutions:
[0070] A method for emotion recognition using an edge-cloud coordination system comprises the following steps:
[0071] Obtaining the pre-processed local user EEG data from the edge storage module, extracting EEG signal features using a signal processing method, and transmitting the extracted EEG signal feature data to the edge storage module for storage;
[0073] Load the extracted user EEG signal feature data and the local emotion recognition model from the edge storage module, and output the emotion type inferred locally to the edge storage module for storage; wherein:
[0074] The time domain unit extracts features from the time domain EEG signal according to the following formula to obtain time domain EEG feature data;
[0075] x D =ReLU(BN(DWConv(x)))
[0076] x P =ReLU(BN(PWConv(x D )))
[0077] Where x represents the input EEG signal, DWConv represents the deep convolution operation, BN represents the batch normalization operation, ReLU represents the activation function, and x D is the feature obtained through deep convolution related operations, PWConv represents point-by-point convolution operation, x P It is the feature obtained through point-by-point convolution related operations;
[0078] The frequency domain unit extracts features from the time domain EEG signal according to the following formula to obtain frequency domain EEG feature data;
[0079] x fD =ReLU(BN(DWConv(FFT(x))))+x
[0080] x fP =ReLU(BN(PWConv(x fD )))
[0081] Where x represents the input EEG signal, FFT represents the fast Fourier transform processing, and x fD Represents the intermediate features of frequency domain information obtained after deep convolution operation and residual connection, x fP It is the feature obtained by performing point-by-point convolution operations on the intermediate features;
[0082] The first lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the first emotion feature data;
[0083] z o =LSA(DWConv(z))
[0084] z d1 =LSA(Down 1 (DWConv(z)))
[0085] z d2 =LSA(Down2 (DWConv(z)))
[0086] z m =z o +Up 1 (z d1 )+Up 2 (z d2 )+z
[0087] z l1 =MLP(DWConv(z m ))+z m
[0088] Among them, z represents the input time domain EEG feature data, LSA represents the lightweight self-attention module, and z o Indicates the result obtained by the input data through the original feature path, Down 1 represents a 4x downsampling module, z d1 Indicates the result obtained by downsampling the input data by 4 times. 2 represents an 8-fold downsampling module, z d2 Indicates the result obtained by downsampling the input data by 8 times. 1 Indicates 4x upsampling module, Up 2 Indicates 8x upsampling module, z m It represents the intermediate features after multi-scale feature extraction, and MLP represents multi-layer perceptron;
[0089] The second lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the second emotion feature data;
[0090] z o =LSA(DWConv(z))
[0091] z d1 =LSA(Down 1 (DWConv(z)))
[0092] z d2 =LSA(Down 2 (DWConv(z)))
[0093] z m =z o +Up 1 (z d1 )+Up 2 (z d2 )+z
[0094] z l2 =MLP(DWConv(z m ))+zm
[0095] The information fusion module uses an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain the emotion type;
[0096] Edge storage module: used to store and manage data collected, processed and analyzed by other modules of the edge computing device and data transmitted from the cloud server via the first data transmission module; wherein:
[0097] The cloud server initializes the training of the emotion recognition model and transmits the initial model to the edge storage module of each edge computing device via the second data transmission module;
[0098] After the edge computing device adds user data, the data is integrated and uploaded to the cloud server via the first data transmission module;
[0099] Based on the distributed user data uploaded by each edge computing device, the cloud server uses the integrated EEG data to retrain the emotion recognition model to obtain a global model, and transmits the global model to each edge computing device;
[0100] Each edge computing device updates the local emotion recognition model based on the global model and repeats the above process to improve the accuracy of the emotion recognition model.
[0101] Furthermore, the information fusion module adopts an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain an emotion type process, including:
[0102] The embedding layer performs dimension conversion on the first emotion feature data and the second emotion feature data according to the following formula to obtain the first emotion feature representation and the second emotion feature representation;
[0103] X e1 =Embedding 1 (X 1 )
[0104] X e2 =Embedding 2 (X 2 )
[0105] The spatial attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion space feature and the second emotion space feature;
[0106] X s1 =σ(Conv([AvgPool(X e1 ));MaxPool(Xe1 )]))
[0107] X s2 =σ(Conv([AvgPool(X e2 ));MaxPool(X e2 )]))
[0108] The channel attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion channel feature and the second emotion channel feature;
[0109] X c1 =σ(MLP(AvgPool(X e1 ))+MLP(MaxPoo l(X e1 )))
[0110] X c2 =σ(MLP(AvgPool(X e2 ))+MLP(MaxPoo l(X e2 )))
[0111] The gating unit processes the first emotion space feature and the first emotion channel feature according to the following formula to obtain the first local emotion feature:
[0112] X f1 =σ(W c1 X c1 +W s1 X s1 +b 1 )
[0113] Where: σ represents the activation function, W c1、 W s1 are all learnable weights in the gated unit, b 1 represents the bias term;
[0114] The gate control unit processes the second emotion space feature and the second emotion channel feature according to the following formula to obtain the second local emotion feature:
[0115] X f2 =σ(W c2 X c2 +W s2 X s2 +b 2 )
[0116] Where: σ represents the activation function, W c2、 W s2 are all learnable weights in the gated unit, b 2 represents the bias term;
[0117] The multi-layer perceptron fuses the first local emotion feature and the second local emotion feature according to the following formula to obtain the emotion type:
[0118] X af =MLP(X f1 +X f2 ). Beneficial Effects
[0119] 1. Considering that EEG signals contain information related to emotional states in both the time domain and the frequency domain, the present invention designs a dual-path feature extraction network based on the Transformer architecture, which is used to extract the features of EEG signals in the time domain and the frequency domain respectively. At the same time, in order to adapt to the limited computing resources of edge computing devices, a lightweight Transformer module is designed as the backbone feature extraction module to perform in-depth feature extraction from different scales. In addition, the designed adaptive cross-domain information fusion module realizes the dynamic fusion of time domain and frequency domain features, which can further enhance the model's ability to understand and express emotional features in EEG signals.
[0120] 2. The present invention realizes the real-time collection and analysis of the subject's EEG signals through edge computing devices, and by deploying emotion recognition algorithms on edge computing devices, it reduces the delay in data transmission to remote cloud servers and the dependence on cloud resources, thereby improving the overall system operation efficiency. The offline processing capability of edge computing devices enables emotion recognition even in environments with unstable or offline networks, can adapt to the needs of different network environments, and improves the flexibility of detection. Edge computing devices process sensitive EEG signal data locally, reducing the risk of data leakage during transmission to the cloud, and helping to protect the privacy and data security of subjects.
[0121] 3. The present invention can realize the collaborative working mode between cloud servers and edge computing devices. Cloud servers have larger computing resources and storage space, can integrate more data processing algorithms and emotion recognition models, and can undertake more complex and large-scale computing and analysis tasks. Edge computing devices have better real-time processing capabilities. The collaborative working mode can effectively improve the overall performance and efficiency of the system. Cloud servers can realize remote monitoring and management of edge computing devices, including remote upgrades, algorithm deployment, troubleshooting and other functions, and can flexibly adjust the distributed computing mode according to different application scenarios and needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0122] Figure 1 This is a schematic diagram of the structure of the edge computing-based emotion recognition system proposed in the present invention;
[0123] Figure 2 This is a schematic diagram of the structure of the lightweight emotion recognition method based on the Transformer architecture proposed in the present invention;
[0124] Figure 3 This is a schematic diagram of the structure of the lightweight Transformer module proposed in the present invention;
[0125] Figure 4 This is a schematic diagram of the structure of the lightweight self-attention module proposed in the present invention;
[0126] Figure 5 This is a schematic diagram of the structure of the adaptive cross-domain information fusion module proposed in the present invention. DETAILED DESCRIPTION
[0127] The following is combined with Figure 1 The present invention is described as follows:
[0128] This invention proposes an emotion recognition method and system based on edge computing, which is expected to achieve considerable social and economic benefits. The best implementation plan is to adopt patent transfer, technical cooperation or product development. This technology can be further developed in cooperation with hospitals, universities, etc. to develop products suitable for emotion recognition and EEG monitoring. This technology can be used in clinical diagnosis, mental health monitoring and other fields, and has important research significance and commercial value.
[0129] Emotion recognition system based on edge computing
[0130] Based on the designed emotion recognition method, the present invention provides an emotion recognition system based on edge computing, such as Figure 1 As shown, the system includes an edge computing device and a cloud server. The edge computing device includes a data acquisition module, a data preprocessing module, a feature extraction module, an emotion recognition module, an edge storage module, an edge interaction module, a device management module, and a first data transmission module, and the cloud server includes a cloud computing module, a cloud storage module, and a second data transmission module.
[0131] Data acquisition module: Use the edge interaction module to record the user's basic information, including but not limited to the user's age, gender, etc. In this module, a portable EEG signal acquisition device can be used to collect the subject's EEG signals, and the collected user data can be transmitted to the edge storage module for storage; when only edge computing devices are used for data collection, the collected user EEG data will be directly stored in the edge storage module, and can be subsequently transmitted to the cloud server for further analysis and processing. When the edge computing device is used to perform emotion recognition on the user, the recognition results can be directly obtained in the edge interaction module.
[0132] Data preprocessing module: obtains the collected local user EEG data from the edge storage module, and performs operations such as re-reference, downsampling, filtering, and artifact removal on the EEG data, extracts the required frequency band data, and transmits the preprocessed EEG data to the edge storage module for storage.
[0133] Feature extraction module: obtains the local user EEG data after preprocessing from the edge storage module, extracts the EEG signal features using signal processing methods, and transmits the extracted EEG signal feature data to the edge storage module for storage.
[0134] Emotion recognition module: Loads the extracted user EEG signal feature data and the local emotion recognition model from the edge storage module, inputs the extracted user EEG signal features into the emotion recognition model, and the model performs inference analysis and outputs the recognition results to the edge storage module for storage.
[0135] Edge storage module: used to store and manage data collected, processed and analyzed by other modules of the edge computing device and data transmitted from the cloud server via the first data transmission module;
[0136] Edge interaction module: used to provide a visual graphical operation interface to guide users to input or confirm their basic information during data collection. After the recognition is completed, the recognition results will be displayed on the graphical interface. Users can provide feedback through the edge interaction module. These feedbacks will be recorded by the system and used to optimize future detection and suggestions.
[0137] Device management module: used to monitor and manage the operating status of edge computing devices in real time, including computing resources, storage space, battery power, data transmission status and network connection status, record device work logs for subsequent task allocation and resource management, and store logs in the edge storage module; automatically detect hardware and software failures of the device when the device is abnormal, and provide troubleshooting guidance; after connecting to the cloud server, support remote configuration of the device's software and hardware; monitor and manage the data transmission process to ensure stable transmission of data between the edge storage module and the cloud server.
[0138] The first data transmission module is used to encrypt the data in the edge storage module and exchange data with the second data transmission module of the cloud server. The transmission methods include wireless transmission and wired transmission. During wireless transmission, the module will first evaluate the communication status of the current wireless connection. When the communication status meets the requirements, the data will be encrypted and transmitted to the second data transmission module to avoid data loss. The data obtained from the second data transmission module will be transmitted to the edge storage module for storage.
[0139] Cloud computing module: used to perform more complex and large-scale calculations and analyses on data transmitted from edge computing devices, integrate distributed user EEG data obtained from various edge computing devices, extract features from the integrated EEG data, retrain the model based on the emotion recognition model of the edge computing device using the integrated user data, and transmit the cloud-integrated user EEG data, extracted EEG signal feature data, and trained emotion recognition model to the cloud storage module for storage; when the edge computing device is stably connected to the cloud server by wireless or wired means, this module can work with the edge computing device to improve the overall performance and efficiency of the system. The specific process of edge-cloud collaboration is as follows:
[0140] 1. The cloud server initializes and trains the emotion recognition model, and transmits the initial model to the edge storage module of each edge computing device via the data transmission module 2;
[0141] 2. After adding user data to the edge computing device, the data is integrated and uploaded to the cloud server via the data transmission module 1; 3. Based on the distributed user data uploaded by each edge computing device, the cloud server uses the integrated EEG data to train the emotion recognition model again to obtain a global model, and transmits the global model to each edge computing device;
[0142] 4. Each edge computing device updates the local emotion recognition model based on the global model and repeats the above process to improve the accuracy of the emotion recognition model;
[0143] Cloud storage module: used to store data transmitted by the edge computing device via the second data transmission module and data processed and analyzed by the cloud computing module. This module is also responsible for data backup.
[0144] The second data transmission module is used to encrypt the data in the cloud storage module and exchange data with the first data transmission module of the edge computing device. The transmission methods include wireless transmission and wired transmission. During wireless transmission, the module will first evaluate the communication status of the current wireless connection. When the communication status meets the requirements, the data will be encrypted and transmitted to the first data transmission module to avoid data loss. The data obtained from the first data transmission module will be transmitted to the cloud storage module for storage.
[0145] In this embodiment, the computing process of each module of the edge computing device is implemented based on an embedded operating system. Security encryption technology is used throughout the process to protect user data privacy and prevent data leakage, security vulnerabilities and malicious attacks. When using edge computing devices and cloud servers, you need to log in to the system through an account and password. Both edge computing devices and cloud servers are equipped with strict data access control mechanisms, allowing only authorized users and applications to access data.
[0146] Edge computing devices process user EEG data locally, improving the flexibility and efficiency of detection; cloud servers can undertake larger-scale computing and analysis tasks, and also have more powerful data storage capabilities; the edge-cloud collaborative working mode combines the advantages of edge computing and cloud computing, and can flexibly allocate computing tasks according to actual computing and storage resources. While protecting data privacy, it further improves data processing and analysis capabilities, providing support for enhancing recognition performance.
[0147] Transformer-based emotion recognition method
[0148] The present invention proposes a lightweight emotion recognition method based on the Transformer architecture. The method adopts a dual-path feature extraction backbone network to extract the features of EEG signals in the time domain and frequency domain respectively. The main structure of the method is as follows: Figure 2 As shown. First, the EEG signal to be identified will be input into the time domain module (TD module) and the frequency domain module (FD module) to obtain the preliminary feature representation of the signal in the time domain and frequency domain. Then, the backbone extraction network including multiple lightweight transformer blocks will perform deep feature extraction on the time domain and frequency domain information. The features obtained by the backbone extraction network will be dynamically adjusted in the adaptive cross-domain fusion module to achieve adaptive fusion of effective information in the time domain and frequency domain. The following is an introduction to each module in the recognition method.
[0149] The time domain module (TD module) mainly uses deep separable convolution to perform preliminary feature extraction on the time domain signal, which is expressed as follows:
[0150] x D =ReLU(BN(DWConv(x)))
[0151] x P =ReLU(BN(PWConv(x D )))
[0152] Among them, x represents the input EEG signal, DWConv represents the deep convolution operation, BN represents the batch normalization operation, ReLU represents the activation function, and x D is the feature obtained through deep convolution related operations, PWConv stands for point-wise convolution operation (Depthwise Convo l ut ion), x P It is the feature obtained through point-by-point convolution related operations.
[0153] The frequency domain module (FD module) first converts the time domain EEG signal into the frequency domain through fast Fourier transform, then uses deep separable convolution to extract the frequency domain features, and combines residual connection to further promote the multi-level interaction of information, as shown below:
[0154] x fD =ReLU(BN(DWConv(FFT(x))))+x
[0155] x fP =ReLU(BN(PWConv(x fD )))
[0156] Where x represents the input EEG signal, FFT represents the fast Fourier transform operation (fast Fourier transform), and x fD Represents the intermediate features of frequency domain information obtained after deep convolution operation and residual connection, x fP It is the feature obtained by performing point-by-point convolution operations on the intermediate features.
[0157] The backbone extraction network contains three stages. In the first stage, the data is first dimensionally divided through the Patch Embedding module, and then the features are extracted through the designed lightweight Transformer module. In the second stage, the Patch Merging module is used to reduce the resolution of the feature map and expand the receptive field of the feature (downsampling), followed by the lightweight Transformer module. The third stage is similar to the second stage and is used to extract deeper features.
[0158] The first lightweight module and the second lightweight module have structures such as Figure 3As shown, the first lightweight module and the second lightweight module contain three parallel paths. After the deep convolution operation, the input data enters the three paths respectively, one of which is the original feature, and the other two paths are downsampled by 4 times and 8 times respectively. Then, they are all processed by the lightweight self-attention module (LSA), and then the downsampled path is upsampled accordingly according to its multiple. Finally, the features of the three paths are spliced. At the same time, residual connections are introduced in this splicing operation. After completing the multi-scale feature extraction, the deep convolution layer with residual connections and the multi-layer perceptron (MLP) are used for deeper integration and extraction. The lightweight Transformer module is expressed as follows:
[0159] z o =LSA(DWConv(z))
[0160] z d1 =LSA(Down 1 (DWConv(z)))
[0161] z d2 =LSA(Down 2 (DWConv(z)))
[0162] z m =z o +Up 1 (z d1 )+Up 2 (z d2 )+z
[0163] z l =MLP(DWConv(z m ))+z m
[0164] Among them, z represents the input feature, LSA represents the lightweight self-attention module, and z o Indicates the result obtained by the input data through the original feature path, Down 1 represents a 4x downsampling module, z d1 Indicates the result obtained by downsampling the input data by 4 times. 2 represents an 8-fold downsampling module, z d2 Indicates the result obtained by downsampling the input data by 8 times. 1 Indicates 4x upsampling module, Up 2 Indicates 8x upsampling module, z mrepresents the intermediate features after multi-scale feature extraction, z l1 represents the features obtained after processing in the lightweight Transformer module, i.e., the first emotion feature data; z l2 It represents the features obtained after processing in the lightweight Transformer module, that is, the second emotion feature data.
[0165] In traditional self-attention modules, there is a quadratic computational complexity, which causes the computational requirements to increase significantly as the input feature size increases. In particular, in the multi-head self-attention mechanism, the attention weights need to be calculated for each head separately. In order to process and combine information from different heads, data reconstruction and normalization operations need to be performed frequently, which leads to a large increase in memory access requirements and affects the operating efficiency of the model. In order to reduce the redundancy of computation and memory access, a lightweight self-attention module is designed. The specific structure is as follows: Figure 4 As shown in the figure, this module adopts a single-head self-attention calculation method, and further reduces the calculation requirements by means of the segmentation operation of the input features, and uses some of the original features retained after the segmentation operation to promote the model's learning of the identity feature mapping. At the same time, due to the permutation invariance of the self-attention mechanism operation, the model will not be able to fully learn the position information of the features, thereby affecting the information interaction of cross-regional features. Therefore, different forms of position encoding mechanisms are added to the Transformer architecture to allow the model to learn local position information and its global dependencies. Absolute position encoding encodes the position information of the element into a fixed position vector through a sine or cosine function, which causes it to ignore the relative position relationship between different elements. This fixed encoding may not be able to adapt to new position information in real time in dynamically changing tasks.
[0166] In order to solve the above problems, a learnable position encoding mechanism is designed to improve the position expression of features and the ability of global information interaction, and use deep convolution operations to generate dynamic position information. The lightweight self-attention module is specifically expressed as follows:
[0167] X att ,X res =Spi lt(X)
[0168] X satt =Se l f-Attention(DWConv(X att W Q ,X att W K ,X att W V ))
[0169] X pe =X satt +LPE(Xatt W V )
[0170] X O =X pe +X res
[0171] Among them, X represents the feature of the input lightweight self-attention module, Spilt represents the segmentation operation, and X att With X res They represent the features used to input the self-attention operation and the residual operation, respectively. Q , W K , W V They represent the weight matrices used to calculate Q, K, and V, respectively. Self-Attention represents the self-attention operation. satt represents the feature obtained through the self-attention operation, LPE represents the learnable position encoding operation, and X pe represents the feature obtained after the self-attention operation combined with the learnable position encoding, X O Represents the features processed by the lightweight self-attention module.
[0172] In order to fuse the time domain and frequency domain features processed by the backbone extraction network, an adaptive cross-domain information fusion module is designed. The specific structure is as follows: Figure 5 As shown in the figure, the input features are first transformed in dimension through the embedding layer, so that the time domain and frequency domain features can enter the unified representation space for subsequent fusion operations, and then the spatial attention and channel attention mechanisms are applied to strengthen the feature information in the channel dimension and the spatial dimension. The two output features are then adaptively fused through the gated unit, which dynamically controls the fusion weight of each feature by learning the correlation between the input time domain and frequency domain features. This adaptive fusion process enables the model to automatically adjust the fusion strategy according to the task requirements and the input features, thereby improving the expression ability of the spatiotemporal features. Finally, the fused features are further processed by the multi-layer perceptron (MLP) to obtain the recognition results. The adaptive cross-domain information fusion module is specifically expressed as follows:
[0173] X e1 =Embedding 1 (X 1 )
[0174] X e2 =Embedding 2 (X 2 )
[0175] X c1 =σ(MLP(AvgPool(X e1 ))+MLP(MaxPoo l(X e1 )))
[0176] X c2 =σ(MLP(AvgPool(X e2 ))+MLP(MaxPoo l(X e2 )))
[0177] X s1 =σ(Conv([AvgPool(X e1 ));MaxPool(X e1 )]))
[0178] X s2 =σ(Conv([AvgPool(X e2 ));MaxPool(X e2 )]))
[0179] X f1 =σ(W c1 X c1 +W s1 X s1 +b 1 )
[0180] X f2 =σ(W c2 X c2 +W s2 X s2 +b 2 )
[0181] X af =MLP(X f1 +X f2 )
[0182] Among them, X e1 With X e2 Respectively represent the features after Embedding operation, X c1 With X c2 They represent the features obtained after channel attention processing, X s1 With X s2 They represent the features obtained after spatial attention processing, X f1 With X f2 They represent the features obtained after adaptive fusion through the gating unit, AvgPool and MaxPool represent the average pooling operation and the maximum pooling operation respectively, Conv represents the convolution operation, σ represents the activation function, and W c1、 Ws1、 W c2、 W s2 are all learnable weights in the gated unit, b 1 With b 2 represents the bias term, X af Represents the final features obtained after the adaptive cross-domain information fusion module.
[0183] Although the present invention has been described above, the present invention is not limited to the above-mentioned specific embodiments. The above-mentioned specific embodiments are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can make many modifications without departing from the purpose of the present invention, which are all within the protection of the present invention.
Claims
1. An edge-cloud collaborative system for emotion recognition, the system comprising an edge computing device and a cloud server; characterized in that: The edge computing device includes a data acquisition module, a data preprocessing module, a feature extraction module, an emotion recognition module, an edge storage module, an edge interaction module, a device management module and a first data transmission module, and the cloud server includes a cloud computing module, a cloud storage module and a second data transmission module, wherein: Feature extraction module: obtains the local user's EEG data after preprocessing from the edge storage module, extracts EEG signal features using signal processing methods, and transmits the extracted EEG signal feature data to the edge storage module for storage; Emotion recognition module: loads the extracted user EEG signal feature data and local emotion recognition model from the edge storage module, and outputs the emotion type inferred locally to the edge storage module for storage; Edge storage module: used to store and manage data collected, processed and analyzed by other modules of the edge computing device and data transmitted from the cloud server via the first data transmission module; The edge interaction module is used to provide a visual graphical operation interface to guide users to input or confirm their basic information during data collection, and to display the recognition results on the graphical interface after the recognition is completed. Users can provide feedback through the edge interaction module; wherein: The cloud server initializes the training of the emotion recognition model and transmits the initial model to the edge storage module of each edge computing device via the second data transmission module; After the edge computing device adds user data, the data is integrated and uploaded to the cloud server via the first data transmission module; Based on the distributed user data uploaded by each edge computing device, the cloud server uses the integrated EEG data to retrain the emotion recognition model to obtain a global model, and transmits the global model to each edge computing device; Each edge computing device updates the local emotion recognition model based on the global model and repeats the above process to improve the accuracy of the emotion recognition model.
2. The edge-cloud collaborative system for emotion recognition according to claim 1, characterized in that: The emotion recognition module adopts a dual-path feature extraction backbone network structure, which includes a frequency domain single element, a time domain unit, a first lightweight module, a second lightweight module and an information fusion module; wherein: The time domain unit extracts features from the time domain EEG signal according to the following formula to obtain time domain EEG feature data; x D =ReLU(BN(DWConv(x))) x P =ReLU(BN(PWConv(x D ))) Where x represents the input EEG signal, DWConv represents the deep convolution operation, BN represents the batch normalization operation, ReLU represents the activation function, and x D is the feature obtained through deep convolution related operations, PWConv represents point-by-point convolution operation, x P It is the feature obtained through point-by-point convolution related operations; The frequency domain unit extracts the features of the time domain EEG signal according to the following formula to obtain the frequency domain EEG feature data; x fD =ReLU(BN(DWConv(FFT(x))))+x x fP =ReLU(BN(PWConv(x fD ))) Where x represents the input EEG signal, FFT represents the fast Fourier transform processing, and x fD Represents the intermediate features of frequency domain information obtained after deep convolution operation and residual connection, x fP It is the feature obtained by performing point-by-point convolution operations on the intermediate features; The first lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the first emotion feature data; With o =LSA(DWConv(z)) With d1 =LSA(Down1(DWConv(z))) With d2 =LSA(Down2(DWConv(z))) from m =from o +Up1(from d1 )+Up2(from d2 )+z With l1 =MLP(DWConv(from m ))+z m Among them, z represents the input time domain EEG feature data, LSA represents the lightweight self-attention module, and z o It indicates the result obtained by the input data through the original feature path, Down1 indicates the 4-fold downsampling module, and z d1 represents the result of the input data being downsampled by 4 times, Down2 represents the 8-fold downsampling module, and z d2 Indicates the result of the input data after the 8-fold downsampling path, Up1 indicates the 4-fold upsampling module, Up2 indicates the 8-fold upsampling module, and z m It represents the intermediate features after multi-scale feature extraction, and MLP represents multi-layer perceptron; The second lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the second emotion feature data; With o =LSA(DWConv(z)) With d1 =LSA(Down1(DWConv(z))) With d2 =LSA(Down2(DWConv(z))) from m =from o +Up1(from d1 )+Up2(from d2 )+z With l2 =MLP(DWConv(from m ))+z m The information fusion module adopts an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain the emotion type.
3. The edge-cloud collaborative system for emotion recognition according to claim 2, characterized in that: The first lightweight module and the second lightweight module are both composed of a Patch Embedding layer, a Patch Merging layer and Transformer layer; the lightweight self-attention module is set in the Transformer layer; the lightweight The self-attention module includes a segmentation unit, a self-attention unit and a position encoding unit; wherein: The segmentation unit performs feature segmentation according to the following formula; X att ,X res =Spi lt(X) Where: X represents the input feature of the lightweight self-attention module, Spilt represents the segmentation operation, and X att With X res Represent the features used to input self-attention operation and residual operation respectively; The self-attention unit calculates the first emotion feature data and the second emotion feature data respectively according to the following formula To obtain deep sentiment features: X satt =Se l f-Attent ion(DWConv(X att W Q ,X att W K ,X att W V )) Where: W Q , W K , W V Respectively represent the weight matrices used to calculate Q, K, and V, and Self-Attention represents the self-attention operation; The position encoding unit processes the deep emotion feature according to the following formula to obtain the global deep emotion feature: levy; X pe =X satt +LPE(X att W V ) X O =X pe +X res Where: LPE represents the learnable position encoding operation, X pe Represents the features obtained after the self-attention operation is combined with the learnable position encoding.
4. The edge-cloud collaborative system for emotion recognition according to claim 2, characterized in that: The information fusion module consists of an embedding layer, a spatial attention layer, a channel attention layer, a gating unit, and a multi-layer perception mechanism. The information fusion module uses an adaptive cross-domain fusion algorithm to combine the first emotion feature data with the second emotion feature data. The process of obtaining the emotion type by fusing the feature data is as follows: The Embedding layer performs dimension conversion on the first emotion feature data and the second emotion feature data according to the following formula to obtain the first emotion feature representation and the second emotion feature representation; X e1 =Embedd ing1(X1) X e2 =Embedd ing2(X2) The spatial attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion space feature and the second emotion space feature; X s1 =σ(Conv([AvgPoo l(X e1 ));MaxPoo l(X e1 )])) X s2 =σ(Conv([AvgPoo l(X e2 ));MaxPoo l(X e2 )])) The channel attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion channel feature and the second emotion channel feature; X c1 =σ(MLP(AvgPoo l(X e1 ))+MLP(MaxPoo l(X e1 ))) X c2 =σ(MLP(AvgPoo l(X e2 ))+MLP(MaxPoo l(X e2 ))) The gating unit processes the first emotion space feature and the first emotion channel feature according to the following formula to obtain the first local emotion feature: X f1 =σ(W c1 X c1 +W s1 X s1 +b1) Where: σ represents the activation function, W c1、 W s1 are all learnable weights in the gated unit, and b1 represents the bias term; The gate control unit processes the second emotion space feature and the second emotion channel feature according to the following formula to obtain the second local emotion feature: X f2 =σ(W c2 X c2 +W s2 X s2 +b2) Where: σ represents the activation function, W c2、 W s2 are all learnable weights in the gated unit, and b2 represents the bias term; The multi-layer perceptron fuses the first local emotion feature and the second local emotion feature according to the following formula to obtain the emotion type: X af =MLP(X f1 +X f2 )。 5. A method for emotion recognition using the edge-cloud coordination system according to any one of claims 1 to 4, characterized in that: The steps include: Obtain the local user's EEG data after preprocessing from the edge storage module, and use the signal processing method to obtain Extract EEG signal features and transmit the extracted EEG signal feature data to the edge storage module for storage Storage; Load the extracted user EEG signal feature data and local emotion recognition model from the edge storage module and The emotion type obtained by local reasoning is output to the edge storage module for storage; wherein: The time domain unit extracts features from the time domain EEG signal according to the following formula to obtain time domain EEG feature data; x D =ReLU(BN(DWConv(x))) x P =ReLU(BN(PWConv(x D ))) Where x represents the input EEG signal, DWConv represents the deep convolution operation, BN represents the batch normalization operation, ReLU represents the activation function, and x D is the feature obtained through deep convolution related operations, PWConv represents point-by-point convolution operation, x P It is the feature obtained through point-by-point convolution related operations; The frequency domain unit extracts features from the time domain EEG signal according to the following formula to obtain frequency domain EEG feature data; x fD =ReLU(BN(DWConv(FFT(x))))+x x fP =ReLU(BN(PWConv(x fD ))) Where x represents the input EEG signal, FFT represents the fast Fourier transform processing, and x fD Represents the intermediate features of frequency domain information obtained after deep convolution operation and residual connection, x fP It is the feature obtained by performing point-by-point convolution operations on the intermediate features; The first lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the first emotion feature data; With o =LSA(DWConv(z)) With d1 =LSA(Down1(DWConv(z))) With d2 =LSA(Down2(DWConv(z))) from m =from o +Up1(from d1 )+Up2(from d2 )+z With l1 =MLP(DWConv(from m ))+z m Among them, z represents the input time domain EEG feature data, LSA represents the lightweight self-attention module, and z o It indicates the result obtained by the input data through the original feature path, Down1 indicates the 4-fold downsampling module, and z d1 represents the result of the input data being downsampled by 4 times, Down2 represents the 8-fold downsampling module, and z d2 Indicates the result of the input data after the 8-fold downsampling path, Up1 indicates the 4-fold upsampling module, Up2 indicates the 8-fold upsampling module, and z m It represents the intermediate features after multi-scale feature extraction, and MLP represents multi-layer perceptron; The second lightweight module extracts the time-domain EEG feature data according to the following formula to obtain the second emotion feature data; With o =LSA(DWConv(z)) With d1 =LSA(Down1(DWConv(z))) With d2 =LSA(Down2(DWConv(z))) from m =from o +Up1(from d1 )+Up2(from d2 )+z With l2 =MLP(DWConv(from m ))+z m The information fusion module uses an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain the emotion type; Edge storage module: used to store and manage data collected, processed and analyzed by other modules of edge computing devices and data transmitted from the cloud server via the first data transmission module; wherein: The cloud server initializes the training emotion recognition model and transmits the initial model via the second data transmission module Transmitted to the edge storage module of each edge computing device; After adding user data to the edge computing device, the data is integrated and uploaded to the cloud via the first data transmission module. End server; Based on the distributed user data uploaded by each edge computing device, the cloud server uses the integrated EEG data to retrain the emotion recognition model to obtain a global model, and transmits the global model to each edge computing device; Each edge computing device updates the local emotion recognition model based on the global model, and repeats the above process to improve the emotion recognition model. The accuracy of the emotion recognition model.
6. The method for emotion recognition using the edge-cloud coordination system of claim 5, characterized in that: The information fusion module uses an adaptive cross-domain fusion algorithm to fuse the first emotion feature data with the second emotion feature data to obtain Emotional type processes, including: The Embedding layer performs dimension conversion on the first emotion feature data and the second emotion feature data according to the following formula to obtain the first emotion feature representation and the second emotion feature representation; X e1 =Embedd ing1(X1) X e2 =Embedd ing2(X2) The spatial attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion space feature and the second emotion space feature; X s1 =σ(Conv([AvgPoo l(X e1 ));MaxPoo l(X e1 )])) X s2 =σ(Conv([AvgPoo l(X e2 ));MaxPoo l(X e2 )])) The channel attention layer processes the first emotion feature representation and the second emotion feature representation according to the following formula to obtain the first emotion channel feature and the second emotion channel feature; X c1 =σ(MLP(AvgPoo l(X e1 ))+MLP(MaxPoo l(X e1 ))) X c2 =σ(MLP(AvgPoo l(X e2 ))+MLP(MaxPoo l(X e2 ))) The gating unit processes the first emotion space feature and the first emotion channel feature according to the following formula to obtain the first local emotion feature: X f1 =σ(W c1 X c1 +W s1 X s1 +b1) Where: σ represents the activation function, W c1、 W s1 are all learnable weights in the gated unit, and b1 represents the bias term; The gate control unit processes the second emotion space feature and the second emotion channel feature according to the following formula to obtain the second local emotion feature: X f2 =σ(W c2 X c2 +W s2 X s2 +b2) Where: σ represents the activation function, W c2、 W s2 are all learnable weights in the gated unit, and b2 represents the bias term; The multi-layer perceptron fuses the first local emotion feature and the second local emotion feature according to the following formula to obtain the emotion type: X af =MLP(X f1 +X f2 )。