An emotion perception assistance system for intelligent communication devices

CN122604378APending Publication Date: 2026-08-21RIZHAO VOCATIONAL & TECHNICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611023407.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

然而,目前的情绪感知技术多局限于视觉和语音信号的分析,尚未充分利用脑电信号

Benefits of technology

[0048] This invention improves the accuracy and robustness of the emotion recognition model by introducing multiple methods for extracting emotion-related features, and also enhances the model's ability to distinguish complex emotional states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122604378A_ABST
    Figure CN122604378A_ABST
Patent Text Reader

Abstract

The application discloses an emotion perception auxiliary system and method for intelligent communication equipment, comprising the following steps: acquiring the electroencephalogram signal of a user through an electroencephalogram acquisition module, applying a data enhancement technology to expand data after signal preprocessing, thereby improving the generalization ability of the model, and then extracting multi-dimensional features of the emotional state from the time domain, the frequency domain and the space domain. The emotion recognition module adopts a deep learning model combining a convolutional neural network and a self-attention network to perform emotion classification on the electroencephalogram feature signal, and the system automatically adjusts the display mode of the equipment or plays prompt information according to the recognition result. The intelligent response module provides adaptive feedback according to the emotional state of the user. The application realizes accurate recognition and personalized feedback of the emotional state of the user, endows the intelligent communication equipment with the emotion perception function, and effectively improves the experience and interactivity of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of emotion recognition and intelligent communication devices, and in particular to an emotion perception assistance system based on EEG signal analysis. Background Technology

[0002] With the widespread adoption of smart devices, users are increasingly demanding personalization and intelligence. In emotion recognition, electroencephalogram (EEG) signals, as a non-invasive physiological signal, can provide real-time feedback on a user's emotional state, making emotional interaction possible for smart devices. However, current emotion perception technologies are largely limited to the analysis of visual and speech signals, and have not yet fully utilized EEG signals. Summary of the Invention

[0003] The main objective of this invention is to provide an emotion perception assistance system that enhances emotional adaptability in real-time communication through accurate emotion recognition technology. This system identifies a series of key features affecting emotion state recognition and constructs an effective multi-feature fusion algorithm to ensure high recognition rates and robustness under different emotional states, thereby improving the naturalness and intelligence of human-computer interaction.

[0004] The technical solution adopted in this invention is: an emotion perception assistance system and method for intelligent communication devices, comprising the following specific steps:

[0005] S1: The EEG acquisition module is used to collect the user's EEG signals. The EEG acquisition module is configured as a multi-channel module to collect emotion-related electrical signals from different brain regions of the user in a wireless transmission manner, with a sampling frequency of not less than 512Hz.

[0006] S2: The signal preprocessing module performs filtering and noise suppression on the EEG signals acquired by the EEG acquisition module;

[0007] S3: The data augmentation module performs time clipping and time shifting, adds Gaussian noise, and stretches and compresses the EEG signal processed by the signal preprocessing module.

[0008] S4: The feature extraction module extracts and integrates multi-dimensional features of EEG signals from the time domain, frequency domain, and spatial domain. In the feature integration stage, the initial domain weight coefficients and feature weight coefficients are preset, and the weight ratio between the time domain, frequency domain, and spatial domain features is automatically and dynamically adjusted according to the user's historical emotional fluctuation data through the meta-learning algorithm to generate an adaptive feature matrix.

[0009] S5: The emotion recognition module uses a deep hybrid model composed of a convolutional neural network (CNN) and a self-attention model (Transformer) to classify emotions based on the multidimensional features of EEG signals.

[0010] The CNN module uses a 1×3 convolutional kernel to extract local features from the adaptive feature matrix. During the extraction process, it avoids disrupting the transient phase relationship between channels in different brain regions and converts the extracted local features into an embedded representation, which is then input into the Transformer module for global dependency modeling.

[0011] The CNN module also includes a phase consistency preservation unit. Before the convolution operation, this unit first calculates the phase-locking value between each pair of EEG channels and introduces this phase-locking value as a regularization term into the loss function of the CNN. During the convolution process, the phase difference of the feature maps of different channels within the same frequency band is forced to not exceed a preset phase distortion threshold. The output feature map of the convolutional layer needs to pass through a phase reconstruction verification layer. This verification layer verifies whether the output features still maintain the cross-channel phase topology relationship at the time of input. If the phase topology similarity is lower than the preset similarity value, feature re-extraction is triggered.

[0012] S6: The intelligent response module has an adaptive steady-state interactive unit, which continuously identifies the results of the user's emotion classification and adjusts the display mode of the communication device or provides voice broadcast reminders based on the identification results;

[0013] The intelligent response module possesses adaptive steady-state interactivity. It continuously identifies the user's emotions and adjusts the display mode of the communication device or provides voice prompts based on the identification results, including the following steps:

[0014] When the system detects that the user is in a persistently low mood, it will automatically play upbeat and cheerful music, and the background color of the display screen will automatically change to a comfortable orange to alleviate the user's mood.

[0015] When the device detects that the user is in a state of sustained agitation, it will automatically play soothing light music, the background color of the display screen will automatically change to a soft blue, and a prompt message will be displayed on the device interface to remind the user to adjust their emotions appropriately.

[0016] Users can customize the response of emotion reminders in the device's settings and can enable or disable the voice broadcast function;

[0017] S7: The intelligent response module has an adaptive transient interaction unit. The adaptive transient interaction unit extracts the nonlinear exponential or differential variance features in the adaptive feature matrix in real time. When the instantaneous change rate of the above features exceeds the preset threshold, it is determined to be an abnormal emotional fluctuation. Before the emotion recognition module outputs the final classification result, it feeds forward to control the underlying driver of the intelligent communication device to quickly reduce the display brightness and volume.

[0018] The intelligent response module possesses adaptive transient interactivity, including the following steps:

[0019] When the system detects an abnormally large increase in the fluctuation of a user's emotional state, it will quickly and automatically reduce the display brightness and volume, automatically play soothing light music, automatically change the background color of the display screen to a soft blue, and display a prompt message on the device interface to remind the user to adjust their emotions appropriately.

[0020] When the user's emotional state fluctuates slightly, the system will maintain normal display brightness and volume, but will provide periodic positive feedback prompts.

[0021] The system provides the function of saving and reviewing user feedback history, allowing users to view the trend of their emotional state changes;

[0022] The adaptive transient interactivity unit also includes a conflict arbitration module: when the transient path detects that the instantaneous rate of change exceeds a threshold, the observation window of the reorganization unit is activated; within the observation window, the second-order difference of the nonlinear exponent and the skin conductance response are extracted simultaneously as supporting evidence; if the sign inversion of the second-order difference exceeds 2 times or the supporting evidence does not support emotional abnormality, the feedforward instruction is suppressed; if the transient instruction contradicts the classification result output by the emotion recognition module after a subsequent delay of 200ms, the system records a transient-steady-state inconsistency event and fine-tunes the transient threshold in the reverse direction; when the number of inconsistency events exceeds a preset number within 1 hour, the system automatically switches the transient response mode from immediate execution to arbitration before execution and generates a prompt asking the user whether to close the transient interaction.

[0023] In S1 above, the EEG acquisition module also includes an adaptive electrode contact status monitoring module, which monitors the electrode contact status in real time, reduces signal noise caused by poor electrode contact, and issues an alert when the contact resistance exceeds 8kΩ. The EEG acquisition module also includes an electrode contact quality-driven data repair channel: when the contact resistance of a certain channel exceeds 10kΩ within 5 consecutive sampling points, the system marks the channel as a low-confidence channel. At this time, the data of the channel is not discarded, but the EEG signals of adjacent spatial channels are used to generate a replacement signal for the low-confidence channel in real time through a spatial interpolation model based on a graph convolutional network. The spatial interpolation model uses the user's historical high-quality data as training samples to learn a nonlinear mapping function for each electrode position. If more than 3 channels are marked as low-confidence at the same time, emotion recognition is paused and the user is prompted to check the electrodes.

[0024] In S2 above, the signal preprocessing module performs filtering and noise suppression on the EEG signal, including the following steps:

[0025] First, a Butterworth high-pass filter and a Butterworth low-pass filter are applied to the original EEG signal to remove extremely low-frequency DC drift and high-frequency electromyography interference.

[0026] Secondly, a Butterworth notch filter is used to suppress power frequency noise;

[0027] Finally, ICA was further applied to separate electrooculography, electromyography and other interference signals from the EEG signal. The decomposition dimension of ICA was set to be consistent with the number of electrode channels.

[0028] In S3 above, the data augmentation module improves the recognition accuracy of the emotion recognition model through data augmentation techniques. The data augmentation methods include three parts: time pruning and time shifting, adding Gaussian noise, and signal stretching and compression, including the following steps:

[0029] First, randomly select a segment from each EEG signal sequence, randomly determine a starting point on the time axis of the signal sequence, and then cut out a signal segment of fixed length from that point.

[0030] Then, a random time shift operation is performed on the cropped signal segment to produce a slight offset on the time axis, thereby improving the model's robustness to time changes.

[0031] Next, Gaussian noise is added to the clipped and shifted signal segments;

[0032] Finally, the noise-enhanced signal segments are slightly stretched or compressed to simulate physiological differences between individuals.

[0033] In S4 above, the feature extraction module calculates the time-domain features, frequency-domain features, and spatial-domain features of the EEG signal respectively, integrates the above features in the time domain, frequency domain, and spatial domain, and then inputs them into the emotion recognition module;

[0034] The time-domain features include mean potential, standard deviation, peak value, peak-to-peak value, root mean square, variance, skewness, kurtosis, interquartile range, maximum value, minimum value, absolute deviation of mean, energy, integral, number of zero crossings, amplitude envelope, autocorrelation coefficient, sum of squared amplitudes, signal entropy, and Lempel-Ziv complexity.

[0035] The frequency characteristics include relative energy, total power, main band power, instantaneous frequency, power spectral density, phase consistency, frequency offset, frequency center, bandwidth, harmonic ratio, frequency entropy, band energy ratio, short-time Fourier transform energy, weighted frequency, band peak value, signal amplitude spectrum, signal phase spectrum, Hilbert transform energy, frequency domain entropy, and resonant frequency.

[0036] The spatial features include spatial filtering features, channel relative strength, local field potential, mutual information, differential entropy, spatial energy, linear discriminant features, symmetry index, channel cross-correlation, channel coherence, brain power source localization, spatial covariance, independent component analysis, spatial pattern recognition, differential variance, channel skewness, channel kurtosis, nonlinear exponent, feature symmetry, and neural signal spatial entropy.

[0037] In S5 above, the emotion recognition module uses a deep hybrid model composed of CNN and Transformer for emotion classification;

[0038] The CNN module includes three convolutional layers. The first layer has 32 convolutional kernels, a kernel size of 1×3, and a stride of 1. The second layer has 64 convolutional kernels, a kernel size of 1×3, and a stride of 1. The third layer has 128 convolutional kernels, a kernel size of 1×3, and a stride of 1. Each convolutional layer is followed by a max pooling layer with a kernel size of 1×2 and a stride of 2. The ReLU activation function is applied after each convolutional layer.

[0039] The Transformer module includes an input embedding layer that converts the feature sequence output by the convolutional layer into an embedding representation with 128 dimensions, and uses a 6-layer Transformer encoder, each layer including 8 attention heads, each head having a dimension of 16. A fully connected layer is added after the last Transformer encoder layer to map the output to the sentiment category space.

[0040] The above also includes S8: when training and testing the emotion recognition module, determining the batch size of the training set and the test set, the optimal optimizer, the number of iterations, the learning rate, and the weight decay coefficient required to train the model; in step S8, when training and testing the emotion recognition model, the batch size of both the training set and the test set is set, the Adam optimizer is used for training, the initial learning rate is set, the learning rate decay coefficient is preset, and each iteration during training is tested on the validation set.

[0041] In S8 above, after completing model training, the model is tested and evaluated, including the following steps:

[0042] The classification process is repeated for each test set sample, and the average of the classification results is calculated as the final prediction result.

[0043] The confusion matrix and F1 score were used to evaluate the classification accuracy of the model on the test set.

[0044] The evaluation results are displayed on the device interface, providing emotion recognition accuracy and confusion matrix.

[0045] The meta-learning algorithm in S4 above includes a dual-timescale memory unit: a long-term memory (LTM) stores the baseline feature weight distribution of the user's emotions over the past 30 days, and a short-term working memory (STM) stores the real-time feature fluctuations over the most recent hour. Every 5 minutes, the system calculates the KL divergence between the current adaptive feature matrix and the corresponding emotional state in the LTM. If the divergence exceeds a dynamic threshold, it is determined that emotional drift has occurred. At this time, the meta-learning algorithm prioritizes updating the weights using the gradients from the STM and uses the weights from the LTM as a regularization term to prevent overfitting to instantaneous noise. The dynamic threshold is adaptively adjusted according to the user's usage time: the initial threshold is used, and the threshold is reduced after more than 7 days of use.

[0046] Between the feature extraction module and the emotion recognition module, an emotion-state-specific feature recombining unit is inserted. This recombining unit pre-stores a feature-emotion correlation matrix pre-trained based on a large amount of population data. For each basic emotion, the correlation matrix provides the corresponding optimal feature subset. The system calculates the feature activation intensity of the current EEG signal on each emotion correlation subset in real time, and only concatenates the features corresponding to the two emotion subsets with the highest activation intensity before sending them into the CNN-Transformer model. Inactive feature subsets are discarded, and the discard ratio is dynamically adjusted to retain 30% of the features to prevent information loss.

[0047] This invention has positive effects:

[0048] This invention improves the accuracy and robustness of the emotion recognition model by introducing multiple methods for extracting emotion-related features, and also enhances the model's ability to distinguish complex emotional states.

[0049] By deeply analyzing EEG signals, we have successfully achieved accurate classification of different emotional states, providing a more reliable foundation for emotion perception and intelligent interaction.

[0050] The system's high performance makes real-time emotion recognition possible, significantly improving the effectiveness of emotion-driven communication applications, such as the emotional adaptability of intelligent assistants in user communication, thereby improving user experience, enhancing the naturalness and intelligence of human-computer interaction, and laying a solid technical foundation for the future development of communication technology.

[0051] This invention overcomes the limitations of traditional equal-weight feature concatenation. By introducing a meta-learning algorithm in the S4 stage, the system can automatically and dynamically optimize the weight ratio of time-domain, frequency-domain, and spatial-domain features based on the unique physiological signal baselines and historical emotional fluctuation data of different users. This makes the generated adaptive feature matrix highly consistent with individual differences, significantly improving the robustness of cross-subject emotion recognition.

[0052] This invention is based on a transient response mechanism of specific EEG features. The system does not need to wait for the deep learning model to complete the full and time-consuming forward propagation calculation. Once the instantaneous change rate of the underlying features exceeds the threshold, it can directly call the underlying driver of the communication device to adjust the screen and volume in a feedforward manner, thus shortening the device's response delay to sudden changes in the user's emotions. Attached Figure Description

[0053] The invention will now be further described with reference to the accompanying drawings.

[0054] Figure 1 This is a diagram showing the overall architecture of the system of the present invention;

[0055] Figure 2 This is a diagram of the intelligent response module of the system of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0057] (Example)

[0058] An emotion perception assistance system and method for intelligent communication devices includes the following specific steps:

[0059] S1: The EEG acquisition module is used to collect the user's EEG signals. The EEG acquisition module is configured as a multi-channel module to collect emotion-related electrical signals from different brain regions of the user in a wireless transmission manner, with a sampling frequency of not less than 512Hz.

[0060] S2: The signal preprocessing module performs filtering and noise suppression on the EEG signals acquired by the EEG acquisition module;

[0061] S3: The data augmentation module performs time clipping and time shifting, adds Gaussian noise, and stretches and compresses the EEG signal processed by the signal preprocessing module.

[0062] S4: The feature extraction module extracts and integrates multi-dimensional features of EEG signals from the time domain, frequency domain, and spatial domain. In the feature integration stage, the initial domain weight coefficients and feature weight coefficients are preset, and the weight ratio between the time domain, frequency domain, and spatial domain features is automatically and dynamically adjusted according to the user's historical emotional fluctuation data through the meta-learning algorithm to generate an adaptive feature matrix.

[0063] S5: The emotion recognition module uses a deep hybrid model composed of a convolutional neural network (CNN) and a self-attention model (Transformer) to classify emotions based on the multidimensional features of EEG signals.

[0064] The CNN module uses a 1×3 convolutional kernel to extract local features from the adaptive feature matrix. During the extraction process, it avoids disrupting the transient phase relationship between channels in different brain regions and converts the extracted local features into an embedded representation, which is then input into the Transformer module for global dependency modeling.

[0065] The CNN module also includes a phase consistency preservation unit. Before the convolution operation, this unit first calculates the phase locking value (PLV) between each pair of EEG channels and introduces this PLV as a regularization term into the loss function of the CNN. During the convolution process, the phase difference variation of the feature maps of different channels within the same frequency band (θ wave, α wave, or β wave) is forcibly constrained to not exceed a preset phase distortion threshold (±5°). The output feature map of the convolutional layer needs to pass through a phase reconstruction verification layer. This verification layer verifies whether the output features still maintain the cross-channel phase topology relationship at the time of input. If the phase topology similarity is less than 0.92, feature re-extraction is triggered.

[0066] S6: The intelligent response module has an adaptive steady-state interactive unit, which continuously identifies the results of the user's emotion classification and adjusts the display mode of the communication device or provides voice broadcast reminders based on the identification results;

[0067] The intelligent response module possesses adaptive steady-state interactivity. It continuously identifies the user's emotions and adjusts the display mode of the communication device or provides voice prompts based on the identification results, including the following steps:

[0068] When the system detects that the user is in a persistently low mood, it will automatically play upbeat and cheerful music, and the background color of the display screen will automatically change to a comfortable orange to alleviate the user's mood.

[0069] When the device detects that the user is in a state of sustained agitation, it will automatically play soothing light music, the background color of the display screen will automatically change to a soft blue, and a prompt message will be displayed on the device interface to remind the user to adjust their emotions appropriately.

[0070] Users can customize the response of emotion reminders in the device's settings and can enable or disable the voice broadcast function;

[0071] S7: The intelligent response module has an adaptive transient interaction unit. The adaptive transient interaction unit extracts the nonlinear exponential or differential variance features in the adaptive feature matrix in real time. When the instantaneous change rate of the above features exceeds the preset threshold, it is determined to be an abnormal emotional fluctuation. Before the emotion recognition module outputs the final classification result, it feeds forward to control the underlying driver of the intelligent communication device to quickly reduce the display brightness and volume.

[0072] The intelligent response module possesses adaptive transient interactivity, including the following steps:

[0073] When the system detects an abnormally large increase in the fluctuation of a user's emotional state, it will quickly and automatically reduce the display brightness and volume, automatically play soothing light music, automatically change the background color of the display screen to a soft blue, and display a prompt message on the device interface to remind the user to adjust their emotions appropriately.

[0074] When the user's emotional state fluctuates slightly, the system will maintain normal display brightness and volume, but will provide periodic positive feedback prompts.

[0075] The system provides the function of saving and reviewing user feedback history, allowing users to view the trend of their emotional state changes;

[0076] The adaptive transient interactivity unit also includes a conflict arbitration module.

[0077] In step S1, in order to reduce signal noise caused by poor electrode contact, the EEG acquisition module also includes an adaptive electrode contact status monitoring module, which can monitor the electrode contact status in real time and issue an alert when the contact resistance exceeds 8kΩ to ensure the signal quality of data acquisition.

[0078] Specifically, the system injects extremely weak high-frequency AC detection signals into the human body and acquires the contact impedance of each electrode in real time through a synchronous demodulation circuit. During interpolation, to overcome the limitations of traditional graph neural networks that rely solely on static physical distance, thereby significantly improving the interpolation accuracy and topological representation of emotional signals, the system constructs a dynamic adjacency matrix. This dynamic matrix not only depends on the physical Euclidean distance between electrodes, but more importantly, it extracts and incorporates the global phase synchronization index between channels under different emotional frequency bands in real time. This means that when two physically distant brain regions exhibit high neural functional synchronization under specific emotional stimulation, their mutual reference weights in the graph convolutional network will show a non-linear surge, thus achieving accurate repair based on brain functional connectivity rather than purely physical connectivity. To adapt to the computing power limitations of mobile communication devices and completely solve the representation ambiguity problem caused by directly using generalized data, the spatial interpolation process abandons the conventional global retraining approach. Instead, it uses the user's personal historical high-quality data as exclusive supporting samples and performs lightweight deployment on the local terminal. During system initialization, a user-specific spatial covariance prior matrix extracted from historical high-quality data is pre-loaded. During real-time interpolation, the graph convolutional network only performs targeted fine-tuning updates on the eigenvalues ​​and key diagonal elements of the prior matrix. This mechanism essentially uses a small amount of real-time, high-quality channel signals to dynamically activate pre-stored individual-specific brain network patterns, ensuring that the interpolated data accurately reproduces the user's current real EEG transient microstate, rather than simply spatial smoothing, blind inference from neighboring points, or rigid averaging of historical data.

[0079] In step S2, the signal preprocessing module performs filtering and noise suppression on the EEG signal, including the following steps: First, a Butterworth high-pass filter with a cutoff frequency of 0.5Hz and a Butterworth low-pass filter with a cutoff frequency of 45Hz are applied to the original EEG signal to remove extremely low-frequency DC drift and high-frequency electromyography interference. Then, a 50Hz Butterworth notch filter with a bandwidth of 48-52Hz is used to suppress power frequency noise. Further, ICA is applied to separate electrooculography, electromyography, and other interfering signals from the EEG signal. The decomposition dimension of the ICA is set to be consistent with the number of electrode channels. Assuming a total of 32 electrode channels are used, the number of independent components in the ICA is also set to 32. This method can typically effectively remove more than 80% of the interfering components from the original EEG signal, improving signal quality and thus enhancing the performance of subsequent emotion recognition models.

[0080] In step S3, the data augmentation module improves the recognition accuracy of the emotion recognition model through data augmentation techniques. The data augmentation methods include three parts: time pruning and shifting, adding Gaussian noise, and signal stretching and compression. First, a segment is randomly selected from each EEG signal sequence in the training set. Specifically, a starting point is randomly determined on the time axis of the signal sequence, and a fixed-length signal segment is pruned from that point. Then, the pruned signal segment undergoes a random time shift operation, causing a slight offset on the time axis to improve the model's robustness to time variations. Next, Gaussian noise is added to the pruned and shifted signal segment. Finally, the noise-enhanced signal segment is slightly stretched or compressed to simulate physiological differences between individuals. This improves the model's generalization ability, allowing it to better adapt to the differences in EEG signals among different individuals. The formulas for the above process are as follows:

[0081] Where X is the original EEG signal sequence, representing a signal sample of the model in the training set, with dimensions C×T, where C represents the number of EEG signal channels and T represents the number of time points; Xaug is the augmented EEG signal sequence, i.e., the sample used for training after data augmentation; Crop(X,s,l) is the pruning function, which randomly selects a starting point s from the signal sequence X, where s is randomly selected from the range [0,T−l] to ensure that a suitable segment is obtained within the time length l, where l=0.6T, to ensure that rich temporal information can be obtained within the time span of the original signal; Shift(·,δ ) is a time shift function that slightly shifts the clipped signal segment along the time axis, δ∈[−0.1T,0.1T], thus avoiding excessive signal distortion; N is the added Gaussian noise, representing a Gaussian noise sequence used to simulate a noisy environment in a real-world scenario. The mean of the Gaussian noise is 0, and σ in the variance σ² is 3%; Resample(·,α) is a stretching / compression function that resamples the noisy signal to simulate physiological differences between individuals. α is randomly selected within the range [0.9,1.1] to control the resampling ratio and achieve the stretching or compression effect.

[0082] In step S4, the feature extraction module calculates the time-domain features (mean potential, standard deviation, peak value, peak-to-peak value, root mean square, variance, skewness, kurtosis, interquartile range, maximum value, minimum value, absolute deviation of mean, energy, integral, number of zero crossings, amplitude envelope, autocorrelation coefficient, sum of squares of amplitude, signal entropy, Lempel-Ziv complexity) and frequency-domain features (relative energy, total power, main frequency band power, instantaneous frequency, power spectral density, phase consistency, frequency offset, frequency center, bandwidth, harmonic ratio, frequency entropy, frequency band energy ratio, short-time Fourier transform energy, etc.) of the EEG signal. The feature extraction module integrates the following features: weighted frequency, peak frequency, signal amplitude spectrum, signal phase spectrum, Hilbert transform energy, frequency domain entropy, and resonant frequency; and spatial features (spatial filtering features, channel relative strength, local field potential, mutual information, differential entropy, spatial energy, linear discriminant features, symmetry index, channel cross-correlation, channel coherence, brain power source localization, spatial covariance, independent component analysis, spatial pattern recognition, differential variance, channel skewness, channel kurtosis, nonlinear exponent, feature symmetry, and neural signal spatial entropy). These features are then combined in the time, frequency, and spatial domains and input into the emotion recognition module. The calculation formula for the feature extraction module is as follows:

[0083] in, The final feature matrix will serve as the input to the emotion recognition module, including all time-domain, frequency-domain, and spatial-domain features; , , These are the weighting coefficients for time-domain, frequency-domain, and spatial-domain features, respectively, used to represent the importance of these features in emotion recognition. These are specific values ​​of time-domain characteristics, including mean potential, standard deviation, peak value, etc. Let be the weight coefficient corresponding to the i-th time-domain feature, representing the importance of that feature in emotion recognition; These are specific values ​​of frequency domain characteristics, including relative energy, total power, and main band power. The weight coefficient corresponding to the i-th frequency domain feature represents the importance of this feature in emotion recognition; These are specific values ​​of spatial characteristics, including spatial filtering characteristics, channel relative intensity, local field potential, etc. Let be the weight coefficient corresponding to the i-th spatial feature, representing the importance of that feature in emotion recognition; in addition, the initial domain weight coefficients... , , Both are 0.33, the initial feature weight coefficients , , All values ​​were initially set to 0.05. Later, based on the initial values, the weights were automatically adjusted using a meta-learning algorithm to improve the model's adaptability and accuracy for emotion recognition tasks.

[0084] After obtaining the feature matrix, the system does not simply concatenate them, but assigns initial domain weight coefficients and feature weight coefficients (e.g., both set to initial constants). Subsequently, the system runs a meta-learning algorithm in the background, which uses the user's recent historical feature inputs and actual emotional feedback as a support set to calculate gradients and update weight parameters in real time.

[0085] The meta-learning algorithm comprises a dual-timescale memory unit: a long-term memory (LCM) storing the baseline feature weight distribution of the user's emotions over the past 30 days, and a short-term working memory (SMM) storing real-time feature fluctuations within the last hour. The SSM is updated using a first-in, first-out (FIFO) sliding window queue. Every 5 minutes, the system calculates the relative entropy between the current adaptive feature matrix and the corresponding emotional state in the LCM. If this value exceeds a dynamic threshold, emotional drift is identified. The dynamic threshold decays exponentially with the number of days the device has been used and is adaptively adjusted by incorporating the divergence standard deviation from the recent working memory. When emotional drift occurs, the system executes a first-order model-independent meta-learning update strategy to reduce computational overhead on mobile devices. The update logic incorporates the weights from the LCM as elastic anchors into the loss function, prioritizing gradient updates using data from the SSM. However, due to the constraints imposed by the LCM anchor values, the magnitude of weight updates is limited, preventing the model from overfitting to instantaneous environmental noise.

[0086] Through this dynamic weighting, when the system faces users with significant differences in activity levels across different brain regions, it can adaptively amplify the weights of high-contribution features (such as high-frequency differential entropy) and suppress redundant features, thereby generating an adaptive feature matrix with extremely high individual specificity.

[0087] Between the feature extraction module and the emotion recognition module, there is also an emotion-state-specific feature recombination unit. This recombination unit pre-stores a feature-emotion correlation matrix pre-trained based on a large amount of population data. For each basic emotion, the correlation matrix provides the corresponding optimal feature subset. The system calculates the feature activation intensity of the current EEG signal on each emotion correlation subset in real time. This activation intensity is obtained through a lightweight attention scorer, and the weight matrix of the scorer adopts an orthogonal initialization strategy to avoid gradient vanishing. The system only concatenates the features corresponding to the two emotion subsets with the highest activation intensity and feeds them into the CNN-Transformer model; inactive feature subsets are discarded, and the discard ratio is dynamically determined by the Shannon entropy of the activation vector. When the Shannon entropy is high, indicating that the user is in a complex mixed emotion edge state, the system automatically increases the feature retention ratio, up to 60% of the feature dimension. To solve the problem of variable feature dimension caused by dynamic discarding, a dynamic projection layer based on a multilayer perceptron is placed before the CNN input to uniformly map the variable-length retained features to a fixed-dimensional latent vector space. This mechanism reduces the computational load at the front end of the system while fundamentally preventing the loss of key emotional boundary information.

[0088] In step S5, the emotion recognition module uses a deep hybrid model consisting of CNN and Transformer for emotion classification. The CNN module includes three convolutional layers: the first layer has 32 kernels (1×3 size) and a stride of 1; the second layer has 64 kernels (1×3 size) and a stride of 1; and the third layer has 128 kernels (1×3 size) and a stride of 1. Each convolutional layer is followed by a max-pooling layer (1×2 kernel size, stride of 2), and a ReLU activation function is applied after each convolutional layer. The Transformer module includes an input embedding layer that converts the feature sequence output from the convolutional layers into a 128-dimensional embedding representation, and uses a 6-layer Transformer encoder, each layer including 8 attention heads, each head having a dimension of 16. A fully connected layer is added after the last Transformer encoder layer to map the output to the emotion category space. In this way, the CNN module can gradually extract local features from the 3×20 feature matrix, and the Transformer module is used to model the global dependencies between features, ultimately completing the emotion classification.

[0089] In deep hybrid models combining CNN and Transformer, the CNN module is specifically limited to using a 1×3 one-dimensional lateral convolution kernel. EEG signals have extremely strong temporal correlation and inter-channel phase synchronization, and traditional two-dimensional convolutions (such as 3×3) are prone to information aliasing in both spatial and temporal dimensions. The 1×3 convolution kernel used in this invention is specifically designed for multi-channel EEG sequences, sliding only along the time step without mixing across channels, thus strictly preserving the transient phase relationship between different EEG channels in the early stages of feature dimensionality reduction;

[0090] The CNN module also includes a phase consistency preservation unit. Before convolution, this unit first calculates the phase-lock value between each pair of EEG channels and introduces this value as a regularization term into the CNN's loss function to ensure that the change in the phase difference matrix before and after convolution is minimized. Specifically, before performing convolution calculation, the system first performs a Hilbert transform on the multi-channel EEG signal to be processed, extracts the transient analytical phase in a specific emotion-related frequency band, and calculates the phase-lock value between any two channel pairs, thereby constructing a baseline phase topology map. The weight coefficient of the regularization term introduced in the loss function is dynamically scaled according to the gradient magnitude of the current backpropagation and the real-time signal-to-noise ratio of the front-end acquired signal. During the forward computation of the convolutional layer, the feature map needs to undergo a nonlinear transformation through a differentiable phase-limited activation function. This activation function can maintain the coherence of the gradient during network backpropagation, thereby forcibly constraining the phase difference change of different channel feature maps within the same frequency band within a preset phase distortion threshold. The output feature map of the convolutional layer needs to pass through a phase reconstruction verification layer, which verifies whether the output features still maintain the cross-channel phase topology relationship at the time of input. This verification layer employs a Siamese network architecture, containing two parallel fully connected mapping paths that are structurally symmetrical and share weights. The first path receives the original input features before convolution, and the second path receives the feature map output from the convolutional layer. Both paths, through a processing procedure involving two layers of linear projection, convert the input features into fixed-dimensional sentiment topological latent vectors, thus mapping features from different stages to the same metric space. The system quantitatively evaluates the topological isomorphism of these two latent vectors on the cross-channel phase topology map by calculating the cosine similarity in this metric space. If the calculated cosine similarity value is lower than 0.92, it indicates that the current convolutional operation has caused distortion in the model's internal representation. In this case, the system immediately triggers closed-loop self-healing adjustment at the current layer and simultaneously initiates micro-step iteration optimization. To ensure the response efficiency of intelligent communication devices when processing real-time EEG signals and avoid the huge computational overhead and high latency caused by backpropagation of the global network, the iterative optimization here is strictly limited to the current non-compliant single convolutional layer. The system directly isolates the weight tensor of this convolutional layer, calculates the local partial derivative only with respect to the current topological isomorphism score, and performs targeted gradient updates on the parameters of the convolutional kernel of this layer with a very small learning rate. The number of iterations for this fast update calculation is forcibly set to no more than 3. If the target is still not met after 3 updates, the current frame data is discarded and feature re-extraction is triggered, thereby fundamentally solving the problem of internal model representation distortion while ensuring low computational consumption on mobile devices.

[0091] In step S6, the intelligent response module has an adaptive steady-state interactivity unit. The intelligent response module continuously identifies the user's emotions and adjusts the display mode of the communication device or provides voice broadcast reminders based on the identification results. Specifically, this includes:

[0092] (1) When the user’s mood is detected to be in a state of continuous depression, youthful and sunny music will be played automatically, and the background color of the display screen will be automatically changed to a comfortable orange to alleviate the user’s mood.

[0093] To overcome the drawback of conventional deep learning models, which consume excessive time for forward inference and cannot meet the immediate need for soothing acute emotional changes, the system establishes a fast response data channel that is completely bypassed from the main classification process before executing transient responses. To eliminate the interference of high-frequency electromyography artifacts on the instantaneous rate of change, the system first uses the Savitzky-Golay multinomial filtering algorithm to smooth low-level features such as nonlinear exponents and difference variances within a 50-millisecond sliding time window, and then calculates the first derivative as the instantaneous rate of change. Once the instantaneous rate of change of a feature exceeds a preset threshold, the transient control command skips the conventional operating system application layer UI rendering queue and directly triggers the hardware interrupt interface of the underlying hardware abstraction layer. This hardware-software collaborative underlying calling mechanism achieves millisecond-level feedforward physical environment control, ensuring that the user receives initial physical relief before the final emotion label is output.

[0094] (2) When the user’s emotions are detected to be in a state of continuous excitement, soothing light music will be played automatically, the background color of the display screen will be automatically changed to a soft blue, and a prompt message will be displayed on the device interface to remind the user to adjust their emotions appropriately.

[0095] (3) Users can customize the response method of emotion reminders in the device settings and can enable or disable the voice broadcast function.

[0096] In step S7, the intelligent response module has an adaptive transient interactivity unit, including the following:

[0097] (1) When the system detects an abnormal increase in the fluctuation of the user's emotional state, it will quickly and automatically reduce the display brightness and volume, automatically play soothing light music, automatically change the background color of the display screen to a soft blue, and display a prompt message on the device interface to remind the user to adjust their emotions appropriately.

[0098] To overcome the drawback of conventional deep learning models, which consume excessive time for forward inference and cannot meet the immediate need for soothing acute emotional changes, the system establishes a fast response data channel that completely bypasses the Transformer main classification process before executing transient responses. To eliminate the interference of high-frequency electromyography artifacts on the instantaneous rate of change, the system first uses the Savitzky-Golay multinomial filtering algorithm to smooth low-level features such as nonlinear exponents and difference variances within a 50-millisecond sliding time window, and then calculates the first derivative as the instantaneous rate of change. Once the instantaneous rate of change of a feature exceeds a system-preset threshold, the transient control command skips the conventional operating system application layer UI rendering queue and directly triggers the hardware interrupt interface of the underlying hardware abstraction layer. This hardware-software collaborative underlying calling mechanism achieves millisecond-level feedforward physical environment control, ensuring that the user receives initial physical relief before the final emotion label is output.

[0099] (2) When the system detects that the fluctuation of the user's emotional state is small, it will maintain the normal display brightness and volume, but provide timely positive feedback prompts, such as "I am very stable", to achieve a positive guidance effect.

[0100] (3) The system provides the function of saving and reviewing user feedback history, so that users can view the trend of their emotional state changes to help regulate their emotions.

[0101] (4) The adaptive transient interactivity unit also includes a conflict arbitration module:

[0102] When the transient path detects that the instantaneous rate of change exceeds the threshold, the observation window of the recombination unit is activated. Within the observation window, the second-order difference of the nonlinear exponent and the skin conductance response are extracted simultaneously as supporting evidence. If the sign inversion of the second-order difference exceeds 2 times or the supporting evidence does not support emotional abnormality, the feedforward instruction is suppressed. If the transient instruction contradicts the classification result output by the emotion recognition module after a subsequent delay of 200ms, the system records a transient-steady-state inconsistency event and fine-tunes the transient threshold in the reverse direction. When the number of inconsistency events exceeds the preset number within 1 hour, the system automatically switches the transient response mode from immediate execution to arbitration before execution and generates a prompt asking the user whether to close the transient interaction.

[0103] To address the risk of feedforward control being erroneously triggered by non-emotional physiological artifacts in human-computer interaction scenarios, the conflict arbitration module employs specific temporal fuzziness suppression logic: It calculates the confidence level of the second-order difference sign inversion using a Gaussian membership function. If this confidence level is lower than a preset fuzzy membership threshold, and the slope characteristics of the skin conductance response signal synchronously acquired through the analog-to-digital converter interface of the intelligent communication device show a negative correlation within the observation window, it indicates that the skin conductance has not significantly increased, lacking the physiological manifestation of genuine sympathetic nerve arousal. In this case, the system determines that the EEG fluctuation is triggered by a non-emotional random artifact and immediately generates a veto command, forcibly intercepting and withdrawing the already issued feedforward control command at the hardware driver level. Furthermore, if the uninterrupted transient command contradicts the final classification result output by the emotion recognition module after a 200ms delay, the system records a transient-steady-state inconsistency event and triggers a reverse fine-tuning of the transient threshold. The fine-tuning of this threshold follows an exponential decay scaling rule based on the frequency of inconsistent events, enabling the system to automatically tighten the feedforward triggering conditions after encountering frequent misjudgments.

[0104] In step S8, the specific configuration used for training and testing the emotion recognition model is as follows: the batch size for both the training and test sets is set to 64 to ensure the stability of batch gradient descent during model training. The Adam optimizer is used for training, with an initial learning rate of 0.0001 and a learning rate decay coefficient of 0.0001 to ensure the model can gradually reduce the step size during convergence, preventing overfitting. The training process involves 1000 iterations, with each iteration tested on the validation set to monitor the model's real-time performance on the test set.

[0105] The adaptive transient interactivity unit was designed to resolve the conflict between the computational latency of deep learning models and the need for immediate reassurance of users experiencing sudden emotional shifts. This unit establishes a rapid response channel that bypasses the main classification process of the Transformer, directly monitoring the nonlinear exponential and difference variance features in the adaptive feature matrix in real time at the front end. When the instantaneous rate of change of these two features exceeds a preset threshold within 100 milliseconds, the system determines that the user is experiencing an acute emotional fluctuation. It then bypasses the upper-layer application and sends control commands forward to the underlying operating system, instantly lowering the screen brightness and volume. This feedforward control mechanism ensures that the user receives initial physical relief before the final emotion label is output. The conflict arbitration module further addresses potential false triggering issues in the feedforward control, improving system reliability and user experience.

[0106] Furthermore, in step S8, after completing model training, the model is tested and evaluated, including the following:

[0107] (1) Repeat the classification 5 times for each test set sample, and calculate the average of the 5 classification results as the final prediction result to reduce the impact of random factors on the test results.

[0108] (2) Use confusion matrix and F1 score to evaluate the classification accuracy of the model on the test set to ensure that the classification model performs evenly across different emotion categories.

[0109] (3) Display the evaluation results on the device interface and provide the emotion recognition accuracy and confusion matrix so that users can understand the recognition performance of the model.

Claims

1. An emotion perception assistance system and method for intelligent communication devices, characterized in that, The specific steps include the following: S1: The EEG acquisition module is used to collect the user's EEG signals. The EEG acquisition module is configured as a multi-channel module to collect emotion-related electrical signals from different brain regions of the user in a wireless transmission manner, with a sampling frequency of not less than 512Hz. S2: The signal preprocessing module filters and suppresses noise in the EEG signals acquired by the EEG acquisition module; S3: The data augmentation module performs time clipping and time shifting, adds Gaussian noise, and stretches and compresses the EEG signal processed by the signal preprocessing module. S4: The feature extraction module extracts and integrates multidimensional features of EEG signals from the time domain, frequency domain, and spatial domain; In the feature integration stage, initial domain weight coefficients and feature weight coefficients are preset, and the weight ratio between the time domain, frequency domain and spatial domain features is automatically and dynamically adjusted based on the user's historical emotional fluctuation data through a meta-learning algorithm to generate an adaptive feature matrix. S5: The emotion recognition module uses a deep hybrid model composed of a convolutional neural network (CNN) and a self-attention model (Transformer) to classify emotions based on the multidimensional features of EEG signals. The CNN module uses a 1×3 convolutional kernel to extract local features from the adaptive feature matrix. During the extraction process, it avoids disrupting the transient phase relationship between channels in different brain regions and converts the extracted local features into an embedded representation, which is then input into the Transformer module for global dependency modeling. The CNN module also includes a phase consistency preservation unit. Before the convolution operation, this unit first calculates the phase-locking value between each pair of EEG channels and introduces this phase-locking value as a regularization term into the loss function of the CNN. During the convolution process, the phase difference of the feature maps of different channels within the same frequency band is forced to not exceed a preset phase distortion threshold. The output feature map of the convolutional layer needs to pass through a phase reconstruction verification layer. This verification layer verifies whether the output features still maintain the cross-channel phase topology relationship at the time of input. If the phase topology similarity is lower than the preset similarity value, feature re-extraction is triggered. S6: The intelligent response module has an adaptive steady-state interactive unit, which continuously identifies the results of the user's emotion classification and adjusts the display mode of the communication device or provides voice broadcast reminders based on the identification results; The intelligent response module possesses adaptive steady-state interactivity. It continuously identifies the user's emotions and adjusts the display mode of the communication device or provides voice prompts based on the identification results, including the following steps: When the system detects that the user is in a persistently low mood, it will automatically play upbeat and cheerful music, and the background color of the display screen will automatically change to a comfortable orange to alleviate the user's mood. When the device detects that the user is in a state of sustained agitation, it will automatically play soothing light music, the background color of the display screen will automatically change to a soft blue, and a prompt message will be displayed on the device interface to remind the user to adjust their emotions appropriately. Users can customize the response of emotion reminders in the device's settings and can enable or disable the voice broadcast function; S7: The intelligent response module has an adaptive transient interaction unit. The adaptive transient interaction unit extracts the nonlinear exponential or differential variance features in the adaptive feature matrix in real time. When the instantaneous change rate of the above features exceeds the preset threshold, it is determined to be an abnormal emotional fluctuation. Before the emotion recognition module outputs the final classification result, it feeds forward to control the underlying driver of the intelligent communication device to quickly reduce the display brightness and volume. The intelligent response module possesses adaptive transient interactivity, including the following steps: When the system detects an abnormally large increase in the fluctuation of a user's emotional state, it will automatically and quickly reduce the display brightness and volume, automatically play soothing light music, automatically change the background color of the display screen to a soft blue, and display a prompt message on the device interface to remind the user to adjust their emotions appropriately. When the user's emotional state fluctuates slightly, the system will maintain normal display brightness and volume, but will provide periodic positive feedback prompts. The system provides the function of saving and reviewing user feedback history, allowing users to view the trend of their emotional state changes; The adaptive transient interactivity unit also includes a conflict arbitration module: when the transient path detects that the instantaneous rate of change exceeds a threshold, the observation window of the reorganization unit is activated; within the observation window, the second-order difference of the nonlinear exponent and the skin conductance response are extracted simultaneously as supporting evidence; if the sign inversion of the second-order difference exceeds 2 times or the supporting evidence does not support emotional abnormality, the feedforward instruction is suppressed; if the transient instruction contradicts the classification result output by the emotion recognition module after a subsequent delay of 200ms, the system records a transient-steady-state inconsistency event and fine-tunes the transient threshold in the reverse direction; when the number of inconsistency events exceeds a preset number within 1 hour, the system automatically switches the transient response mode from immediate execution to arbitration before execution and generates a prompt asking the user whether to close the transient interaction.

2. The emotion perception assistance system according to claim 1, characterized in that, In S1, the EEG acquisition module also includes an adaptive electrode contact status monitoring module, used to monitor the electrode contact status in real time, reduce signal noise caused by poor electrode contact, and issue an alert when the contact resistance exceeds 8kΩ; the EEG acquisition module also includes a data repair channel driven by electrode contact quality. When the contact resistance of a certain channel exceeds 10kΩ within 5 consecutive sampling points, the system marks that channel as a low-confidence channel; Instead of discarding the data for that channel, the EEG signals from adjacent spatial channels are used to generate a replacement signal for the low-confidence channel in real time using a spatial interpolation model based on a graph convolutional network. The spatial interpolation model uses high-quality historical data from the user as training samples to learn a nonlinear mapping function for each electrode position. If more than three channels are marked as low-confidence at the same time, emotion recognition is paused and the user is prompted to check the electrodes.

3. The emotion perception assistance system according to claim 1, characterized in that, In step S2, the signal preprocessing module performs filtering and noise suppression on the EEG signal, including the following steps: First, a Butterworth high-pass filter and a Butterworth low-pass filter are applied to the original EEG signal to remove extremely low-frequency DC drift and high-frequency electromyography interference. Secondly, a Butterworth notch filter is used to suppress power frequency noise; Finally, ICA was further applied to separate electrooculography, electromyography and other interference signals from the EEG signal. The decomposition dimension of ICA was set to be consistent with the number of electrode channels.

4. The emotion perception assistance system according to claim 1, characterized in that, In step S3, the data augmentation module improves the recognition accuracy of the emotion recognition model through data augmentation techniques. The data augmentation method includes three parts: time pruning and time shifting, adding Gaussian noise, and signal stretching and compression, and includes the following steps: First, randomly select a segment from each EEG signal sequence, randomly determine a starting point on the time axis of the signal sequence, and then cut out a signal segment of fixed length from that point. Then, a random time shift operation is performed on the cropped signal segment to produce a slight offset on the time axis, thereby improving the model's robustness to time changes. Next, Gaussian noise is added to the clipped and shifted signal segments; Finally, the noise-enhanced signal segments are slightly stretched or compressed to simulate physiological differences between individuals.

5. The emotion perception assistance system according to claim 1, characterized in that, In step S4, the feature extraction module calculates the time-domain features, frequency-domain features, and spatial-domain features of the EEG signal, integrates the above features according to the time domain, frequency domain, and spatial domain, and then inputs them into the emotion recognition module. The time-domain features include mean potential, standard deviation, peak value, peak-to-peak value, root mean square, variance, skewness, kurtosis, interquartile range, maximum value, minimum value, absolute deviation of mean, energy, integral, number of zero crossings, amplitude envelope, autocorrelation coefficient, sum of squared amplitudes, signal entropy, and Lempel-Ziv complexity. The frequency characteristics include relative energy, total power, main band power, instantaneous frequency, power spectral density, phase consistency, frequency offset, frequency center, bandwidth, harmonic ratio, frequency entropy, band energy ratio, short-time Fourier transform energy, weighted frequency, band peak value, signal amplitude spectrum, signal phase spectrum, Hilbert transform energy, frequency domain entropy, and resonant frequency. The spatial features include spatial filtering features, channel relative strength, local field potential, mutual information, differential entropy, spatial energy, linear discriminant features, symmetry index, channel cross-correlation, channel coherence, brain power source localization, spatial covariance, independent component analysis, spatial pattern recognition, differential variance, channel skewness, channel kurtosis, nonlinear exponent, feature symmetry, and neural signal spatial entropy.

6. The emotion perception assistance system according to claim 1, characterized in that, In S5, the emotion recognition module uses a deep hybrid model composed of CNN and Transformer for emotion classification; The CNN module includes three convolutional layers. The first layer has 32 convolutional kernels, a kernel size of 1×3, and a stride of 1. The second layer has 64 convolutional kernels, a kernel size of 1×3, and a stride of 1. The third layer has 128 convolutional kernels, a kernel size of 1×3, and a stride of 1. Each convolutional layer is followed by a max pooling layer with a kernel size of 1×2 and a stride of 2. The ReLU activation function is applied after each convolutional layer. The Transformer module includes an input embedding layer that converts the feature sequence output by the convolutional layer into an embedding representation with 128 dimensions, and uses a 6-layer Transformer encoder, each layer including 8 attention heads, each head having a dimension of 16. A fully connected layer is added after the last Transformer encoder layer to map the output to the sentiment category space.

7. The emotion perception assistance system according to claim 1, characterized in that, It also includes S8: when training and testing the emotion recognition module, determining the batch size of the training set and the test set, the optimal optimizer, the number of iterations, the learning rate, and the weight decay coefficient required to train the model; in step S8, when training and testing the emotion recognition model, the batch size of both the training set and the test set is set, the Adam optimizer is used for training, the initial learning rate is set, the learning rate decay coefficient is preset, and each iteration during training is tested on the validation set.

8. The emotion perception assistance system according to claim 7, characterized in that, In step S8, after model training is completed, the model is tested and evaluated, including the following steps: The classification process is repeated for each test set sample, and the average of the classification results is calculated as the final prediction result. The confusion matrix and F1 score were used to evaluate the classification accuracy of the model on the test set. The evaluation results are displayed on the device interface, providing emotion recognition accuracy and confusion matrix.

9. The emotion perception assistance system according to claim 1, characterized in that, The meta-learning algorithm in S4 includes a dual-timescale memory unit: a long-term memory (LTM) stores the baseline feature weight distribution of the user's emotions over the past 30 days, and a short-term working memory (STM) stores the real-time feature fluctuations over the most recent hour. Every 5 minutes, the system calculates the KL divergence between the current adaptive feature matrix and the corresponding emotional state in the LTM. If the divergence exceeds a dynamic threshold, emotional drift is identified. In this case, the meta-learning algorithm prioritizes updating the weights using the gradient from the STM and uses the weights from the LTM as a regularization term to prevent overfitting to instantaneous noise. The dynamic threshold is adaptively adjusted based on the user's usage time: an initial threshold is used, and the threshold decreases after more than 7 days of use.

10. The emotion perception assistance system according to claim 1, characterized in that, Between the feature extraction module and the emotion recognition module, an emotion-state-specific feature recombination unit is inserted: this recombination unit pre-stores a feature-emotion correlation matrix pre-trained based on a large amount of population data. For each basic emotion, the correlation matrix provides the corresponding optimal feature subset. The system calculates the feature activation intensity of the current EEG signal on each emotion correlation subset in real time, and only concatenates the features corresponding to the two emotion subsets with the highest activation intensity before sending them into the CNN-Transformer model. Inactive feature subsets are discarded, and the discard ratio is dynamically adjusted to retain 30% of the features to prevent information loss.