Bioelectric emotion recognition method and device based on AR glasses and electronic equipment

By integrating sensor devices on AR glasses, EEG, and electrocutaneous skin reaction data were obtained and emotional recognition was used using CNN+LSTM models, the problem that single modal data was difficult to capture emotional diversity was solved, and high accuracy and portability of emotion recognition was achieved.

CN120131015APending Publication Date: 2025-06-13GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510260628.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing emotion recognition methods rely on single modal data, making it difficult to fully capture the diversity of emotions, resulting in low recognition accuracy.

Method used

Using a bioelectric emotion recognition method based on AR glasses, we can obtain EEG, electromyography and skin electroreaction data and combine the CNN+LSTM fusion model to identify relative energy, time-frequency characteristics and electrical reaction characteristics.

Benefits of technology

Through multimodal data fusion, the accuracy and robustness of emotion recognition are improved, the risk of misjudgment of a single signal is reduced, and the miniaturization and portability of the emotion recognition system is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120131015A_ABST
    Figure CN120131015A_ABST
Patent Text Reader

Abstract

The invention provides a bioelectric emotion recognition method and device based on AR glasses and electronic equipment, and relates to the field of data processing. In the method, bioelectricity data for a user sent by sensor equipment is obtained, the sensor equipment is located on AR glasses worn by the user, and the bioelectricity data comprises electroencephalogram data, myoelectricity data and skin electric response data; calculating relative energy corresponding to the preset wave band according to the electroencephalogram data; according to the myoelectricity data, time-frequency characteristics of myoelectricity activities are determined; determining electric reaction characteristics according to the skin electric reaction data; and adopting a CNN + LSTM fusion model to perform emotion recognition on the relative energy, the time-frequency characteristics and the electric reaction characteristics to obtain a recognition result. By implementing the technical scheme provided by the invention, the accuracy of emotion recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a bioelectric emotion recognition method, device and electronic device based on an AR glasses. Background Art

[0002] Emotion recognition technology is an important direction in current human-computer interaction, intelligent health management and psychological research, and is widely used in scenarios such as emotion regulation, stress monitoring and personalized services.

[0003] Currently, most existing emotion recognition methods rely on single-modal data sources. However, since emotion is a complex mental activity that often involves multi-dimensional physiological and behavioral changes, it is difficult to comprehensively capture the diversity of emotions only relying on single-modal data, resulting in low accuracy of emotion recognition.

[0004] Therefore, there is an urgent need for a bioelectric emotion recognition method, device and electronic device based on an AR glasses. Summary of the Invention

[0005] This application provides a bioelectric emotion recognition method, device and electronic device based on an AR glasses, which is convenient for improving the accuracy of emotion recognition.

[0006] In the first aspect of this application, a bioelectric emotion recognition method based on an AR glasses is provided. The method includes: obtaining bioelectric data for a user sent by a sensor device, where the sensor device is located on the AR glasses worn by the user, and the bioelectric data includes electroencephalogram data, electromyogram data and skin conductance response data; calculating the relative energy corresponding to a preset frequency band according to the electroencephalogram data; determining the time-frequency characteristics of electromyogram activity according to the electromyogram data; determining the electroresponse characteristics according to the skin conductance response data; using a CNN+LSTM fusion model to perform emotion recognition on the relative energy, the time-frequency characteristics and the electroresponse characteristics to obtain a recognition result.

[0007] By adopting the above technical solutions, by fusing electroencephalogram (EEG) data, electromyogram (EMG) data, and skin conductance response (SCR) data, this method can capture the emotional changes of users from multiple dimensions. Compared with emotion recognition based on single-modal data, this multi-modal fusion strategy greatly improves the comprehensiveness and accuracy of the description of emotional states, and reduces the risk of misjudgment caused by the absence or distortion of a single signal. By using a CNN+LSTM fusion model, combining the spatial feature extraction ability of CNN and the temporal dynamic modeling ability of LSTM, the model can not only efficiently extract static features but also capture the dynamic characteristics of emotions changing over time, thus significantly improving the accuracy and robustness of emotion recognition. Integrating the sensor device into the AR glasses worn by users realizes the miniaturization and portability of the emotion recognition system. Combining the real-time data acquisition and processing capabilities, this solution can quickly respond to the dynamic changes of users' emotions, providing the possibility for real-time emotion monitoring and feedback. The combination of multi-modal features effectively addresses the impacts of environmental noise and individual differences. Therefore, it is convenient to improve the accuracy of emotion recognition.

[0008] Optionally, the obtaining of the bioelectrical data of the user sent by the sensor device specifically includes: receiving the first data of the user sent by the EEG electrode, where the EEG electrode is used for EEG detection of the forehead area and the area behind the ear of the user; receiving the second data of the user sent by the EMG sensor, where the EMG sensor is used for EMG detection of the front face area of the user; receiving the third data of the user sent by the SCR sensor, where the SCR sensor is used for SCR detection of the side face area of the user; and performing denoising, data segmentation, and baseline correction on the first data, the second data, and the third data to obtain the bioelectrical data.

[0009] By adopting the above technical solutions, by integrating electroencephalogram (EEG) electrodes, electromyography (EMG) sensors, and galvanic skin response (GSR) sensors, the forehead area, the area behind the ears, the front and side areas of the face of the user are respectively detected, covering multiple key physiological signal sources during emotional changes. Such a layout design ensures comprehensive monitoring of the user's emotional state, enhancing the accuracy and depth of recognition. The optimized arrangement of each sensor for a specific part can more accurately collect bioelectric signals related to emotions. This targeted layout can reduce the interference of irrelevant signals and improve the quality and reliability of signal acquisition. Denoising, data segmentation, and baseline correction are performed on the received first, second, and third data, significantly improving the purity and consistency of the bioelectric signals. The preprocessing process can effectively eliminate noise interference and baseline drift problems, thus providing more stable input data for subsequent emotional feature extraction and recognition. This method ensures the synchronization and consistency of multimodal data by uniformly processing EEG, EMG, and GSR signals. Even if there are fluctuations or noises in one modal signal, other modal data can still provide supplementary information, thereby enhancing the robustness and reliability of the emotion recognition system.

[0010] Optionally, calculating the relative energy corresponding to a preset frequency band according to the EEG data specifically includes: parsing the EEG data to obtain an EEG signal; determining the preset frequency band according to the frequency band corresponding to the EEG signal; converting the EEG signal into a spectrum through fast Fourier transform according to the preset frequency band; and determining the ratio between the spectrum density corresponding to the preset frequency band and the total spectrum density according to the spectrum to obtain the relative energy corresponding to the preset frequency band.

[0011] By adopting the above technical solutions, by calculating the relative energy of a preset frequency band, this method can extract features with high emotional correlation from EEG signals. This targeted extraction method can more accurately reflect the emotional state and improve the effectiveness and pertinence of emotion recognition. Using fast Fourier transform to convert the EEG signal in the time domain into a frequency domain spectrum makes signal analysis more intuitive and efficient. The calculation is fast and stable, suitable for real-time signal processing scenarios, meeting the high requirements of the emotion recognition system for response speed. By calculating the ratio of the spectrum density of the preset frequency band to the total spectrum density to obtain the relative energy, this method effectively avoids the deviation caused by the amplitude difference of individual EEG signals. As a normalized feature, the relative energy has better robustness and adaptability, suitable for multi-user emotion recognition scenarios. This solution supports flexible setting of the preset frequency band according to specific application requirements, meeting the recognition requirements for different emotional states. In this way, the most relevant frequency band can be selected according to the task scenario to further optimize the emotion recognition performance.

[0012] Optionally, determining the time-frequency characteristics of the myoelectric activity according to the myoelectric data specifically includes: determining the average absolute value of the muscle activity intensity and the myoelectric peak according to the myoelectric data; generating time-domain characteristics based on the average absolute value of the muscle activity intensity and the myoelectric peak; calculating the frequency range and the frequency center of the muscle activity through fast Fourier transform according to the myoelectric data to generate frequency-domain characteristics; and obtaining the time-frequency characteristics according to the time-domain characteristics and the frequency-domain characteristics.

[0013] By adopting the above technical solution, by extracting the time-domain characteristics and the frequency-domain characteristics of the myoelectric data, this method can comprehensively characterize the muscle activity characteristics from two dimensions of time and frequency. This combination method avoids the insufficient information that may be caused by single-dimensional characteristics and improves the description accuracy of the myoelectric activity. Time-domain characteristics such as the average absolute value and the peak value of the muscle activity intensity can intuitively reflect the intensity and dynamic changes of muscle contraction. By extracting the frequency range and the frequency center of the muscle activity through fast Fourier transform, the frequency-domain characteristics can reveal the dynamic law of the muscle activity, and these frequency information can help identify the micro myoelectric signal characteristics in the emotional changes. Combining the time-domain characteristics and the frequency-domain characteristics to generate the time-frequency characteristics can comprehensively reflect the intensity and dynamic change law of the muscle activity. This fusion feature provides a more complete information basis for emotion recognition and improves the accuracy and robustness of the emotion recognition model.

[0014] Optionally, determining the electroresponse characteristics according to the skin electroresponse data specifically includes: calculating the peak conductivity and the conductivity change rate according to the skin electroresponse data; and determining the electroresponse characteristics based on the peak conductivity and the conductivity change rate.

[0015] By adopting the above technical solution, by calculating the peak conductivity and the rate of change of conductivity in the skin conductance response data, this method extracts the core features closely related to emotional changes. For example, when a person is emotionally excited or tense, the skin conductivity will increase significantly, and the rate of change of conductivity can reflect the intensity and speed of emotional fluctuations. These features intuitively reflect the physiological manifestations of emotions. The peak conductivity and the rate of change of conductivity are the direct features of the skin conductance response. The calculation process is simple and efficient, without complex data conversion or modeling steps, suitable for real-time emotion monitoring scenarios, and at the same time reducing the consumption of system resources. As normalized and dynamic features, the peak conductivity and the rate of change of conductivity can effectively reduce the influence of individual skin conductance baseline differences and improve the robustness of emotion recognition. This enables this method to adapt to the application requirements of multiple users and multiple scenarios, such as stress monitoring, anxiety recognition, etc. Both the peak conductivity and the rate of change of conductivity have clear physical meanings: the peak conductivity reflects the maximum level of skin conductance, while the rate of change of conductivity reflects the speed and trend of skin conductance changes. This clear feature definition facilitates establishing a direct connection between physiological changes and emotional states, making the emotion recognition results more interpretable. The skin conductance response has a high time resolution and can quickly respond to emotional changes.

[0016] Optionally, using the CNN+LSTM fusion model to perform emotion recognition on the relative energy, the time-frequency features, and the electro-response features to obtain an identification result, specifically including: through the CNN+LSTM fusion model, performing dimensionality reduction, filtering, and feature flattening on the time-frequency features to obtain a first feature vector; through the CNN+LSTM fusion model, judging the dependence relationship between the relative energy and the electro-response features to determine the short-term dependence relationship and the long-term dependence relationship; generating a second feature vector according to the short-term dependence relationship and the long-term dependence relationship; fusing the first feature vector and the second feature vector and integrating them into a high-dimensional feature through a fully connected layer; mapping the high-dimensional feature into a probability distribution of emotion categories through a Softmax activation function; generating an identification result according to the probability distribution of the emotion categories.

[0017] By adopting the above technical solution, through the CNN+LSTM fusion model, this method utilizes the powerful feature extraction ability of CNN and the temporal dependence modeling ability of LSTM to achieve efficient processing and analysis of multi-modal emotion features. This fusion structure can simultaneously mine the spatial features and temporal characteristics of the data, comprehensively improving the accuracy of emotion recognition. This method respectively performs dimensionality reduction, filtering, flattening, and dependence relationship modeling on relative energy, time-frequency features, and electro-response features, effectively avoiding information redundancy while retaining the key features related to emotions. This refined processing improves the utilization efficiency of multi-modal features and provides a more reliable basis for emotion recognition. Through the temporal modeling ability of LSTM, this method can identify short-term and long-term dynamic dependence relationships in emotion changes. By fusing the first feature vector and the second feature vector and integrating them into high-dimensional features through a fully connected layer, this method realizes the deep association of multi-modal data. High-dimensional features can better capture the differences between different emotion categories, improving the discrimination and accuracy of emotion classification. The Softmax activation function is used to map the high-dimensional features into the probability distribution of emotion categories, making the emotion recognition results more intuitive and easy to interpret.

[0018] Optionally, the method further includes: if it is determined that the recognition result of the user at the current moment indicates a negative and tense emotion, controlling the AR glasses to display a suggestion for deep breathing to prompt the user to take a deep breath to relieve the negative and tense emotion; if it is determined that the recognition result of the user after a preset time period indicates a continuous tension, controlling the AR glasses to display a suggestion to pause watching and take a rest.

[0019] By adopting the above technical solution, by displaying real-time suggestions for deep breathing when detecting the user's negative and tense emotion, this solution can quickly intervene and help the user self-regulate when the emotion just starts to fluctuate, thus preventing the emotion from deteriorating further. This instant feedback mechanism improves the user experience and the humanization level of the system. If it is detected that the user's tense emotion lasts for a period of time, the system further provides a suggestion of "pause watching and take a rest" to help the user withdraw from the scenario that may cause emotional fatigue. This long-term guidance is of great significance for the user's health management, especially in high-intensity work or learning scenarios, which helps prevent stress accumulation. This solution provides different suggestions in stages according to the dynamic changes of the user's emotion state, reflecting the personalized adaptation to the user's needs. The hierarchical intervention in different emotion states improves the pertinence and effectiveness of the suggestions, helping the user better adjust the emotion in specific scenarios.

[0020] In a second aspect of the present application, a bioelectric emotion recognition device based on an AR glasses is provided. The bioelectric emotion recognition device includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire bioelectric data of a user sent by a sensor device, and the sensor device is located on the AR glasses worn by the user. The bioelectric data includes electroencephalogram data, electromyogram data, and skin conductance response data. The processing module is used to calculate the relative energy corresponding to a preset frequency band according to the electroencephalogram data. The processing module is further used to determine the time-frequency characteristics of electromyogram activity according to the electromyogram data. The processing module is further used to determine the electroresponse characteristics according to the skin conductance response data. The processing module is further used to perform emotion recognition on the relative energy, the time-frequency characteristics, and the electroresponse characteristics by using a CNN+LSTM fusion model to obtain a recognition result.

[0021] In a third aspect of the present application, an electronic device is provided. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described above.

[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described above is executed.

[0023] In summary, one or more technical solutions provided in the present application have at least the following technical effects or advantages: By fusing electroencephalogram data, electromyogram data, and skin conductance response data, this method can capture the emotional changes of users from multiple dimensions. Compared with emotion recognition using single-modal data, this multi-modal fusion strategy greatly improves the comprehensiveness and accuracy of the description of emotional states, and reduces the risk of misjudgment caused by the absence or distortion of a single signal. By using a CNN+LSTM fusion model, the spatial feature extraction ability of CNN and the temporal dynamic modeling ability of LSTM are combined, enabling the model to efficiently extract static features and capture the dynamic characteristics of emotions changing over time, thereby significantly improving the accuracy and robustness of emotion recognition. Integrating the sensor device into the AR glasses worn by the user realizes the miniaturization and portability of the emotion recognition system. Combining the real-time data acquisition and processing capabilities, this solution can quickly respond to the dynamic changes of users' emotions, providing the possibility for real-time emotion monitoring and feedback. The combination of multi-modal features effectively addresses the influence of environmental noise and individual differences. Therefore, it is convenient to improve the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1Schematic flowchart of a bioelectric emotion recognition method based on an AR glasses provided by an embodiment of the present application; Figure 2 Another schematic flowchart of a bioelectric emotion recognition method based on an AR glasses provided by an embodiment of the present application; Figure 3 Schematic module diagram of a bioelectric emotion recognition device based on an AR glasses provided by an embodiment of the present application; Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0025] Explanation of reference numerals: 31, acquisition module; 32, processing module; 41, processor; 42, communication bus; 43, user interface; 44, network interface; 45, memory. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0027] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0028] In the description of the embodiments of the present application, the meaning of the term "a plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0029] Emotion recognition technology occupies an important position in current human-computer interaction, intelligent health management and psychological research, and has been widely applied in many fields such as emotion regulation, stress monitoring and personalized services.

[0030] However, most of the existing emotion recognition methods rely only on single-modal data sources, and there are certain limitations in the emotion recognition process. Emotion itself is a complex psychological process, involving changes in multiple dimensions such as physiology and behavior. Therefore, it is difficult to comprehensively and accurately capture the diversity of emotions based on single-modal data alone, thus affecting the accuracy and reliability of emotion recognition.

[0031] To solve the above technical problems, this application provides a bioelectric emotion recognition method based on AR glasses. Referring to Figure 1 , Figure 1 is a schematic flowchart of a bioelectric emotion recognition method based on AR glasses provided by an embodiment of this application. This method is applied to a server and includes steps S110 to S150. The above steps are as follows: S110. Obtain bioelectric data of a user sent by a sensor device. The sensor device is located on an AR glasses worn by the user, and the bioelectric data includes electroencephalogram data, electromyogram data, and skin conductance response data.

[0032] Specifically, the server refers to a central computing system that processes data. It receives bioelectric data from sensors, and processes, analyzes, and applies it. The tasks of this server include data collection, storage, and subsequent emotion analysis, etc. The server needs to receive real-time bioelectric data from the AR glasses sensor device worn on the user through a certain communication protocol (such as wireless connection, Bluetooth, etc.). The AR glasses refer to a glasses device integrated with augmented reality technology. After the user wears it, the glasses can not only display virtual information but also integrate various sensors. The sensors on the glasses can monitor the user's physiological state, such as electroencephalogram, electromyogram, and skin conductance response. The sensor device is installed in the AR glasses and is used to collect the user's bioelectric data in real time, such as electroencephalogram signals, muscle activities, and skin conductance response signals.

[0033] Among them, the electroencephalogram data is the electrical signal of brain activity. In different emotional or cognitive states, the electrical activity patterns of the brain are different, and the electroencephalogram can reflect the user's emotional changes. For example, when in a negative emotion, it may be observed that the brain waves in a specific frequency band are enhanced. The electromyogram data records the electrical signals of muscle activities. Electrical activities are generated when the user makes facial expressions or body movements, and these activities can reflect the user's emotional changes. For example, when nervous, the facial muscles may tighten involuntarily. The skin conductance response is to evaluate the body's stress response by measuring the skin conductivity. When the user experiences emotional changes, especially in a state of anxiety or stress, the skin conductivity will change, usually manifested as an increase in skin sweating and an increase in conductivity.

[0034] In a possible implementation, obtaining biometric data of a user sent by a sensor device specifically includes: receiving first data of the user sent by an electroencephalogram (EEG) electrode, where the EEG electrode is used to perform EEG detection on the forehead area and the area behind the ear of the user; receiving second data of the user sent by an electromyogram (EMG) sensor, where the EMG sensor is used to perform EMG detection on the front face area of the user; receiving third data of the user sent by a galvanic skin response (GSR) sensor, where the GSR sensor is used to perform GSR detection on the side face area of the user; and denoising, segmenting the data, and performing baseline correction on the first data, the second data, and the third data to obtain the biometric data.

[0035] Specifically, an EEG electrode is a sensor used to collect brain electrical activities. By placing electrode patches on the scalp, it can capture brain waves. This electrode is mainly used to perform EEG detection on the forehead area and the area behind the ear of the user. The forehead and the area behind the ear are often selected for EEG detection because these positions can relatively clearly capture the electrical signals of brain activities, especially when detecting brain waves related to emotions and stress. The first data refers to the EEG signal data collected from these EEG electrodes, which contain information about the electrical activities of the brain. An EMG sensor is a sensor used to measure muscle electrical activities. It is attached to the muscle area to detect muscle electrical signals. Here, the EMG sensor is specifically used to perform EMG detection on the front face area of the user (such as the forehead, around the eyes, eyebrows, etc.) because the muscle activities in these areas are closely related to emotional changes. The second data refers to the EMG signal data collected by the EMG sensor. EMG signals reflect facial expression changes, such as frowning, smiling, etc., and are often strongly correlated with emotional states (such as tension, pleasure, anger, etc.). A GSR sensor is used to monitor the electrical conductivity of the skin. Usually, it detects changes in the skin through tiny currents on the skin surface, especially the activities of sweat glands. Galvanic skin response is often related to the user's emotional responses (such as tension, fear, or excitement). Especially when the emotion is high or tense, the electrical conductivity of the skin increases. The third data refers to the GSR data collected by the GSR sensor. These data can reflect the changes in the user's physiological responses when facing emotional stress.

[0036] Among them, since bioelectrical signals are easily affected by noise (such as environmental interference, electrical equipment noise, etc.), it is necessary to denoise the data. Common denoising methods include filter technology, waveform analysis, etc. The purpose is to remove the irrelevant components in the signal and leave a clearer bioelectrical signal. To better analyze bioelectrical data, it is usually necessary to divide the data into several time periods, and each time period represents a cycle of emotional changes. By segmenting the data, the emotional state changes of the user at different time periods can be identified more accurately. Baseline correction is to standardize the signal to eliminate individual differences between different users or different time periods. For example, the skin conductance responses of different users may vary, and through baseline correction, it can be adjusted to a unified standard, making the measurement data at different time periods comparable. The data processed through the above steps (denoising, data segmentation, baseline correction) is the final bioelectrical data, which includes preprocessed electroencephalogram data, electromyogram data, and skin conductance response data. This dataset will be used for subsequent emotion recognition, analysis, and feedback.

[0037] S120. Calculate the relative energy corresponding to the preset band according to the electroencephalogram data.

[0038] Specifically, electroencephalogram data is collected from the scalp through electroencephalogram electrodes and represents the electrical signals of brain neuron activities. These signals can be decomposed into different frequency components, and different frequency components are related to different electroencephalogram activities (such as relaxation, concentration, sleep, etc.). The frequency range of brain electrical activities can be divided into multiple different bands (frequency bands). The preset bands include: delta wave: 0.5 - 4 Hz, related to the deep sleep state; theta wave: 4 - 8 Hz, related to relaxation, meditation, or light sleep; alpha wave: 8 - 13 Hz, related to relaxation, wakefulness, and when eyes are closed; beta wave: 13 - 30 Hz, related to high concentration, anxiety, excitement, etc.; gamma wave: 30 - 100 Hz, related to high-frequency cognitive processes (such as learning, memory processing).

[0039] Among them, when the server processes electroencephalogram data, it will pre-define one or more bands as the objects of analysis. Relative energy refers to the proportion of the energy of a specific band in the total energy. In electroencephalogram signal processing, energy refers to the magnitude of the spectral density. Specifically, the server will transfer the electroencephalogram signal from the time domain to the frequency domain through spectral analysis methods to obtain the energy distribution in different frequency ranges. Then, calculate the spectral density value within this band and compare it with the total spectral density of the entire electroencephalogram signal to obtain the relative energy of this band. This proportional value can reflect the importance or proportion of this band in the total electroencephalogram activity.

[0040] In a possible implementation, according to the electroencephalogram (EEG) data, the relative energy corresponding to a preset frequency band is calculated, which specifically includes: parsing the EEG data to obtain an EEG signal; determining the preset frequency band according to the frequency band corresponding to the EEG signal; converting the EEG signal into a spectrum through fast Fourier transform (FFT) according to the preset frequency band; and determining the ratio between the spectral density corresponding to the preset frequency band and the total spectral density according to the spectrum, so as to obtain the relative energy corresponding to the preset frequency band.

[0041] Specifically, the fast Fourier transform is a mathematical algorithm used to convert a time-domain signal into a frequency-domain signal. The EEG signal shows a continuously changing voltage waveform in the time domain, and the FFT can convert these waveforms into a spectrum (i.e., the amplitude or energy of each frequency component). In this process, the server uses the FFT to transform the EEG signal to obtain the energy distribution of the signal at different frequencies. These frequency components correspond to different frequency bands of brain activity. For example, energy information in frequency ranges such as 0 - 4 Hz, 4 - 8 Hz, 8 - 13 Hz, etc. may be obtained. The spectral density refers to the energy magnitude of the signal within each frequency range. Through FFT analysis, an energy distribution including all frequencies can be obtained. The spectral density corresponding to the preset frequency band refers to the energy of the signal within the selected frequency band range (such as the alpha wave, beta wave, etc.). For example, the energy calculated within the frequency band of 8 - 13 Hz. The total spectral density is the sum of the energies of all frequency bands, that is, the energy of the entire EEG signal. The ratio is calculated by comparing the spectral density of the preset frequency band with the total spectral density to obtain the proportion of the energy of this band in the entire EEG signal. This proportion is the relative energy. Finally, the system calculates the relative energy of the preset frequency band through the above steps, and this value represents the importance or proportion of this band in the overall EEG activity. This relative energy can be used to analyze the physiological or psychological state of the user. For example, a high relative energy of the alpha wave may indicate that the user is in a relaxed state, and a high relative energy of the beta wave may indicate that the user is in a tense or anxious state.

[0042] S130. Determine the time-frequency characteristics of the electromyogram (EMG) activity according to the EMG data.

[0043] Specifically, the time-frequency characteristics combine the time-domain and frequency-domain characteristics to comprehensively describe the characteristics of the EMG signal. It can reflect the time variation and frequency component variation of the signal. For example, the time-domain characteristics can reveal the intensity of muscle activity, while the frequency-domain characteristics can reveal the frequency characteristics of muscle activity. By combining the two, a more comprehensive description can be obtained. For example, assume that the time-domain characteristics of a segment of EMG signal are: the mean is 2, the variance is 1, and the root mean square (RMS) is 1.5, while the frequency-domain characteristics are: the frequency range is from 20 Hz to 50 Hz, and the center frequency is 30 Hz. After combining the time-domain and frequency-domain characteristics, the obtained time-frequency characteristics may be a comprehensive characteristic set including these two aspects.

[0044] In a possible implementation, according to the electromyography data, determine the time-frequency characteristics of the electromyographic activity, specifically including: according to the electromyography data, determine the average absolute value of the muscle activity intensity and the electromyographic peak value; based on the average absolute value of the muscle activity intensity and the electromyographic peak value, generate time-domain characteristics; according to the electromyography data, calculate the frequency range and the frequency center of the muscle activity through fast Fourier transform to generate frequency-domain characteristics; according to the time-domain characteristics and the frequency-domain characteristics, obtain the time-frequency characteristics.

[0045] Specifically, the average absolute value refers to the average value of the absolute values of the electromyographic signals, which can reflect the intensity of the muscle activity. First, take the absolute values of the electromyographic signals within a certain period of time, and then calculate the average value of these absolute values. The larger the average absolute value, the stronger the muscle activity. Suppose a set of electromyographic signal data is: [2, -3, 5, -4, 6]. First, take the absolute values: [2, 3, 5, 4, 6], and then calculate their average value: (2 + 3 + 5 + 4 + 6) / 5 = 4. Thus, the average absolute value = 4. The electromyographic peak value refers to the maximum value of the electromyographic signal within a certain period of time. The peak value usually reflects the highest intensity of the muscle activity. If the signal is: [2, -3, 5, -4, 6], then the peak value is 6.

[0046] The time-domain characteristics are the characteristics directly extracted from the original electromyographic signals and are completed by calculating the statistical quantities of the signals, such as the mean, variance, RMS, etc. The time-domain characteristics reflect the time variation of the signals. Through FFT analysis, the energy distribution of the signals at different frequencies is obtained. The frequency range refers to the interval where the main frequency components of the electromyographic signals are located. For example, a certain muscle activity may be mainly concentrated in the frequency range of 10 Hz to 50 Hz. The frequency center is the weighted average frequency of all frequencies in the frequency spectrum and is usually used to describe the main frequency components of the signal. For example, if the frequency spectrum distribution of the signal is relatively concentrated, its frequency center can be calculated to represent the concentration degree of the frequency spectrum. Suppose through FFT, the frequency spectrum of the signal has strong energy concentration between 20 Hz and 50 Hz, and the frequency center is 35 Hz. This means that the main frequency components of this muscle activity are concentrated around 35 Hz. The time-frequency characteristics combine the time-domain characteristics and the frequency-domain characteristics and comprehensively describe the characteristics of the electromyographic signals. The time-domain characteristics reveal the change trend of the signals in time, while the frequency-domain characteristics reflect the energy distribution of the signals at different frequencies. By combining the time-domain characteristics and the frequency-domain characteristics, the time-frequency characteristics can be obtained. These characteristics provide a comprehensive description of the muscle activity and can more accurately judge the state of the muscle, such as whether it is in a tense, relaxed state, etc.

[0047] S140. According to the galvanic skin response data, determine the electroresponse characteristics.

[0048] Specifically, in emotion monitoring, the characteristics of the skin conductance response can help determine whether the user is in an anxious or stressed state. For example, if the peak conductivity is high and the rate of change in conductivity is large, it may indicate that the user is in a highly tense state, and the system can provide suggestions for emotion regulation, such as taking deep breaths, resting, etc. In health management, the characteristics of the skin conductance response can be used to detect the user's physiological and psychological states, provide timely feedback, and suggest appropriate relaxation methods to avoid excessive emotional fluctuations.

[0049] In a possible implementation, based on the skin conductance response data, the electrical response characteristics are determined, specifically including: calculating the peak conductivity and the rate of change in conductivity from the skin conductance response data; and determining the electrical response characteristics based on the peak conductivity and the rate of change in conductivity.

[0050] Specifically, the peak conductivity refers to the maximum conductivity value in the skin conductance response signal, which is used to reflect the intensity of the user's emotional response at a specific moment. The collected skin conductance response data is a time series (such as the conductivity values recorded per second). This data is traversed to find the maximum value. Suppose the skin conductance response data of a certain user over a period of time is: [0.4, 0.6, 1.2, 1.5, 1.3, 1.0, 0.8] µS. The peak of this data is 1.5 µS, that is, the peak conductivity of the user during this time period is 1.5 µS. A higher peak usually indicates that the user is in a higher state of tension or emotional arousal. The rate of change in conductivity represents the rate of change of skin conductivity over time, reflecting the rapidity and intensity of the emotional response. First, calculate the change value (difference) in conductivity between adjacent time points from the skin conductance response data. Second, divide these differences by the time interval to obtain the rate of change. Finally, the maximum rate of change or the average rate of change can be taken as the characteristic.

[0051] For example, a user wears an AR glasses, and the AR glasses record the skin conductance response data through a sensor device. During the process of the user watching a horror movie, the recorded data shows that the conductivity starts to rise from 0.4 µS, reaches a peak of 1.5 µS, and the rate of change in conductivity also increases significantly. At this time, it shows that the user is in a highly tense emotional state. Based on these characteristics, the server determines that the user may need to relax and prompts the user to take a deep breath or pause watching by controlling the display of the AR glasses. In stress monitoring, these characteristics can be used to detect whether the user's psychological burden exceeds the normal range. In emotion recognition, when combined with data such as electroencephalogram and electromyogram, these characteristics can more accurately identify the user's emotion categories, such as relaxation, anxiety, excitement, thereby further improving the accuracy of user emotion recognition.

[0052] S150. Use a CNN + LSTM fusion model to perform emotion recognition on the relative energy, time-frequency characteristics, and electrical response characteristics to obtain the recognition result.

[0053] Specifically, CNN is good at extracting local features of input data and is particularly suitable for processing time-frequency images, spectral data, etc. It can automatically extract spatial patterns (such as feature peaks, band distributions). LSTM is suitable for processing sequence data and can capture temporal dependencies, such as the time dynamics of emotional changes. By fusing the two, the spatial features (local patterns) and temporal features (dynamic dependencies) of the data can be learned simultaneously. The result output by the CNN+LSTM model is the probability distribution of emotion categories.

[0054] Therefore, by fusing electroencephalogram data, electromyogram data, and galvanic skin response data, this method can capture the emotional changes of users from multiple dimensions. Compared with emotion recognition using single-modal data, this multi-modal fusion strategy greatly improves the comprehensiveness and accuracy of the description of emotional states and reduces the risk of misjudgment caused by the absence or distortion of a single signal. By adopting the CNN+LSTM fusion model, the spatial feature extraction ability of CNN and the temporal dynamic modeling ability of LSTM are combined, enabling the model to efficiently extract static features and capture the dynamic characteristics of emotions changing over time, thus significantly improving the accuracy and robustness of emotion recognition. Integrating sensor devices into the AR glasses worn by users realizes the miniaturization and portability of the emotion recognition server. Combined with the real-time data acquisition and processing capabilities, this solution can quickly respond to the dynamic changes of users' emotions, providing the possibility for real-time emotion monitoring and feedback. The combination of multi-modal features effectively addresses the impacts of environmental noise and individual differences. Therefore, it is convenient to improve the accuracy of emotion recognition.

[0055] In a possible implementation, the CNN+LSTM fusion model is used to perform emotion recognition on relative energy, time-frequency features, and electro-response features to obtain the recognition result. Specifically, it includes: through the CNN+LSTM fusion model, reducing the dimension, filtering, and flattening the time-frequency features to obtain the first feature vector; through the CNN+LSTM fusion model, judging the dependency relationships of relative energy and electro-response features to determine short-term and long-term dependency relationships; generating the second feature vector according to the short-term and long-term dependency relationships; fusing the first feature vector and the second feature vector and integrating them into high-dimensional features through a fully connected layer; mapping the high-dimensional features to the probability distribution of emotion categories through the Softmax activation function; and generating the recognition result according to the probability distribution of emotion categories.

[0056] Specifically, by processing the time-frequency features, they are made suitable for input into the model. Dimensionality reduction is used to remove redundant information, such as removing frequency ranges with low correlation. Filtering removes noise through filters to improve data quality. Feature flattening unfolds matrices or high-dimensional features into one-dimensional vectors for convenient subsequent processing. The LSTM is used to capture the dynamic changes of relative energy and electro-response features in the time series. Short-term dependencies capture the instantaneous fluctuations of features. For example, a rapid increase in conductivity within the past few seconds reflects nervousness. Long-term dependencies analyze the overall trends of features. For example, continuous high-frequency EEG activity may indicate persistent anxiety.

[0057] For example, short-term dependency: within the past 5 seconds, the changes in relative energy are [0.3, 0.35, 0.4, 0.45, 0.5][0.3, 0.35, 0.4, 0.45, 0.5][0.3, 0.35, 0.4, 0.45, 0.5], indicating that the user's mood gradually changes from calm to tense. Long-term dependency: within the past 1 minute, the conductivity gradually increases from 1.01.01.0 microsiemens to 1.81.81.8 microsiemens, indicating that the user may be in a state of long-term anxiety. The server fuses the extracted feature vectors (the first feature vector and the second feature vector) to form a unified high-dimensional feature. The first feature vector is obtained by processing time-frequency features through a CNN. The second feature vector is obtained by capturing the dependencies of relative energy and electro-response features through an LSTM. The two are concatenated and further integrated into a high-dimensional feature through a fully connected layer for convenient classification.

[0058] For example, the first feature vector: [0.1, 0.2, 0.3][0.1, 0.2, 0.3][0.1, 0.2, 0.3].

[0059] The second feature vector: [0.4, 0.5, 0.6][0.4, 0.5, 0.6][0.4, 0.5, 0.6].

[0060] After fusion: [0.1, 0.2, 0.3, 0.4, 0.5, 0.6][0.1, 0.2, 0.3, 0.4, 0.5, 0.6][0.1, 0.2, 0.3, 0.4, 0.5, 0.6].

[0061] High-dimensional feature: [0.15, 0.45, 0.55, 0.65][0.15, 0.45, 0.55, 0.65][0.15, 0.45, 0.55, 0.65].

[0062] Among them, according to the high-dimensional features, the probability distribution of emotion categories is generated through the Softmax activation function. The Softmax function converts the high-dimensional features into normalized probabilities. For example: the input features = [0.5, 1.2, 0.8], and the Softmax output is: the probability distribution = [0.2, 0.5, 0.3]. Select the category with the highest probability as the recognition result. Suppose the classified emotion categories are {relaxed, anxious, tense}, and the Softmax output probability distribution = [0.1, 0.7, 0.2]. The final recognition result is: anxious.

[0063] In a possible implementation manner, referring to Figure 2 , Figure 2 is another process schematic diagram of a bioelectric emotion recognition method based on an AR glasses provided by an embodiment of the present application. It includes steps S210 to step S220, and the above steps are as follows: S210. If it is determined that the recognition result of the user at the current moment indicates a negative tense emotion, control the AR glasses to display a deep breathing suggestion to prompt the user to take a deep breath to relieve the negative tense emotion; S220. If it is determined that the recognition result of the user after a preset time period indicates a continuous tense emotion, control the AR glasses to display a suggestion to pause watching and rest.

[0064] Specifically, bioelectric data is obtained through a sensor, combined with a CNN+LSTM fusion model, to judge the user's emotional state in real time, ensuring the accuracy and timeliness of the feedback. The server provides hierarchical feedback according to the duration and intensity of the emotion. Immediate feedback: When there is short-term tense emotion (such as detected at the current moment), help the user quickly adjust through a simple prompt (deep breathing). Reinforced feedback: When there is continuous tense emotion, provide stronger intervention measures (rest suggestion). The AR glasses display information through a virtual interface, making the prompt intuitive and clear, and at the same time not interrupting the user's line of sight and tasks.

[0065] Therefore, by displaying a deep breathing suggestion in real time when detecting the user's negative tense emotion, this solution can quickly intervene and help the user self-regulate when the emotion just starts to fluctuate, thereby avoiding the further deterioration of the emotion. This immediate feedback mechanism improves the user experience and the humanization level of the server. If it is detected that the user's tense emotion lasts for a period of time, the server further provides a suggestion of "pause watching and rest" to help the user withdraw from the scenario that may cause emotional fatigue. This long-term guidance is of great significance to the user's health management, especially in high-intensity work or learning scenarios, and helps prevent stress accumulation. This solution provides different suggestions in stages according to the dynamic changes of the user's emotional state, reflecting the personalized adaptation to the user's needs. The hierarchical intervention in different emotional states improves the pertinence and effectiveness of the suggestions, and helps the user better adjust the emotion in a specific scenario.

[0066] The present application also provides a bioelectric emotion recognition device based on an AR glasses. Refer to Figure 3 , Figure 3 which is a schematic module diagram of a bioelectric emotion recognition device based on an AR glasses provided by an embodiment of the present application. The device is a server, and the server includes an acquisition module 31 and a processing module 32. Among them, the acquisition module 31 acquires bioelectric data of a user sent by a sensor device. The sensor device is located on the AR glasses worn by the user, and the bioelectric data includes electroencephalogram (EEG) data, electromyogram (EMG) data, and skin conductance response (SCR) data. The processing module 32 calculates the relative energy corresponding to a preset frequency band according to the EEG data. The processing module 32 determines the time-frequency characteristics of the EMG activity according to the EMG data. The processing module 32 determines the electroresponse characteristics according to the SCR data. The processing module 32 uses a CNN+LSTM fusion model to perform emotion recognition on the relative energy, time-frequency characteristics, and electroresponse characteristics to obtain a recognition result.

[0067] In a possible implementation manner, the acquisition module 31 acquires bioelectric data of a user sent by a sensor device, which specifically includes: the acquisition module 31 receives first data of a user sent by an EEG electrode, and the EEG electrode is used for performing EEG detection on the forehead area and the area behind the ear of the user; the acquisition module 31 receives second data of a user sent by an EMG sensor, and the EMG sensor is used for performing EMG detection on the front face area of the user; the acquisition module 31 receives third data of a user sent by an SCR sensor, and the SCR sensor is used for performing SCR detection on the side face area of the user; the processing module 32 performs denoising, data segmentation, and baseline correction on the first data, the second data, and the third data to obtain bioelectric data.

[0068] In a possible implementation manner, the processing module 32 calculates the relative energy corresponding to a preset frequency band according to the EEG data, which specifically includes: the processing module 32 analyzes the EEG data to obtain an EEG signal; the processing module 32 determines a preset frequency band according to the frequency band corresponding to the EEG signal; the processing module 32 converts the EEG signal into a spectrum by fast Fourier transform according to the preset frequency band; the processing module 32 determines the ratio between the spectrum density corresponding to the preset frequency band and the total spectrum density according to the spectrum to obtain the relative energy corresponding to the preset frequency band.

[0069] In a possible implementation manner, the processing module 32 determines the time-frequency characteristics of the EMG activity according to the EMG data, which specifically includes: the processing module 32 determines the average absolute value of the muscle activity intensity and the EMG peak value according to the EMG data; the processing module 32 generates time-domain characteristics based on the average absolute value of the muscle activity intensity and the EMG peak value; the processing module 32 calculates the frequency range and the frequency center of the muscle activity by fast Fourier transform according to the EMG data to generate frequency-domain characteristics; the processing module 32 obtains the time-frequency characteristics according to the time-domain characteristics and the frequency-domain characteristics.

[0070] In a possible implementation, the processing module 32 determines the electrophysiological response characteristics according to the galvanic skin response data, specifically including: the processing module 32 calculates the peak conductivity and the rate of change of conductivity according to the galvanic skin response data; the processing module 32 determines the electrophysiological response characteristics based on the peak conductivity and the rate of change of conductivity.

[0071] In a possible implementation, the processing module 32 uses a CNN+LSTM fusion model to perform emotion recognition on the relative energy, time-frequency characteristics, and electrophysiological response characteristics to obtain an identification result, specifically including: the processing module 32 reduces the dimension, filters, and flattens the time-frequency characteristics through the CNN+LSTM fusion model to obtain a first feature vector; the processing module 32 determines the short-term and long-term dependencies by judging the dependency relationship between the relative energy and the electrophysiological response characteristics through the CNN+LSTM fusion model; the processing module 32 generates a second feature vector according to the short-term and long-term dependencies; the processing module 32 fuses the first feature vector and the second feature vector and integrates them into high-dimensional features through a fully connected layer; the processing module 32 maps the high-dimensional features to the probability distribution of emotion categories through a Softmax activation function; the processing module 32 generates an identification result according to the probability distribution of emotion categories.

[0072] In a possible implementation, if the processing module 32 determines that the emotion indicated by the identification result of the user at the current moment is negative tension, it controls the AR glasses to display a suggestion to take a deep breath to prompt the user to take a deep breath to relieve the negative tension; if the processing module 32 determines that the emotion indicated by the identification result of the user after a preset time period is continuous tension, it controls the AR glasses to display a suggestion to pause watching and take a rest.

[0073] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be repeated here.

[0074] This application also provides an electronic device. Refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.

[0075] Among them, the communication bus 42 is used to realize the connection and communication between these components.

[0076] Among them, the user interface 43 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 43 may also include a standard wired interface and a wireless interface.

[0077] Among them, the network interface 44 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface).

[0078] Among them, the processor 41 may include one or more processing cores. The processor 41 uses various interfaces and lines to connect all parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 45, and by calling the data stored in the memory 45, it executes various functions of the server and processes data. Optionally, the processor 41 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 41 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 41 and may be implemented separately through a single chip.

[0079] Among them, the memory 45 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 45 includes a non-transitory computer-readable storage medium. The memory 45 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 45 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 45 may also be at least one storage device located far from the aforementioned processor 41. As Figure 4 shown, in the memory 45 as a computer storage medium, it may include an operating system, a network communication module, a user interface module, and an application program of a bioelectric emotion recognition method based on AR glasses.

[0080] In Figure 4 the electronic device shown, the user interface 43 is mainly used to provide an input interface for the user to obtain the data input by the user; while the processor 41 can be used to call the application program of a bioelectric emotion recognition method based on AR glasses stored in the memory 45. When executed by one or more processors, the electronic device executes the methods in one or more of the above embodiments.

[0081] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0082] This application also provides a computer-readable storage medium, and the computer-readable storage medium stores instructions. When executed by one or more processors, the electronic device executes the methods in one or more of the above embodiments.

[0083] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0084] In several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in electrical or other forms.

[0085] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0087] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0088] The above are only exemplary embodiments of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the specification and the disclosure of the practical truth. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A bioelectric emotion recognition method based on AR glasses, characterized in that: The method comprises: Acquire bioelectric data for a user sent by a sensor device, where the sensor device is located on the AR glasses worn by the user, and the bioelectric data includes electroencephalogram data, electromyography data, and galvanic skin response data; Calculating the relative energy corresponding to the preset band according to the EEG data; Determining the time-frequency characteristics of electromyographic activity according to the electromyographic data; determining electrical response characteristics according to the electrical skin response data; The CNN+LSTM fusion model is used to perform emotion recognition on the relative energy, the time-frequency characteristics and the electrical response characteristics to obtain a recognition result.

2. The bioelectric emotion recognition method based on AR glasses according to claim 1 is characterized in that: The step of obtaining the bioelectric data of the user sent by the sensor device specifically includes: receiving first data for the user sent by an EEG electrode, wherein the EEG electrode is used to perform EEG detection on a forehead area and an area behind an ear of the user; receiving second data for the user sent by an electromyographic sensor, wherein the electromyographic sensor is used to perform electromyographic detection on a front facial area of ​​the user; receiving third data for the user sent by a skin galvanic response sensor, wherein the skin galvanic response sensor is used to detect skin galvanic response of a side facial area of ​​the user; The first data, the second data and the third data are subjected to denoising, data segmentation and baseline correction to obtain the bioelectric data.

3. The bioelectric emotion recognition method based on AR glasses according to claim 1 is characterized in that: The calculating, according to the EEG data, the relative energy corresponding to the preset band specifically includes: Analyze and obtain an EEG signal according to the EEG data; Determining the preset band according to the frequency band corresponding to the EEG signal; According to the preset band, converting the EEG signal into a frequency spectrum by fast Fourier transform; According to the frequency spectrum, a ratio between the frequency spectrum density corresponding to the preset band and the total frequency spectrum density is determined to obtain a relative energy corresponding to the preset band.

4. The bioelectric emotion recognition method based on AR glasses according to claim 1 is characterized in that: Determining the time-frequency characteristics of electromyographic activity according to the electromyographic data specifically includes: Determine the average absolute value of muscle activity intensity and the peak value of myoelectricity according to the myoelectricity data; Generate a time domain feature based on the average absolute value of the muscle activity intensity and the electromyographic peak value; According to the electromyographic data, the frequency range and frequency center of the muscle activity are calculated by fast Fourier transform to generate frequency domain features; The time-frequency feature is obtained according to the time-domain feature and the frequency-domain feature.

5. The bioelectric emotion recognition method based on AR glasses according to claim 1 is characterized in that: Determining the electrical response characteristics according to the skin electrical response data specifically includes: Calculating the peak conductivity and the conductivity change rate according to the skin electrical response data; The electrical response characteristic is determined based on the peak conductivity and the rate of change of conductivity.

6. The bioelectric emotion recognition method based on AR glasses according to claim 1, characterized in that: The CNN+LSTM fusion model is used to perform emotion recognition on the relative energy, the time-frequency features, and the electrical response features to obtain a recognition result, which specifically includes: Through the CNN+LSTM fusion model, the time-frequency features are reduced in dimension, filtered, and flattened to obtain a first feature vector; By using the CNN+LSTM fusion model, the dependency relationship between the relative energy and the electrical reaction characteristics is judged to determine the short-term dependency relationship and the long-term dependency relationship; Generate a second feature vector according to the short-term dependency and the long-term dependency; The first feature vector and the second feature vector are fused and integrated into a high-dimensional feature through a fully connected layer; Mapping the high-dimensional features into probability distribution of emotion categories through a Softmax activation function; A recognition result is generated according to the probability distribution of the emotion category.

7. The bioelectric emotion recognition method based on AR glasses according to claim 1, characterized in that: The method further comprises: If it is determined that the recognition result of the user at the current moment indicates that the emotion is a negative nervous emotion, controlling the AR glasses to display a suggestion for deep breathing to prompt the user to take a deep breath to relieve the negative nervous emotion; If it is determined that the recognition result of the user after the preset time period indicates that the emotion is continuously tense, the AR glasses are controlled to display a suggestion to pause watching and take a rest.

8. A bioelectric emotion recognition device based on AR glasses, characterized in that: The bioelectric emotion recognition device comprises an acquisition module (31) and a processing module (32), wherein: The acquisition module (31) is used to acquire bioelectric data of a user sent by a sensor device, wherein the sensor device is located on the AR glasses worn by the user, and the bioelectric data includes electroencephalogram data, electromyography data, and galvanic skin response data; The processing module (32) is used to calculate the relative energy corresponding to the preset band according to the EEG data; The processing module (32) is further used to determine the time-frequency characteristics of the electromyographic activity based on the electromyographic data; The processing module (32) is further used to determine electrical response characteristics based on the skin electrical response data; The processing module (32) is also used to use a CNN+LSTM fusion model to perform emotion recognition on the relative energy, the time-frequency characteristics, and the electrical response characteristics to obtain a recognition result.

9. An electronic device, characterized in that: The electronic device comprises a processor (41), a memory (45), a user interface (43) and a network interface (44), wherein the memory (45) is used to store instructions, the user interface (43) and the network interface (44) are both used to communicate with other devices, and the processor (41) is used to execute the instructions stored in the memory (45) so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.

Citation Information

Patent Citations

  • Wearable multi-mode emotional state monitoring device

    CN112120716A

  • Feature fusion method based on multilevel electroencephalogram signal expression

    CN113128459A

  • Electroencephalogram data processing method, device and system, computer equipment and storage medium

    CN114847975A

  • Multi-modal emotion recognition method and system based on regularization fusion

    CN118656745A

  • Fault monitoring device and method for aviation servo actuation test bench

    CN119099872A