Emotion recognition method and device based on multi-modal signals, edge computing device and storage medium

By performing quality assessment and baseline state assessment on multimodal signals and dynamically adjusting the fusion weights, the problem of insufficient multimodal signal fusion strategies in existing technologies is solved, achieving high-precision and high-reliability emotion recognition, which is suitable for edge computing devices.

CN122096798APending Publication Date: 2026-05-29KINGFAR INTERNATIONAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KINGFAR INTERNATIONAL INC
Filing Date
2025-12-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing emotion recognition methods lack dynamic fusion strategies for multimodal signals, resulting in insufficient accuracy and reliability of the system in complex tasks or practical applications, inability to effectively identify modal validity, and complex system structure that is difficult to deploy.

Method used

By evaluating the signal quality and baseline state of multimodal signals, dynamically adjusting the fusion weights, and using edge computing devices for real-time signal processing, dynamic fusion of multimodal signals and emotion recognition can be achieved.

Benefits of technology

It improves the accuracy and reliability of emotion recognition, enhances the robustness and adaptability of the system, simplifies the system structure, and meets the application requirements of real-time performance and low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122096798A_ABST
    Figure CN122096798A_ABST
Patent Text Reader

Abstract

The present disclosure provides an emotion recognition method and device based on multi-modal signals, an edge computing device and a storage medium, and relates to the technical field of human factors engineering. The method comprises: acquiring multi-modal signals of a user, the multi-modal signals comprising a plurality of single-modal signals of different modalities collected for the user; respectively performing signal quality evaluation on the plurality of single-modal signals to obtain an overall quality score of each single-modal signal; respectively performing baseline state evaluation on the plurality of single-modal signals to obtain an overall baseline similarity score of each single-modal signal; respectively performing emotion recognition based on the plurality of single-modal signals to obtain an emotion recognition result of each single-modal signal; calculating a fusion weight of each single-modal signal according to the overall quality score and the overall baseline similarity score; and performing weighted fusion based on the emotion recognition result of each single-modal signal and the fusion weight of each single-modal signal to obtain a global emotion recognition result of the user. The present disclosure helps to improve the accuracy and reliability of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of human factors engineering technology, and in particular to an emotion recognition method, apparatus, edge computing device and storage medium based on multimodal signals. Background Technology

[0002] Current emotion recognition methods mainly fall into two categories. One relies on a single modality (such as EEG signals, facial expressions, or heart rate) for state recognition, lacking a collaborative fusion mechanism for multimodal signals. The other is based on multimodal signals for emotion recognition, but it mostly uses a fixed fusion structure or feature splicing for multimodal signals, and cannot dynamically adjust the fusion strategy according to the quality of the modal signals, task type, or usage scenario. This makes the system prone to failure in complex tasks or practical applications, affecting the accuracy and reliability of emotion recognition. Summary of the Invention

[0003] One of the technical problems this disclosure aims to solve is to improve the accuracy and reliability of emotion recognition.

[0004] To address the aforementioned technical problems, this disclosure provides an emotion recognition method based on multimodal signals, comprising: Acquire the user's multimodal signals, which include multiple single-modal signals with different modes collected from the user; Signal quality is evaluated for multiple single-mode signals, and an overall quality score is obtained for each single-mode signal. Baseline state assessments were performed on multiple single-mode signals to obtain an overall baseline similarity score for each single-mode signal; Emotion recognition is performed based on multiple single-modal signals, and the emotion recognition result for each single-modal signal is obtained; The fusion weights for each single-mode signal are calculated based on the overall quality score and the overall baseline similarity score. The user's global emotion recognition result is obtained by weighted fusion based on the emotion recognition result of each single modal signal and the fusion weight of each single modal signal.

[0005] In some embodiments, the plurality of single-mode signals include a first-mode signal. Before performing baseline state assessments on multiple single-mode signals separately, the method further includes: The overall quality score of the first modal signal is compared with a first threshold to determine the effectiveness of the first modal signal; When the overall quality score of the first modal signal is greater than or equal to the first threshold, the step of baseline state assessment of the first modal signal is performed.

[0006] In some embodiments, the plurality of single-mode signals include a second-mode signal. Before performing emotion recognition based on multiple single-modal signals, the method also includes: The overall baseline similarity score of the second modal signal is compared with a second threshold to determine whether the second modal signal is a signal in a non-baseline state. When the overall baseline similarity score of the second modality signal is less than the second threshold, the step of emotion recognition based on the second modality signal is performed.

[0007] In some embodiments, signal quality assessment is performed on multiple single-mode signals to obtain an overall quality score for each single-mode signal, including: Extract the first feature corresponding to each single-mode signal, where the first feature is related to the signal quality; The first feature is mapped to a confidence score according to the preset confidence score mapping rule; Generate an overall quality score for each single-mode signal based on the confidence score; and / or Baseline state assessments were performed on multiple single-mode signals to obtain an overall baseline similarity score for each single-mode signal, including: Multiple second features are extracted for each single-modal signal, where the second features are related to the emotional state; Obtain the baseline template corresponding to each single-mode signal; Baseline matching is performed based on multiple second features and baseline templates to obtain baseline similarity scores corresponding to multiple second features; The baseline similarity scores corresponding to multiple second features are weighted and calculated to obtain the overall baseline similarity score for each single-mode signal.

[0008] In some embodiments, emotion recognition is performed based on multiple single-modal signals to obtain the emotion recognition result for each single-modal signal, including: The feature weights corresponding to multiple second features are determined based on the baseline similarity scores corresponding to multiple second features. A weighted feature vector is generated based on multiple second features and their corresponding feature weights. The weighted feature vectors are input into the recognition model corresponding to each single-modal signal to perform emotion recognition, and the emotion recognition result corresponding to each single-modal signal is obtained.

[0009] In some embodiments, the method further includes: Within the judgment period, the emotion recognition result of the target modal signal is compared with the emotion recognition results of other modal signals one by one, and the similarity of the recognition results is calculated. The target modal signal is any one of multiple single modal signals. If the similarity of the identification results is less than the third threshold, an inconsistency alarm is triggered, and the fusion weight of the target modal signal is reset to zero.

[0010] In some embodiments, when the single-modal signal is an electroencephalogram (EEG) signal, the first feature includes at least one of electrode impedance, amplitude over-limit ratio, and spectral energy concentration. When the single-mode signal is a near-infrared signal, the first characteristic includes at least one of signal stability, heartbeat frequency band energy proportion, and dual-wavelength consistency; When the single-mode signal is a PPG signal, the first feature includes at least one of the following: average heart rate, electrode detachment, amplitude abnormality, and NN interval fluctuation. When the single-modal signal is an eye-tracking signal, the first feature includes the eye-tracking coordinate loss rate; When the single-mode signal is an image signal, the first feature includes at least one of average brightness, contrast and sharpness; When the single-modal signal is a speech signal, the first feature includes at least one of the following: the effective speech segment ratio and the signal-to-noise ratio. When the single-modal signal is an electrical skin signal, the first feature includes amplitude validity.

[0011] This disclosure also provides an emotion recognition device based on multimodal signals, comprising: The acquisition module is used to acquire the user's multimodal signals, which include multiple single-mode signals with different modes collected from the user. The signal quality assessment module is used to assess the signal quality of multiple single-mode signals separately and obtain the overall quality score of each single-mode signal. The baseline state assessment module is used to assess the baseline state of multiple single-mode signals separately and obtain the overall baseline similarity score for each single-mode signal. The emotion recognition module is used to perform emotion recognition based on multiple single-modal signals, and obtain the emotion recognition result for each single-modal signal; The calculation module is used to calculate the fusion weight of each single-mode signal based on the overall quality score and the overall baseline similarity score; The recognition result fusion module is used to perform weighted fusion based on the emotion recognition result of each single modal signal and the fusion weight of each single modal signal to obtain the user's global emotion recognition result.

[0012] This disclosure also provides an edge computing device, a processor, and a memory. The memory stores a computer program, and the processor runs the computer program to implement the emotion recognition method based on multimodal signals as shown in the above embodiments.

[0013] This disclosure also provides a computer-readable storage medium storing a computer program that, when run on a computer, implements the emotion recognition method based on multimodal signals as described in the above embodiments.

[0014] Through the above technical solution, the emotion recognition method, apparatus, edge computing device, and storage medium based on multimodal signals provided in this disclosure perform quality analysis and baseline state evaluation on each single-modal signal to obtain an overall baseline similarity score for the overall quality score of each single-modal signal. The fusion weight of each single-modal signal is then calculated by combining the overall baseline similarity score with the overall quality score. When the signal quality deteriorates, or the deviation between the current state and the baseline state is small, the fusion weights of the signals are dynamically adjusted to dynamically regulate the contribution of each modality signal to the final emotion recognition result, thereby improving the accuracy and reliability of emotion recognition and enhancing the emotion recognition system's ability to recognize subtle differences in emotion categories. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the emotion recognition method based on multimodal signals provided in this embodiment of the disclosure; Figure 2 This is a schematic diagram of the signal quality assessment process provided in the embodiments of this disclosure; Figure 3 This is a schematic diagram of the baseline status assessment process provided in the embodiments of this disclosure; Figure 4 This is a schematic diagram of the structure of the edge computing device provided in the embodiments of this disclosure; Figure 5 This is a schematic diagram of the structure of the emotion recognition device based on multimodal signals provided in the embodiments of this disclosure. Detailed Implementation

[0017] The embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings and examples. The detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of this disclosure by way of example, but should not be used to limit the scope of this disclosure. This disclosure can be implemented in many different forms and is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

[0018] These embodiments are provided to make the disclosure thorough and complete, and to fully express the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specifically stated, the relative arrangement of components and steps, material composition, numerical expressions, and values ​​set forth in these embodiments should be interpreted as exemplary only and not as limiting.

[0019] All terms used in this disclosure have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein.

[0020] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.

[0021] The physiological data involved in this disclosure are various measurable data signals in the human body, including but not limited to electrocardiogram (ECG) signals, skin temperature (SKT) signals, photoplethysmogram (PPG) signals, electrodermal activity (EDA) signals, heart rate (HR) signals, electromyogram (EMG) signals, electroencephalogram (EEG) signals, and peripheral capillary oxygen saturation (SpO2) signals.

[0022] Current multimodal emotion recognition technologies suffer from at least the following technical problems: 1. Weak modality fusion capability and poor dynamic adaptability: Most multimodal signals adopt fixed fusion structures or feature splicing, and cannot dynamically adjust the fusion strategy according to the quality of modal signals, task type or usage scenario. This makes the system prone to failure in complex tasks or practical applications, affecting the accuracy and reliability of emotion recognition.

[0023] 2. Inability to control modal validity: The lack of a mechanism for evaluating real-time signal quality makes it impossible to identify and manage modal failure states, such as signal loss, artifact interference, and facial occlusion. This results in low-quality modalities continuously interfering with the recognition results during fusion, severely impacting the robustness and accuracy of the emotion recognition system.

[0024] 3. Lack of state triggering mechanism and baseline reference judgment: After receiving multimodal data, the system directly enters the classification stage, ignoring the difference analysis between the data and the user's baseline state. The lack of a mechanism to judge "whether a significant state change has occurred" leads to misjudgments even when the system is in a resting or inactive state, affecting interpretability and application credibility.

[0025] 4. Complex system structure or poor real-time performance, making it difficult to deploy: Some studies have attempted to introduce deep learning or large models for state recognition, but these methods are usually highly dependent on computing resources, have large inference latency, lack lightweight structure and engineering portability, and are difficult to meet the application requirements of wearable devices, edge platforms and other applications that require real-time performance and low power consumption.

[0026] 5. Weak system closed-loop capability and difficulty in implementation: Existing solutions mostly focus on the recognition model itself and lack a complete system process from modal data acquisition, signal quality assessment, state judgment, result output to feedback optimization. This makes it difficult to achieve automated operation and product-level deployment, which limits its promotion value in real-world scenarios.

[0027] Based on this, this disclosure provides an emotion recognition method based on multimodal signals, which helps to improve the accuracy and reliability of emotion recognition.

[0028] Figure 1 An emotion recognition method based on multimodal signals is provided in the embodiments of this disclosure, such as... Figure 1 As shown, the specific steps include: Step S11: Acquire the user's multimodal signals, which include multiple single-mode signals with different modes collected from the user.

[0029] Multimodal signals are physiological and behavioral signals related to emotions. Multimodal signals include multiple monomodal signals, which include, but are not limited to, electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), electrodermal conductance (EDA), eye tracking, speech audio, and facial images.

[0030] Specifically, signal acquisition devices enable real-time synchronous acquisition of various emotion-related physiological and behavioral signals, such as electroencephalogram (EEG), near-infrared spectroscopy (fNIRS), electrical skin activity (EDA), heart rate variability (HRV), speech audio, facial images, and eye-tracking signals. The signal acquisition devices can interface with the main system of the edge computing device via wired (USB, serial port) or wireless (Bluetooth, Wi-Fi) methods, supporting dynamic configuration and plug-and-play mode access.

[0031] In some embodiments, after acquiring multimodal signals, it is also necessary to perform time alignment on the multimodal signals to form a time-consistent input data stream.

[0032] Specifically, by integrating a high-precision Network Time Protocol (NTP) synchronization mechanism and a clock drift compensation algorithm based on a local crystal oscillator within the system, real-time time calibration and correction can be performed on each modal signal acquisition channel. This ensures that each modal signal is extracted and fused in a unified time domain during subsequent feature extraction and fusion judgment, which is helpful for the synchronous perception and recognition of multimodal emotional states.

[0033] Step S12: Perform signal quality assessment on multiple single-mode signals respectively to obtain the overall quality score of each single-mode signal.

[0034] Specifically, this disclosure employs a unified sliding window mechanism (e.g., a unified window width of 10 seconds and a step size of 2 seconds) to perform real-time signal quality assessment on multiple single-modal signals. For each single-modal signal, the system presets key feature indicators and converts each indicator into a confidence score in the 0–1 range. Finally, a weighted fusion is used to generate an overall quality score for each single-modal signal within the current time window. This overall quality score is used to dynamically adjust the participation weight and validity judgment of each modality in the emotion recognition result fusion process, improving the robustness and adaptability of the system. Furthermore, this disclosure adopts a pluggable design, supporting the loading of new modalities and their corresponding quality assessment strategies through configuration files, demonstrating good flexibility and platform adaptability.

[0035] In some embodiments, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the signal quality assessment process provided in the embodiments of this disclosure.

[0036] Step S12 above specifically includes the following steps: Step S121: Extract the first feature corresponding to each single-mode signal, wherein the first feature is related to the signal quality.

[0037] Step S122: Map the first feature to a confidence score according to the preset confidence score mapping rule.

[0038] Step S123: Generate an overall quality score for each single-mode signal based on the confidence score.

[0039] Specifically, the system pre-defines corresponding key feature indicators (i.e., the first feature) for each single-mode signal. After acquiring multiple single-mode signals, when performing real-time quality analysis on the multi-mode signals in the current time window, the system extracts the first feature corresponding to each single-mode signal and then maps the first feature to a confidence score according to the pre-defined confidence score mapping rules.

[0040] It should be noted that if the number of first features of a certain modal signal is 1, then the confidence score of the first feature is used as the overall quality score of the modal signal; if the number of first features of a certain modal signal is multiple, then the confidence scores of the multiple first features are weighted and fused (either by equal weight or by preset weight) and used as the overall quality score of the modal signal.

[0041] If a certain modal signal (such as EEG signal or near-infrared signal) includes signals from multiple channels, the quality of each channel signal is evaluated separately to obtain the overall quality score of each channel. The overall quality scores of each channel are then weighted and fused (either by equal weight or by preset weight) to obtain the overall quality score of the modal signal.

[0042] In some embodiments, when the single-modal signal is an EEG signal, the first feature includes at least one of electrode impedance, amplitude over-limit ratio, and spectral energy concentration; when the single-modal signal is a near-infrared signal, the first feature includes at least one of signal stability, heart rate band energy ratio, and dual-wavelength consistency; when the single-modal signal is a PPG signal, the first feature includes at least one of average heart rate, electrode detachment, amplitude abnormality, and normal-to-normal interval (NN interval) fluctuation; when the single-modal signal is an eye-tracking signal, the first feature includes eye-tracking coordinate loss rate; when the single-modal signal is an image signal, the first feature includes at least one of average brightness, contrast, and sharpness; when the single-modal signal is a speech signal, the first feature includes at least one of effective speech segment ratio and signal-to-noise ratio; when the single-modal signal is an electrodermal signal, the first feature includes amplitude validity.

[0043] Specifically, the quality assessment of multiple single-mode signals includes: 1. Quality assessment of electroencephalogram (EEG) signals Within the current time window, the EEG signal is assessed channel by channel in real time, and a standardized confidence score (0–1) is output to dynamically control the participation weight and validity judgment of this modality in the emotion recognition process. The overall quality score of the EEG signal integrates the confidence scores of the following three primary features, specifically including: (1) Confidence score of electrode impedance The impedance value of each channel is acquired in real time. The preset confidence score mapping rule for the electrode impedance is as follows: electrode impedance < 10kΩ, score is 1; 10 kΩ ≤ electrode impedance < 20 kΩ, score is 0.75; 20 kΩ ≤ electrode impedance < 50 kΩ, score is 0.5; 50 kΩ ≤ electrode impedance ≤ 100 kΩ, score is 0.25; electrode impedance > 100kΩ or measurement failure, score is 0.

[0044] (2) Confidence score of the percentage of amplitude exceeding the limit The proportion of sample points exceeding ±100μV in the current window is statistically analyzed to determine whether the signal has high-amplitude artifact interference. The preset confidence score mapping rule corresponding to the amplitude exceeding the limit is as follows: amplitude exceeding the limit < 5%, score is 1; 5% ≤ amplitude exceeding the limit ≤ 20%, score is 0.5; amplitude exceeding the limit > 20%, score is 0.

[0045] (3) Confidence score of spectral energy concentration Short-time Fourier transform (STFT) was performed on the EEG signal to calculate the total energy in the frequency bands 0.5–45 Hz and 0–70 Hz, and the energy percentage was determined. The calculation formula is as follows:

[0046] Where R represents the spectral energy concentration. This represents the total energy in the frequency band 0.5–45Hz. This represents the total energy in the frequency band 0–70Hz.

[0047] The preset confidence score mapping rule corresponding to the spectral energy concentration is as follows: if R≥0.5, it means that the main energy of the signal is concentrated in the physiologically effective frequency band, and the score is 1; otherwise, it is 0.

[0048] It should be noted that the frequency band 0.5–45Hz is the effective frequency band for EEG signals, and the frequency band 0–70Hz represents the frequency band with a sampling rate of 0–1 / 2. In other embodiments, the corresponding frequency band can be selected according to the actual application. This disclosure only uses the frequency bands 0.5–45Hz and 0–70Hz as examples.

[0049] 2. Quality assessment of near-infrared (fNIRS) signals Within the current time window, the raw light intensity signal of fNIRS is evaluated channel by channel, and a standardized confidence score (0–1) is output to dynamically adjust the participation weight and effectiveness judgment of this modality in the emotion recognition process. The overall quality score of the near-infrared signal integrates the confidence scores of the following three primary features, specifically including: (1) Confidence score of signal stability Calculate the coefficient of variation CV (CV = standard deviation / mean) of the original optical intensity signal for each channel within the current window. The coefficient of variation CV is used to measure the signal stability. The preset confidence score mapping rule corresponding to the signal stability is as follows: If CV ≤ 15, the score is 1; 15 < CV ≤ 30, the score is 0.5; CV > 30, the score is 0.

[0050] (2)Confidence score of the energy proportion in the heartbeat frequency band Extract the energy index of the heartbeat frequency band after converting the original optical intensity signal into the frequency domain. Calculate the ratio between the energy of the target frequency band (i.e., the heartbeat frequency band, such as 0.8–1.5 Hz) and the total energy (such as the energy of the 0.5–2 Hz frequency band), which reflects whether the channel has successfully captured the physiological rhythm signal.

[0051] The preset confidence score mapping rule corresponding to the energy proportion in the heartbeat frequency band is as follows: The ratio ≥ 1.5, the score is 1; the ratio < 0.8, the score is 0; 1.4 ≤ ratio < 1.5, the score is 0.9; 1.3 ≤ ratio < 1.4, the score is 0.8; 1.2 ≤ ratio < 1.3, the score is 0.7; 1.1 ≤ ratio < 1.2, the score is 0.6; 1.0 ≤ ratio < 1.1, the score is 0.5; 0.9 ≤ ratio < 1.0, the score is 0.4; 0.8 ≤ ratio < 0.9, the score is 0.3; (3)Confidence score of the dual - wavelength consistency Perform cross - correlation analysis on the filtered red light and near - infrared light in the heartbeat frequency band to obtain the cross - correlation coefficient and extract the peak power.

[0052] The preset confidence score mapping rule corresponding to the dual - wavelength consistency is as follows: If the cross - correlation coefficient ≥ 0.7 and the peak power ≥ 0.2, the score is 1; if 0.5 ≤ cross - correlation coefficient < 0.7 and 0.1 ≤ peak power < 0.2, the score is 0.5; otherwise, the score is 0.

[0053] 3. Quality assessment of the PPG signal Perform real - time analysis on the PPG signal within the current time window, extract the first feature, evaluate the signal rhythm and integrity, and output the standardized confidence score (0–1) for dynamically controlling the participation weight and effectiveness judgment of this modality in the emotion recognition process. The overall quality score of the PPG signal combines the confidence scores of the following 4 first features, specifically including: (1)Confidence score of the average heart rate Based on the detection of the pulse wave within the window, extract the NNI sequence and calculate the average heart rate (HR). The preset confidence score mapping rule corresponding to the average heart rate is as follows: If 40 bpm ≤ HR ≤ 180 bpm, the score is 1; otherwise, the score is 0.

[0054] (2) Confidence score of electrode detachment Static segments are identified by calculating the first-order difference of the PPG signal. If the fluctuation amplitude (mean absolute value of the first-order difference of the PPG signal) within the current window is less than 0.01, it is judged as electrode detachment or severe distortion. The preset confidence score mapping rule for electrode detachment is as follows: normal fluctuation (mean absolute value of the first-order difference of the PPG signal is greater than or equal to 0.01), the score is 1; signs of detachment appear (mean absolute value of the first-order difference of the PPG signal is less than 0.01), the score is 0.

[0055] (3) Confidence score for anomalous amplitude This function checks whether the signal amplitude within the current window exceeds the 5%–95% amplitude percentile range of the entire segment (of collected data). The preset confidence score mapping rule for amplitude anomalies is as follows: if the signal amplitude within the current window is within the range, the score is 1; if the signal amplitude within the current window exceeds the range, the score is 0.

[0056] (4) Confidence score of NN interval fluctuation The proportion of outliers in the current window's NN interval sequence is analyzed to assess the NN interval volatility. Outliers are defined as points that deviate from the mean of the entire segment (of collected data) by more than three times the standard deviation of the entire segment. The preset confidence score mapping rule for NN interval volatility is as follows: outlier proportion ≤ 10%, score 1; 10% < outlier proportion ≤ 30%, score 0.5; outlier proportion > 30%, score 0.

[0057] 4. Quality assessment of eye-tracking signals The quality of eye-tracking data is evaluated in real time within the current time window. A first feature is extracted from the eye-tracking signals, and a standardized confidence score (0–1) is generated to dynamically control the participation weight and validity judgment of this modality in the emotion recognition process. The first feature is the eye-tracking coordinate loss rate, which can be used to assess the data completeness of the eye-tracking signals. The overall quality score of the eye-tracking signals is the confidence score of the eye-tracking coordinate loss rate.

[0058] Specifically, the percentage of sample points with missing (e.g., null or invalid) gaze (eye movement) trajectories within the current window is calculated (i.e., the eye movement coordinate loss rate). The preset confidence score mapping rule for the eye movement coordinate loss rate is as follows: eye movement coordinate loss rate < 10%, score is 1; 10% ≤ eye movement coordinate loss rate ≤ 30%, score is 0.5; eye movement coordinate loss rate > 30%, score is 0.

[0059] 5. Image signal quality assessment Based on continuous frame image data within the current time window, the average brightness, contrast, and sharpness are evaluated frame by frame. The overall quality score of each frame is calculated, and then a weighted average is performed to obtain the overall quality score of the image signal within the current time window. The overall quality score of each frame incorporates the confidence scores of the following three primary features: (1) Confidence score of average brightness The average grayscale value of the image (reflecting average brightness) is statistically analyzed. The preset confidence score mapping rule for the average brightness is: if the average grayscale value is in the range of 80-180, the score is 1; otherwise, the score is 0.

[0060] (2) Confidence score of contrast The standard deviation of image brightness (reflecting contrast) is calculated. The preset confidence score mapping rule for contrast is: if the standard deviation is >30, the score is 1; otherwise, the score is 0.

[0061] (3) Confidence score of sharpness The Laplacian operator is used to calculate the image sharpness index (variance). The preset confidence score mapping rule for sharpness is: if the variance is >100, the score is 1; otherwise, the score is 0.

[0062] 6. Quality assessment of speech signals Real-time quality assessment of the speech signal is performed within the current time window, extracting two feature indicators: speech validity and signal-to-noise ratio (SNR) to evaluate the usability and clarity of the speech data. A confidence score in the 0–1 range is generated. The overall quality score of the speech signal integrates the confidence scores of the following two primary features: (1) Confidence score of the percentage of effective speech segments The proportion of short-term energy exceeding a threshold energy (reflecting the proportion of effective speech segments) is calculated. The threshold energy can be represented by the silence baseline plus 10 dB of speech segment energy. The preset confidence score mapping rule corresponding to the proportion of effective speech segments is as follows: proportion > 90%, score is 1; 60% ≤ proportion ≤ 90%, score is 0.5; proportion < 60%, score is 0.

[0063] (2) Confidence score of signal-to-noise ratio Estimate the signal-to-noise ratio (SNR) of the overall speech signal in the current window. The preset confidence score mapping rule for the SNR is as follows: SNR > 20 dB, score is 1; 10 dB ≤ SNR ≤ 20 dB, score is 0.5; SNR < 10 dB, score is 0.

[0064] 7. Quality assessment of electrodermal signals Real-time quality assessment of electrodermal signals is performed within the current time window. The first feature is extracted from the amplitude validity dimension, and a standardized confidence score (0–1) is generated to dynamically control the participation weight and validity judgment of this modality in the emotion recognition process. The overall quality score of the electrodermal signal is the confidence score of amplitude validity.

[0065] Specifically, it is determined whether the average skin conductance level (SCL) of the current window is within a reasonable physiological range (e.g., 0.01–60 μS). The preset confidence score mapping rule for amplitude validity is: if it is within the range, the score is 1; otherwise, the score is 0.

[0066] It should be noted that if a single-mode signal has multiple first features, the overall quality score of the modal signal can be generated based on the confidence scores of one or more first features. For example, the confidence scores of average brightness, contrast, or sharpness can be used as the overall quality score of the image signal, or the overall quality score of the image signal can be generated by weighted fusion of the confidence scores of average brightness and contrast, or by weighted fusion of the confidence scores of average brightness, contrast, and sharpness.

[0067] In this disclosure, the signal quality scores for all modalities are based on a unified sliding window mechanism, for example, a window width of 10 seconds and a step size of 2 seconds, to ensure consistent alignment across modal time scales. Multiple index scores within each modality are normalized to the 0–1 interval and fused using an equal-weighted averaging method to generate a modal-level confidence score (i.e., the overall quality score of each single-modal signal) for each time window.

[0068] By assessing the quality of real-time signals, the system continuously monitors the stability and integrity of data from various modalities (such as EEG, speech, facial images, and electrodermal conductance). It extracts metrics such as artifact rate, signal-to-noise ratio (SNR), image sharpness, and packet loss rate to generate confidence scores. When the quality of a modality's signal deteriorates (e.g., facial occlusion or speech mute), the system automatically reduces the fusion weight of that modality or dynamically masks it, ensuring that low-quality signals do not mislead emotion judgments and improving recognition robustness and contextual adaptability.

[0069] To ensure the stability and data reliability of subsequent recognition processes, this disclosure further proposes setting unified gating conditions: only when the signal quality score of a certain modality signal within the current window is ≥0.6, is the modality data considered valid and included in the subsequent baseline state assessment and emotion recognition processes. Furthermore, if a modality signal continuously exhibits low quality (e.g., a multi-window confidence score consistently below 0.6), the system will automatically downweight or temporarily remove it to ensure the robustness and effectiveness of the multimodal fusion structure.

[0070] In some embodiments of this disclosure, the plurality of single-mode signals include a first mode signal, wherein the first mode signal can be any one of the plurality of single-mode signals. Before performing baseline state assessment on the plurality of single-mode signals respectively, the method further includes: comparing the overall quality score of the first mode signal with a first threshold to determine the validity of the first mode signal; when the overall quality score of the first mode signal is greater than or equal to the first threshold, performing the step of baseline state assessment on the first mode signal.

[0071] For example, if the overall quality score of the speech signal in the current window is calculated to be 0.7, which is greater than the first threshold (such as 0.6), it proves that the speech signal in the current window is a valid signal and can be used for subsequent baseline state assessment and emotion recognition processes.

[0072] In some embodiments, for single-channel signals, the validity of the modality signal can be determined based on the overall quality score of the modality signal. For multi-channel signals (such as EEG signals and near-infrared signals), the validity of the modality signal can be determined not only based on the overall quality score of the modality signal but also by the validity of each channel. For example, for an EEG signal with multiple channels, the overall quality score of each channel is calculated. If the overall quality score of a certain channel is greater than or equal to a first threshold, it indicates that the channel is usable and the data in that channel is valid. The ratio of usable channels to total channels is calculated. If the ratio is greater than or equal to a preset threshold (such as 0.5 or 0.6), it indicates that the modality signal is valid.

[0073] This disclosure improves recognition robustness and situational adaptability by setting a gating threshold for signal quality scoring to ensure that low-quality signals do not mislead emotional judgments.

[0074] Step S13: Baseline state assessment is performed on multiple single-mode signals to obtain the overall baseline similarity score for each single-mode signal.

[0075] The system incorporates multimodal resting state or neutral emotional state templates. Based on large-scale sample data, it constructs emotional baseline references by grouping individuals by age, gender, etc., for comparing the deviation (or similarity) of the current input state with the individual's or group's normal state. Multiple unimodal signals are compared with their corresponding baseline templates (resting state or neutral emotional state templates) to obtain an overall baseline similarity score for each unimodal signal.

[0076] In some embodiments, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the baseline status assessment process provided in the embodiments of this disclosure.

[0077] Step S13 above specifically includes the following steps: Step S131: Extract multiple second features corresponding to each single-modal signal, wherein the second features are related to the emotional state.

[0078] Step S132: Obtain the baseline template corresponding to each single-mode signal.

[0079] Step S133: Baseline matching is performed based on multiple second features and baseline templates to obtain baseline similarity scores corresponding to multiple second features.

[0080] Step S134: The baseline similarity scores corresponding to multiple second features are weighted and calculated to obtain the overall baseline similarity score of each single-mode signal.

[0081] Specifically, the second feature extracted from each single-modal signal is identical to the features contained in the corresponding baseline template; both are core indicators highly sensitive to emotional states, meaning the second feature is correlated with emotional state. Feature selection is based on existing emotional neuroscience literature, behavioral physiology research findings, and cross-validation results from multiple publicly available emotion datasets (such as DEAP, AMIGOS, MAHNOB-HCI, SEED, etc.) to ensure its efficiency and reliability in emotion recognition. Each modality baseline template contains the following features: Electroencephalography (EEG): power characteristics of 5 classic frequency bands (δ, θ, α, β, γ), among which α and β waves are often used to distinguish between tension and calmness, and θ and δ are significantly correlated with arousal level; Near-infrared spectroscopy (fNIRS): Changes in oxygenated (HbO) and deoxygenated (HbR) hemoglobin concentrations in the frontal lobe region, used to reflect the level of cortical activation during emotional evoked emotions; HRV (Heart Rate Variability): average heart rate, SDNN (Standard Deviation of NN Intervals), PNN50 and LF / HF (Low Frequency / High Frequency) ratio, representing changes in autonomic nervous system regulation and emotional tension. Electrodermal conductance (EDA): The average skin conductance level (SCL) and the skin conductance response (SCR) count within a time window (e.g., 10 s) are used to determine the intensity of sympathetic nerve arousal. Eye movement: Mean pupil diameter, often used to assess the intensity of emotional activation and changes in attentional focus; Speech features: emotion-related speech features such as formants, fundamental frequency (F0), spectral centroid and energy changes; Images (facial expressions): Key facial action units (AUs) (such as AU4, AU6, AU12, AU17, etc.) and their intensity values ​​are used to characterize different types of emotional expressions.

[0082] Whenever a new multimodal signal is acquired within a 10-second sliding window, for any single-modal signal, the system first extracts multiple second features of the single-modal feature, obtains the baseline template corresponding to the single-modal feature, then compares the multiple second features and the baseline template dimension by dimension to obtain the baseline similarity score corresponding to the multiple second features, and finally performs a weighted calculation on the baseline similarity scores corresponding to the multiple second features to obtain the overall baseline similarity score of the single-modal signal.

[0083] In some embodiments, baseline matching is performed based on multiple second features and a baseline template to obtain baseline similarity scores corresponding to multiple second features, including: using the confidence score mapping rule under the Gaussian distribution assumption to calculate the baseline similarity scores corresponding to multiple second features based on the mean and standard deviation of the baseline template.

[0084] The matching method between multiple second features and the baseline template is based on the Gaussian distribution assumption. The baseline similarity score corresponding to multiple second features is calculated using the confidence score mapping rule under the Gaussian distribution assumption. Specifically, for any input second feature value x, its deviation from the baseline mean μ and standard deviation σ is mapped to a confidence score S (i.e., the baseline similarity score, ranging from 0 to 1), calculated as follows:

[0085] The confidence score S represents the degree of consistency between the current state of the second feature and the baseline state.

[0086] For each single-modal signal, the baseline similarity scores corresponding to multiple second features are calculated separately, and then the scores are averaged with equal weight or preset weight to obtain the modal-level baseline similarity score, thus obtaining the overall baseline similarity score for each single-modal signal.

[0087] To prevent the system from making misjudgments in a resting or inactive state, which would affect interpretability and application reliability, a gating threshold is set for the baseline state evaluation of each modality signal before emotion recognition. Only when the overall baseline similarity score of a modality signal is less than the gating threshold is the modality signal determined to be a non-baseline state signal, and the subsequent emotion recognition process for that modality signal is activated.

[0088] In some embodiments, the plurality of single-modal signals includes a second modal signal, wherein the second modal signal is any one of the plurality of single-modal signals. Before performing emotion recognition based on the plurality of single-modal signals respectively, the method further includes: comparing the overall baseline similarity score of the second modal signal with a second threshold to determine whether the second modal signal is a signal in a non-baseline state; when the overall baseline similarity score of the second modal signal is less than the second threshold, performing the step of emotion recognition based on the second modal signal.

[0089] Specifically, the system sets a second threshold (e.g., 0.3). If the overall baseline similarity score of the second modality signal is greater than or equal to the second threshold, it indicates that the current state has not significantly deviated from the resting interval, and the system does not enter the emotion recognition stage to avoid invalid calculations. If the overall baseline similarity score of the second modality signal is less than the second threshold, the second modality signal is determined to be a signal of a non-baseline state, and the step of emotion recognition based on the second modality signal can be executed.

[0090] This disclosure calculates whether there is a significant emotional deviation based on the multi-feature matching degree between the current window data of each modal signal and the corresponding baseline template. The subsequent emotion recognition process is initiated only when the emotion state is determined to be "non-baseline", that is, when there is significant emotional activation (such as anger, tension, pleasure, etc.). This avoids misjudging emotional fluctuations in resting, emotionally stable, or unexcited conditions, improves the system's response accuracy and scene interpretability, and enhances the system's stability and practical interpretability. It is suitable for application scenarios that require continuous emotion monitoring, such as teaching, driving, and customer service.

[0091] In some embodiments, a global gating threshold (e.g., 0.3) can be set to average the overall baseline similarity scores of all modal signals to obtain the global baseline matching confidence. When the global baseline matching confidence is less than the global gating threshold, the subsequent emotion recognition process is triggered.

[0092] This disclosure, through the fusion and consistency verification of multiple modal information, ensures that the system makes a high-confidence judgment of "non-baseline state" only when multiple modalities collaboratively and mutually corroborate each other to point to the same state change. This can effectively filter single-modal noise and greatly improve the accuracy and reliability of the judgment.

[0093] In some embodiments, this disclosure provides an adaptive baseline update mechanism. When the user's accumulated resting data exceeds 20% of the current baseline template data, the system will automatically recalculate the baseline mean μ and standard deviation σ corresponding to each modality and update the current baseline template, thereby adapting to slow changes in individual states or scene migrations over the long term. Furthermore, the system can dynamically add new modalities, new features, and user tags through configuration files, exhibiting high scalability and practical deployment flexibility.

[0094] Step S14: Perform emotion recognition based on multiple single-modal signals to obtain the emotion recognition result for each single-modal signal.

[0095] The system independently constructs and trains an emotion recognition model for each monomodal signal (such as EEG, fNIRS, EDA, HRV, speech, image, eye tracking, etc.). Multiple monomodal signals are input into their respective emotion recognition models to obtain the emotion recognition result for each monomodal signal.

[0096] Furthermore, this disclosure embeds a baseline-driven attention layer within each modal signal emotion recognition model for dynamic weighting at the feature level, thereby improving the accuracy of emotion recognition.

[0097] Specifically, emotion recognition is performed based on multiple single-modal signals to obtain the emotion recognition result for each single-modal signal. This includes: determining the feature weights corresponding to multiple second features based on the baseline similarity scores corresponding to multiple second features; generating a weighted feature vector based on the multiple second features and the feature weights corresponding to the multiple second features; and inputting the weighted feature vector into the recognition model corresponding to each single-modal signal to perform emotion recognition, thereby obtaining the emotion recognition result for each single-modal signal.

[0098] The features of the emotion recognition model input corresponding to each modality signal are consistent with the features of the baseline template to achieve a direct correspondence with the baseline evaluation results, ensuring that the baseline matching probability can be directly used for attention modulation and dynamic weighting within the model.

[0099] In the above embodiments, for each single-modal signal, key physiological or behavioral indicators reflecting emotional state (i.e., second features) are extracted from the original signal. Examples include the power of each frequency band in EEG, HbO / HbR fluctuations in fNIRS, SCL and SCR in EDA, SDNN and LF / HF in HRV, fundamental frequency and energy changes in speech, AU intensity in facial images, and pupil diameter. Multiple second features are matched with a baseline template to calculate the baseline similarity score corresponding to the second feature (values ​​in the range of 0-1, with larger values ​​indicating closer proximity to the baseline). Based on the baseline similarity scores corresponding to multiple second features, an "anti-consistency weight" is calculated for each second feature. This is obtained by subtracting the baseline similarity score from 1, and this anti-consistency weight is the feature weight corresponding to the second feature. Intuitively, the more a second feature deviates from the baseline (the smaller the baseline similarity score), the greater its anti-consistency weight, and the more "attention" it receives from the system in emotion recognition. Conversely, if a second feature is very close to the baseline (the high baseline similarity score), its anti-consistency weight is smaller, and its influence is relatively suppressed. The input values ​​of multiple second features are multiplied by their corresponding feature weights to form a weighted feature vector. This weighted feature vector is then input into the recognition model (e.g., the Softmax classification layer in the model) corresponding to each single-modal signal to perform emotion recognition, resulting in the emotion recognition result for each single-modal signal. The emotion recognition result is, for example, a three-class emotion probability vector [Ppos, Pneu, Pneg], where Ppos, Pneu, and Pneg represent the predicted probabilities of positive, neutral, and negative emotions, respectively.

[0100] In some embodiments, the method provided in this disclosure further includes: if the feature weight is less than a fourth threshold, then setting the feature weight to a preset value.

[0101] To improve robustness, a "weight lower limit" can be set. When the "anti-consistency weight" is too small and the feature weight is close to zero, that is, when the feature weight is less than the fourth threshold (close to 0), it is forced to be raised to a pre-set minimum value (i.e., a preset value, such as a constant between 0 and 1). This can prevent the feature weight from being excessively suppressed and losing potentially effective information.

[0102] Step S15: Calculate the fusion weight of each single-mode signal based on the overall quality score and the overall baseline similarity score.

[0103] After completing emotion recognition for each single-modal signal, the system introduces a global weighting mechanism at the modality layer based on baseline deviation and signal quality scores to integrate multimodal recognition results and generate the final emotion judgment. This mechanism dynamically evaluates the current validity and credibility of each modality, ensuring that the multi-source data adaptively reflects the differences in contributions from different modalities during the fusion process.

[0104] The system first receives the output results from the emotion recognition model corresponding to each single-modal signal. =[Ppos,Pneu,Pneg], and simultaneously read the overall quality score corresponding to each single-mode signal. Similarity score with overall baseline The overall quality score reflects the stability and integrity of the modal data within the current window, while the overall baseline similarity score measures whether the feature distribution of that modality deviates significantly from the resting state. Both scores jointly determine the weight calculation for that modality in the fusion process. The formula for calculating the fusion weight of each single-modal signal is as follows:

[0105] in, This represents the fusion weight, where M is the number of valid modes currently participating in the fusion. This represents the overall quality score corresponding to the k-th single-mode signal. This represents the overall baseline similarity score corresponding to the k-th single-modal signal. When the signal quality score of a certain modality is lower than 0.6, its output is automatically masked and does not participate in this round of fusion calculation. Step S16: Perform weighted fusion based on the emotion recognition result of each single-modal signal and the fusion weight of each single-modal signal to obtain the user's global emotion recognition result. All modal outputs that pass the validity judgment are weighted and summed to obtain the system's global sentiment prediction vector Z:

[0106] Final Emotion Category The category is determined by the maximum value among the three probabilities:

[0107] For example, if Z=[0.2,0.5,0.3], then the final emotion category is... It corresponds to the category with a probability of 0.5, namely neutral sentiment.

[0108] This disclosure adaptively adjusts the contribution of each modality and its internal features in the multimodal emotion recognition process, achieving dynamic fusion and output from signal quality control and baseline deviation assessment to the final three-emotion classification (positive, neutral, and negative). The system comprehensively considers the baseline differences at the modality level and the feature level. At the feature level, a baseline-driven attention layer is embedded within the emotion recognition model for dynamic weighting at the feature level. At the modality level, an adaptive weighting based on signal quality and baseline deviation probability is implemented, achieving a two-layer attention control mechanism, thereby maintaining high recognition accuracy and system robustness in complex environments. Through the modality-level fusion mechanism, multimodal dynamic balance is achieved in complex scenarios: when the signal quality of a certain modality decreases or its difference from the baseline is insufficient, its weight in the overall decision-making is automatically reduced; conversely, when a certain modality exhibits significant emotional activation signals, its weight increases accordingly, thereby improving the accuracy and robustness of emotion recognition.

[0109] The system simultaneously records the weight distribution and historical trends of each modality within the current time window for subsequent modality performance tracking and adaptive updates. For example, the system continuously monitors the consistency of output results between different modalities and adaptively adjusts the weights based on the dynamic distribution characteristics of cross-modal outputs. When a modality's judgment results are significantly inconsistent with those of other modalities within a certain operating period and show a continuous reverse trend, the system automatically determines that its reliability has decreased and modulates the weight of that modality to 0 during emotion recognition result fusion to prevent abnormal modalities from interfering with the overall emotion recognition results.

[0110] In some embodiments, the method provided in this disclosure further includes: within a determination period, comparing the emotion recognition result of the target modal signal with the emotion recognition results of other modal signals one by one, and calculating the similarity of the recognition results, wherein the target modal signal is any one of a plurality of single modal signals; if the similarity of the recognition results is less than a third threshold, an inconsistency alarm is triggered, and the fusion weight of the target modal signal is reset to zero.

[0111] The system employs a sliding consistency monitoring mechanism, which does not rely on a fixed time window but rather makes dynamic judgments based on real-time statistical characteristics. Specifically, the system uses a single real-time prediction output as the judgment period (each result is recorded once). For each output, the target modal signal is compared with the recognition results of each other modal signal: if the binary labels are the same (i.e., the recognition results are the same), it is recorded as 1; otherwise, it is recorded as 0. The average of these comparison scores is taken to obtain the consistency of this output (the value is between 0 and 1). The system archives 30 outputs as a batch: at the end of each batch, the average consistency of the 30 outputs in that batch is recorded as the average consistency rate of the target modal signal (i.e., the similarity of the recognition results, the value is between 0 and 1); simultaneously, the percentage of times the target modal signal's result is 0 compared with other modal signals in that batch is recorded as the opposite result ratio (the value is between 0 and 1). If the target modality signal is detected to have an average consistency rate lower than a threshold (i.e., the third threshold, such as 0.3) for three or more consecutive batches, an inconsistency alarm is triggered; the system multiplies the fusion weight of the target modality signal by a mask at the emotion recognition result fusion layer. m =0, temporarily setting its fusion weight upper limit to 0, so that the identification result of the target modal signal no longer participates in the fusion, and only data monitoring is retained. When its average consistency rate recovers to above the threshold (e.g., 0.6) and remains so for 3 consecutive batches, the system automatically removes the mask and restores its original weight upper limit.

[0112] Through the aforementioned consistency adaptive mechanism, the system achieves dynamic reliability assessment without fixed time window constraints, can automatically identify and isolate abnormal outputs that continuously deviate from other modalities, realizes real-time self-adjustment and automatic recovery of fusion weights, and significantly improves the robustness and long-term operational stability of the multimodal emotion recognition system.

[0113] In summary, the emotion recognition method based on multimodal signals provided in this disclosure has the following technical advantages: This disclosure introduces dynamic signal quality assessment, baseline difference gating, modal attention modulation, and adaptive consistency adjustment mechanisms into the field of multimodal emotion recognition. It constructs an intelligent emotion recognition system with a closed-loop operating logic of "acquisition-evaluation-recognition-adjustment-feedback," achieving adaptive optimization throughout the entire process from signal management to recognition decision-making. Compared with existing technologies, this disclosure achieves significant technological advancements in the following aspects: 1. A dynamic closed-loop structure for multimodal emotion recognition has been formed. By sharing confidence parameters, baseline deviation probability, and consistency statistics, the data layer and decision layer can interact in real time. During operation, the system can continuously monitor signal quality, determine emotion activation, fuse recognition results, and self-adjust weights, thereby constructing a multimodal recognition framework with high robustness, self-learning, and self-repair capabilities.

[0114] 2. A modal reliability management mechanism based on result consistency is proposed. This disclosure introduces a consistency judgment model between modal results in the fusion layer. By automatically calculating the matching rate between the output of each modality and the results of other modalities, it suppresses modalities that deviate or reverse their output over a long period of time. This achieves dynamic modality screening and weight zeroing strategy without the need for a fixed time window, which significantly improves the stability and anti-interference ability of the recognition system.

[0115] 3. An emotion activation gating mechanism based on individualized baselines was implemented.

[0116] The system can build resting or neutral emotion templates for different users and compare the deviation of the current state from the baseline features in real time input. The recognition process is triggered only when significant emotional activation is detected, which effectively reduces misjudgment in the resting state and enhances the interpretability and application reliability of the system.

[0117] 4. A multi-level adaptive adjustment strategy for fusion weights was established.

[0118] The system integrates signal quality scores and baseline deviation probabilities to achieve multi-layered weighted regulation from the feature layer, modality layer, to the decision layer. This mechanism can automatically adjust modality contribution under different emotional states and signal conditions, achieving adaptive optimal fusion of emotion recognition results.

[0119] 5. It has good lightweight and deployability.

[0120] The system adopts a modular and pluggable design, with high computing efficiency and low latency. It can be flexibly deployed on wearable devices, mobile terminals and edge computing platforms, meeting the requirements of real-time emotion monitoring and low power consumption operation, and has good engineering feasibility.

[0121] This disclosure also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods provided in the embodiments shown in this disclosure.

[0122] The following is combined with Figure 4 The exemplary edge computing devices provided in the embodiments of this disclosure are further described. Figure 4 A schematic diagram of the edge computing device 4000 is shown.

[0123] The aforementioned edge computing device 4000 may include: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the multi-process emotion recognition method based on multimodal signals provided in the embodiments of this disclosure by calling the program instructions.

[0124] Figure 4 A block diagram is shown of an exemplary edge computing device 4000 suitable for implementing embodiments of the present disclosure. Figure 4 The edge computing device 4000 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0125] like Figure 4 As shown, the edge computing device 4000 is presented in the form of a general-purpose computing device. The components of the edge computing device 4000 may include, but are not limited to: one or more processors 4010, memory 4020, communication bus 4040 connecting different system components (including memory 4020 and processor 4010), and communication interface 4030.

[0126] The 4040 communication bus represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0127] Edge computing devices 4000 typically include a variety of computer system-readable media. These media can be any available media that can be accessed by the edge computing device, including volatile and non-volatile media, and removable and non-removable media.

[0128] Memory 4020 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. Edge computing devices may further include other removable / non-removable, volatile / non-volatile computer system storage media. Although Figure 4As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the communication bus 4040 via one or more data media interfaces. The memory 4020 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0129] A program / utility having a set (at least one) of program modules may be stored in memory 4020. Such program modules include, but are not limited to, an operating system, one or more applications, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules typically perform the functions and / or methods described in the embodiments of this disclosure.

[0130] The edge computing device 4000 can also communicate with one or more external devices (e.g., keyboard, pointing device, display, etc.), one or more devices that enable users to interact with the edge computing device, and / or any device that enables the edge computing device to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). This communication can be performed through the communication interface 4030. Furthermore, the edge computing device 4000 can also communicate through a network adapter (… Figure 4 (Not shown) communicates with one or more networks (e.g., Local Area Network (LAN), Wide Area Network (WAN), and / or public networks, such as the Internet). The aforementioned network adapter can communicate with other modules of the edge computing device via the communication bus 4040. It should be understood that, although... Figure 4 As not shown, the edge computing device 4000 can be used with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Drives (RAID) systems, tape drives, and data backup storage systems.

[0131] The processor 4010 executes various functional applications and data processing by running programs stored in the memory 4020, such as implementing the methods provided in the embodiments of this disclosure.

[0132] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this disclosure are merely illustrative and do not constitute a structural limitation on the edge computing device 4000. In other embodiments of this disclosure, the edge computing device 4000 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0133] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0134] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0135] In the embodiments provided in this disclosure, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0136] This disclosure also provides an emotion recognition device based on multimodal signals, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of an emotion recognition device based on multimodal signals provided in this disclosure embodiment. The emotion recognition device 50 based on multimodal signals includes: The acquisition module 51 is used to acquire the user's multimodal signals, which include multiple single-mode signals with different modes collected from the user. The signal quality assessment module 52 is used to assess the signal quality of multiple single-mode signals respectively and obtain the overall quality score of each single-mode signal; The baseline state assessment module 53 is used to assess the baseline state of multiple single-mode signals respectively and obtain the overall baseline similarity score of each single-mode signal. The emotion recognition module 54 is used to perform emotion recognition based on multiple single-modal signals respectively, and obtain the emotion recognition result of each single-modal signal; The calculation module 55 is used to calculate the fusion weight of each single-mode signal based on the overall quality score and the overall baseline similarity score; The recognition result fusion module 56 is used to perform weighted fusion based on the emotion recognition result of each single modal signal and the fusion weight of each single modal signal to obtain the user's global emotion recognition result.

[0137] Figure 5 The emotion recognition device 50 based on multimodal signals provided in the illustrated embodiment can be used to execute the technical solution of the method embodiment shown in this application. Its implementation principle and technical effect can be further referred to the relevant description in the method embodiment.

[0138] The embodiments of this disclosure have now been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.

[0139] While specific embodiments of this disclosure have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments or equivalent substitutions can be made to some technical features without departing from the scope and spirit of this disclosure. In particular, as long as there is no structural conflict, the technical features mentioned in the various embodiments can be combined in any manner.

Claims

1. An emotion recognition method based on multimodal signals, characterized in that, include: Acquire the user's multimodal signals, which include multiple single-modal signals with different modes collected for the user; Signal quality is evaluated for each of the multiple single-mode signals to obtain an overall quality score for each single-mode signal; Baseline state assessments are performed on the multiple single-mode signals to obtain an overall baseline similarity score for each single-mode signal; Emotion recognition is performed based on the multiple single-modal signals respectively, and the emotion recognition result of each single-modal signal is obtained; The fusion weights for each single-modal signal are calculated based on the overall quality score and the overall baseline similarity score. The user's global emotion recognition result is obtained by weighted fusion based on the emotion recognition result of each single modal signal and the fusion weight of each single modal signal.

2. The method according to claim 1, characterized in that, The plurality of single-mode signals includes a first-mode signal. Before performing baseline state assessment on the plurality of single-mode signals respectively, the method further includes: The overall quality score of the first modal signal is compared with a first threshold to determine the validity of the first modal signal; When the overall quality score of the first modal signal is greater than or equal to the first threshold, the step of performing a baseline state assessment on the first modal signal is executed.

3. The method according to claim 1, characterized in that, The plurality of single-mode signals includes a second-mode signal. Before performing emotion recognition based on the plurality of single-modal signals respectively, the method further includes: The overall baseline similarity score of the second modal signal is compared with a second threshold to determine whether the second modal signal is a signal in a non-baseline state. When the overall baseline similarity score of the second modality signal is less than the second threshold, the step of emotion recognition based on the second modality signal is performed.

4. The method according to claim 1, characterized in that, The step of evaluating the signal quality of each of the multiple single-mode signals to obtain an overall quality score for each single-mode signal includes: Extract the first feature corresponding to each single-mode signal, wherein the first feature is related to signal quality; The first feature is mapped to a confidence score according to a preset confidence score mapping rule; Generate an overall quality score for each single-mode signal based on the confidence score; and / or The baseline state assessment of the multiple single-mode signals is performed respectively to obtain an overall baseline similarity score for each single-mode signal, including: Extract multiple second features corresponding to each single-modal signal, wherein the second features are related to the emotional state; Obtain the baseline template corresponding to each single-mode signal; Baseline matching is performed based on the plurality of second features and the baseline template to obtain the baseline similarity score corresponding to the plurality of second features; The baseline similarity scores corresponding to the multiple second features are weighted and calculated to obtain the overall baseline similarity score of each single-mode signal.

5. The method according to claim 4, characterized in that, The step of performing emotion recognition based on the multiple single-modal signals to obtain the emotion recognition result for each single-modal signal includes: The feature weights corresponding to the multiple second features are determined based on the baseline similarity scores corresponding to the multiple second features. A weighted feature vector is generated based on the plurality of second features and the feature weights corresponding to the plurality of second features; The weighted feature vector is input into the recognition model corresponding to each single-modal signal to perform emotion recognition, and the emotion recognition result corresponding to each single-modal signal is obtained.

6. The method according to claim 1, characterized in that, The method further includes: Within the judgment period, the emotion recognition result of the target modal signal is compared with the emotion recognition results of other modal signals one by one, and the similarity of the recognition results is calculated. The target modal signal is any one of the multiple single modal signals. If the similarity of the identification result is less than the third threshold, an inconsistency alarm is triggered, and the fusion weight of the target modal signal is reset to zero.

7. The method according to claim 4, characterized in that, When the single-modal signal is an electroencephalogram (EEG) signal, the first feature includes at least one of electrode impedance, amplitude over-limit ratio, and spectral energy concentration. When the single-mode signal is a near-infrared signal, the first feature includes at least one of signal stability, heartbeat frequency band energy ratio, and dual-wavelength consistency. When the single-mode signal is a PPG signal, the first feature includes at least one of the following: average heart rate, electrode detachment, amplitude abnormality, and NN interval fluctuation. When the single-modal signal is an eye-tracking signal, the first feature includes the eye-tracking coordinate loss rate; When the single-mode signal is an image signal, the first feature includes at least one of average brightness, contrast, and sharpness; When the single-modal signal is a speech signal, the first feature includes at least one of the following: effective speech segment ratio and signal-to-noise ratio. When the single-mode signal is an electrical skin signal, the first feature includes amplitude validity.

8. An emotion recognition device based on multimodal signals, characterized in that, include: The acquisition module is used to acquire the user's multimodal signals, which include multiple single-modal signals with different modes collected for the user; The signal quality assessment module is used to assess the signal quality of the multiple single-mode signals respectively, and obtain the overall quality score of each single-mode signal; The baseline state assessment module is used to assess the baseline state of the multiple single-mode signals respectively, and obtain the overall baseline similarity score of each single-mode signal; An emotion recognition module is used to perform emotion recognition based on the multiple single-modal signals respectively, and obtain the emotion recognition result for each single-modal signal; The calculation module is used to calculate the fusion weight of each single-mode signal based on the overall quality score and the overall baseline similarity score; The recognition result fusion module is used to perform weighted fusion based on the emotion recognition result of each single modal signal and the fusion weight of each single modal signal to obtain the user's global emotion recognition result.

9. An edge computing device, characterized in that, include: A processor and a memory, the memory being used to store a computer program; the processor being used to run the computer program to implement the emotion recognition method based on multimodal signals as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, implements the emotion recognition method based on multimodal signals as described in any one of claims 1-7.