Emotion speculation system and method based on multi-modal emotion monitoring
By constructing a multimodal emotion inference system, adjusting modal emotion probability values and fusing evidence bodies using consistency coefficients, the problem of multimodal conflict resolution was solved, achieving high accuracy in emotion monitoring and personalized decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANXI EDUCATIONAL AUDIO & VIDEO PUBLISHING HOUSE CO LTD
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing multimodal emotion monitoring systems cannot accurately capture users' true emotional state when dealing with conflicts between multimodalities, leading to critical misjudgments and an inability to provide personalized emotion intervention.
A multimodal sentiment inference system is constructed. Through modal data collection, processing, analysis and probability generation units, the modal sentiment probability value is adjusted using the consistency coefficient, and conflict resolution and evidence correction are realized through the evidence fusion module, and the modal evidence weight is dynamically adjusted.
It improves the robustness and accuracy of emotion monitoring, can adapt to different users and scenarios, avoids misleading results caused by traditional weighted averages, and enhances the practicality of decision-making.
Smart Images

Figure CN121867783A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emotion inference technology, specifically to an emotion inference system and method based on multimodal emotion monitoring. Background Technology
[0002] An emotion inference system and method based on multimodal emotion monitoring is disclosed. This system intelligently integrates facial, voice, and physiological signals to accurately infer emotional states. Patent application number 202510687499.6 discloses "an emotion monitoring model generation method, an emotion monitoring method, and an emotion intervention method. The emotion monitoring model generation method includes, in response to an emotion monitoring model generation request, acquiring first feature data and second feature data from multiple users. The first feature data includes physiological feature data and environmental feature data; determining initial modal feature data based on the physiological feature data of the first feature data; correcting the initial modal feature data based on the environmental feature data of the first feature data to determine corrected modal feature data, and determining an associated dataset based on the corrected modal feature data and the second feature data; training a pre-constructed initial emotion monitoring model based on the associated dataset from multiple users until the initial emotion monitoring model meets preset requirements, and determining a general emotion monitoring model."
[0003] The aforementioned existing technologies have solved the problems of low accuracy and individualization in emotion monitoring, as well as untimely emotion intervention and lack of personalized intervention plans. However, during the use of the system, due to the inability to see the fundamental conflicts between multimodalities, the simple weighted average can only output the average emotion, which is completely unable to capture the user's true state, which can easily lead to key misjudgments. Furthermore, it can only provide an isolated emotion label and cannot point out the contradictions between the sensor signals. Summary of the Invention
[0004] The purpose of this invention is to provide an emotion inference system and method based on multimodal emotion monitoring to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an emotion inference system based on multimodal emotion monitoring, comprising;
[0006] The modal data acquisition unit reads the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal, and uses timestamps to synchronize the data within the window.
[0007] The modal data processing unit processes the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and constructs them into feature tuples.
[0008] The modal data analysis unit inputs feature tuples into the value encoder, extracts feature vectors, and analyzes the preliminary emotion probability distribution of different modalities based on the feature vectors.
[0009] A probability generation unit calculates the consistency coefficient between different modalities and then adjusts the emotion probability value corresponding to each modality based on the consistency coefficient.
[0010] The emotion inference unit constructs a comprehensive evidence body based on the emotion probability values adjusted for each modality, and uses the comprehensive evidence body to determine the emotion inference result for the current time window.
[0011] Preferably, the modal data processing unit includes an image processing module and an audio processing module;
[0012] The image processing module obtains the current window. Multiple images from the original facial video frame sequence captured internally. ,in Indicates the first Zhang Image This represents the total number of images, analyzed using a deep neural network, and the output window... Facial image sequence after internal alignment ,in , Indicates the aligned first Zhang Image Indicates the serial number;
[0013] The audio processing module processes the raw audio time-domain signal. The noise-reduced audio signal is obtained. ,right Frame segmentation is performed using preset frame lengths and frame shifts. Divided into multiple overlapping short-time frames , for the A short time frame Processing is performed to obtain a windowed frame. ,in , This represents the Hamming window function. Indicates the total number of short-time frames. Indicates the first Each short time frame is traversed sequentially. Then, output all windowed frames. ,Will Stored to audio feature set middle.
[0014] Preferably, the modal data processing unit further includes an electrocardiogram (ECG) processing module, an electroencephalogram (EEG) processing module, and a feature construction module;
[0015] The ECG processing module uses a bandpass filter on the current window. Baseline drift and high-frequency noise were removed from the raw ECG signals acquired internally. The Pan-Tompkins algorithm was used to locate each heartbeat cycle in the filtered signal, and the window was analyzed based on the detected intervals of consecutive R peaks. Heart rate variability feature set ;
[0016] The EEG processing module uses bandpass filtering to filter the window. The raw EEG signals acquired internally are processed for each channel. The resulting signals are segmented, and a Fast Fourier Transform is performed on each segment. The average power of the standard EEG frequency band is calculated by integrating on the PSD. The power values of each channel and frequency band are combined and standardized to form a window. EEG feature vector ;
[0017] The feature construction module statistically analyzes the first Synchronization Time Window Facial image sequence within Audio feature set Heart rate variability feature set and EEG feature vectors They are then constructed into feature tuples.
[0018] Preferably, the modal data analysis unit includes a sequence analysis module, a speech analysis module, and a physiological analysis module;
[0019] The sequence analysis module inputs the aligned facial image sequence from the feature tuple into the visual encoder. After extracting the corresponding high-level visual feature vectors through forward propagation, it maps them to the emotion category space through a fully connected classification head to generate a preliminary visual modal emotion probability distribution.
[0020] The speech analysis module inputs the aligned audio feature set from the feature tuple into the speech encoder, extracts the high-level speech feature vector, and uses it through the classification head to obtain the preliminary speech modal emotion probability distribution.
[0021] The physiological analysis module concatenates the aligned heart rate variability feature set and EEG feature vector in the feature tuple to form a combined physiological feature vector, which is then input into the physiological encoder to obtain the high-level physiological feature vector. After that, the preliminary physiological modality emotion probability distribution is obtained through the classification head.
[0022] Preferably, the probability generation unit includes a matrix construction module, a probability adjustment module, a deviation coefficient calculation module, and a coefficient output module;
[0023] The matrix construction module receives high-level feature vectors corresponding to different modalities. Then, select Calculate its relationship with Consistency coefficient between ,in , , Represents a relational evaluation network. This represents the Sigmoid function, which calculates the consistency coefficient between each high-level eigenvector and other eigenvectors in turn, and then uses all the consistency coefficients to construct the corresponding matrix.
[0024] The probability adjustment module calculates the mean consistency value of each mode based on the consistency matrix, and selects the mode. Read the consistency mean of this mode. And the sentiment probability given by the initial classification head ,according to and Calculate the adjusted emotional probability value ,in , Indicates the scaling factor. Indicates the adjusted number The probability value of each emotion. Indicates the first The probability value of each emotion. Indicates the total number of emotion categories;
[0025] The deviation coefficient calculation module receives the adjusted emotional probability value. Using consistent mean Analyze the deviation coefficient within the current time window ,in , Indicates the penalty factor;
[0026] The coefficient output module repeats the operation until the adjusted emotional probability value and deviation coefficient corresponding to each modality are calculated.
[0027] Preferably, the emotion inference unit includes an evidence fusion module;
[0028] The evidence fusion module statistically analyzes the modality-adjusted emotional probability value and deviation coefficient for each mode, combines them into a set of evidence for the corresponding mode, and selects the mode. and modality evidence Then, traverse All propositions within, for any two propositions and ,like Then and The product value is added to the corresponding composite proposition probability value, if Then and The product value according to and The original probability values are assigned, and the assigned values are multiplied by the original probability values. The resulting value is used as the new value of the current proposition. After calculating all the proposition combinations, each proposition is summed and normalized to obtain the fused evidence body.
[0029] Preferably, the emotion inference unit further includes an evidence output module and a result determination module;
[0030] The evidence body output module merges the evidence body into a new evidence body, and then merges it with other modal evidence bodies. This process is repeated until all modal evidence bodies are merged to obtain a comprehensive evidence body.
[0031] The result determination module selects the emotion category with the highest probability value from the comprehensive evidence and uses it as the final emotion prediction result for the current time window.
[0032] The aforementioned emotion inference method based on multimodal emotion monitoring includes the following steps:
[0033] S1. Read the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal. Use the timestamp to synchronize each data within the window.
[0034] S2. Process the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and construct them into feature tuples;
[0035] S3. Input the feature tuple into the value encoder, extract the feature vector, analyze the preliminary emotion probability distribution of different modalities based on the feature vector, calculate the consistency coefficient between different modalities, and adjust the emotion probability value corresponding to each modality based on the consistency coefficient.
[0036] S4. Construct a comprehensive evidence body based on the modal-adjusted sentiment probability values, and use the comprehensive evidence body to determine the sentiment prediction result for the current time window.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] This invention constructs a unified computing framework capable of proactively managing, analyzing, and utilizing multimodal conflict information. It achieves conflict-driven dynamic reliability assessment and evidence correction. By quantifying intermodal consistency in real time through a relational evaluation network, the weights of each modality's evidence are dynamically adjusted accordingly. Consistent evidence is enhanced, while conflicting evidence is reasonably weakened. This mechanism enables the system to adapt to different users and scenarios, rather than relying on fixed weights, greatly improving robustness in real-world complex environments. Furthermore, by introducing a conflict resolution and fusion mechanism based on evidence theory, it achieves refined processing of contradictory information. This preserves the reasonable influence of high-confidence claims while avoiding the risk of misleading results from traditional weighted averages in modal conflict, making the final decision more realistic. Attached Figure Description
[0039] Figure 1 A schematic diagram of the overall system flow is provided for embodiments of the present invention. Detailed Implementation
[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] Please see Figure 1 The present invention provides a technical solution: an emotion inference system based on multimodal emotion monitoring, comprising;
[0042] The modal data acquisition unit reads the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal. The data within the synchronization window are synchronized using timestamps.
[0043] The modal data processing unit processes the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and constructs them into feature tuples.
[0044] The modal data analysis unit inputs feature tuples into the value encoder, extracts feature vectors, and analyzes the preliminary emotion probability distribution of different modalities based on the feature vectors.
[0045] The probability generation unit calculates the consistency coefficient between different modalities and then adjusts the emotion probability value corresponding to each modality based on the consistency coefficient.
[0046] The emotion inference unit constructs a comprehensive evidence body based on the emotion probability values adjusted for each modality, and uses the comprehensive evidence body to determine the emotion inference result for the current time window.
[0047] The modal data processing unit includes an image processing module and an audio processing module;
[0048] Image processing module obtains the current window Multiple images from the original facial video frame sequence captured internally. ,in Indicates the first Zhang Image This represents the total number of images, analyzed using a deep neural network, and the output window... Facial image sequence after internal alignment ,in , Indicates the aligned first Zhang Image This represents the sequence number. The specific process of analysis using a deep neural network is to select the first... Zhang Image Using deep neural networks to obtain The corresponding facial bounding box is used. If a face is detected, the set of facial key points is located, and the image is analyzed according to the coordinates of the eye center points in the key point set. Rotation correction is performed, and a similarity transformation is applied based on the corrected key points. The facial region is then cropped and scaled to a standard size to obtain an aligned facial image. Repeat the operation until All were selected;
[0049] The audio processing module processes the raw audio time-domain signal. The noise-reduced audio signal is obtained. ,right Frame segmentation is performed using preset frame lengths and frame shifts. Divided into multiple overlapping short-time frames , for the A short time frame Processing is performed to obtain a windowed frame. ,in , This represents the Hamming window function. Indicates the total number of short-time frames. Indicates the first Each short time frame is traversed sequentially. Then, output all windowed frames. ,Will Stored to audio feature set middle;
[0050] The modal data processing unit also includes an electrocardiogram (ECG) processing module, an electroencephalogram (EEG) processing module, and a feature construction module;
[0051] The ECG processing module uses bandpass filtering on the current window. Baseline drift and high-frequency noise were removed from the raw ECG signals acquired internally. The Pan-Tompkins algorithm was used to locate each heartbeat cycle in the filtered signal, and the window was analyzed based on the detected intervals of consecutive R peaks. Heart rate variability feature set ;
[0052] The EEG processing module uses bandpass filtering to filter the window. The raw EEG signals acquired internally are processed for each channel. The resulting signals are segmented, and a Fast Fourier Transform is performed on each segment. The average power of the standard EEG frequency band is calculated by integrating on the PSD. The power values of each channel and frequency band are combined and standardized to form a window. EEG feature vector ;
[0053] Feature building module statistics Synchronization Time Window Facial image sequence within Audio feature set Heart rate variability feature set and EEG feature vectors They are then constructed into feature tuples;
[0054] The modal data analysis unit includes a sequence analysis module, a speech analysis module, and a physiological analysis module;
[0055] The sequence analysis module inputs the aligned facial image sequence from the feature tuple into the visual encoder. After extracting the corresponding high-level visual feature vectors through forward propagation, it maps them to the emotion category space through a fully connected classification head to generate a preliminary visual modal emotion probability distribution.
[0056] The speech analysis module inputs the aligned audio feature set from the feature tuple into the speech encoder, extracts the high-level speech feature vector, and uses it through the classification head to obtain the preliminary speech modal emotion probability distribution.
[0057] The physiological analysis module concatenates the aligned heart rate variability feature set and EEG feature vector from the feature tuple to form a combined physiological feature vector, which is then input into the physiological encoder to obtain the high-level physiological feature vector. After that, the preliminary physiological modality emotion probability distribution is obtained through the classification head.
[0058] The probability generation unit includes a matrix construction module, a probability adjustment module, a deviation coefficient calculation module, and a coefficient output module;
[0059] The matrix construction module receives high-level feature vectors corresponding to different modalities. Then, select Calculate its relationship with Consistency coefficient between ,in , , Represents a relational evaluation network. This represents the Sigmoid function, which calculates the consistency coefficient between each high-level eigenvector and other eigenvectors in turn, and then uses all the consistency coefficients to construct the corresponding matrix.
[0060] The probability adjustment module calculates the mean consistency value for each mode based on the consistency matrix and selects the mode. Read the consistency mean of this mode. And the sentiment probability given by the initial classification head ,according to and Calculate the adjusted emotional probability value ,in , Indicates the scaling factor. Indicates the adjusted number The probability value of each emotion. Indicates the first The probability value of each emotion. Indicates the total number of emotion categories;
[0061] The deviation coefficient calculation module receives the adjusted emotion probability value. Using consistent mean Analyze the deviation coefficient within the current time window ,in , Indicates the penalty factor;
[0062] The coefficient output module repeats the operation until the adjusted emotional probability value and deviation coefficient corresponding to each modality are calculated;
[0063] The emotion inference unit includes an evidence fusion module;
[0064] The evidence fusion module calculates the modality-adjusted sentiment probability value and deviation coefficient for each modality, combines them into a set of evidence for the corresponding modality, and selects the modality. and modality evidence Then, traverse All propositions within, for any two propositions and ,like Then and The product value is added to the corresponding composite proposition probability value, if Then and The product value according to and The original probability values are assigned, and the assigned values are multiplied by the original probability values. The resulting value is used as the new value of the current proposition. After calculating all the combinations of propositions, each proposition is summed and normalized to obtain the fused evidence body.
[0065] The emotion inference unit also includes an evidence output module and a result determination module;
[0066] The evidence body output module merges the evidence body into a new evidence body, which is then merged with other modal evidence bodies. This process is repeated until all modal evidence bodies are merged to obtain a comprehensive evidence body.
[0067] The result determination module selects the emotion category with the highest probability value from the comprehensive evidence and uses it as the final emotion prediction result for the current time window;
[0068] A sentiment inference method based on multimodal sentiment monitoring includes the following steps:
[0069] S1. Read the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal. Use the timestamp to synchronize each data within the window.
[0070] S2. Process the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and construct them into feature tuples;
[0071] S3. Input the feature tuple into the value encoder, extract the feature vector, analyze the preliminary emotion probability distribution of different modalities based on the feature vector, calculate the consistency coefficient between different modalities, and adjust the emotion probability value corresponding to each modality based on the consistency coefficient.
[0072] S4. Construct a comprehensive evidence body based on the modal-adjusted sentiment probability values, and use the comprehensive evidence body to determine the sentiment prediction result for the current time window.
[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0074] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An emotion inference system based on multi-modal emotion monitoring, characterized in that, include: The modal data acquisition unit reads the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal, and uses timestamps to synchronize the data within the window. The modal data processing unit processes the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and constructs them into feature tuples. The modal data analysis unit inputs feature tuples into the value encoder, extracts feature vectors, and analyzes the preliminary emotion probability distribution of different modalities based on the feature vectors. A probability generation unit calculates the consistency coefficient between different modalities and then adjusts the emotion probability value corresponding to each modality based on the consistency coefficient. The emotion inference unit constructs a comprehensive evidence body based on the emotion probability values adjusted for each modality, and uses the comprehensive evidence body to determine the emotion inference result for the current time window.
2. The emotion inference system based on multi-modal emotion monitoring according to claim 1, characterized in that: The modal data processing unit includes an image processing module and an audio processing module; The image processing module obtains the current window. Multiple images from the original facial video frame sequence captured internally. ,in Indicates the first Zhang Image This represents the total number of images, analyzed using a deep neural network, and the output window... Facial image sequence after internal alignment ,in , Indicates the aligned first Zhang Image Indicates the serial number; The audio processing module processes the raw audio time-domain signal. The noise-reduced audio signal is obtained. ,right Frame segmentation is performed using preset frame lengths and frame shifts. Divided into multiple overlapping short-time frames , for the A short time frame Processing is performed to obtain a windowed frame. ,in , This represents the Hamming window function. Indicates the total number of short-time frames. Indicates the first Each short time frame is traversed sequentially. Then, output all windowed frames. ,Will Stored to audio feature set middle.
3. The emotion inference system based on multi-modal emotion monitoring according to claim 2, characterized in that: The modal data processing unit also includes an electrocardiogram (ECG) processing module, an electroencephalogram (EEG) processing module, and a feature construction module; The ECG processing module uses a bandpass filter on the current window. Baseline drift and high-frequency noise were removed from the raw ECG signals acquired internally. The Pan-Tompkins algorithm was used to locate each heartbeat cycle in the filtered signal, and the window was analyzed based on the detected intervals of consecutive R peaks. Heart rate variability feature set ; The EEG processing module uses bandpass filtering to filter the window. The raw EEG signals acquired internally are processed for each channel. The resulting signals are segmented, and a Fast Fourier Transform is performed on each segment. The average power of the standard EEG frequency band is calculated by integrating on the PSD. The power values of each channel and frequency band are combined and standardized to form a window. EEG feature vector ; The feature construction module statistically analyzes the first Synchronization Time Window Facial image sequence within Audio feature set Heart rate variability feature set and EEG feature vectors They are then constructed into feature tuples.
4. The emotion inference system based on multi-modal emotion monitoring according to claim 1, characterized in that: The modal data analysis unit includes a sequence analysis module, a speech analysis module, and a physiological analysis module; The sequence analysis module inputs the aligned facial image sequence from the feature tuple into the visual encoder. After extracting the corresponding high-level visual feature vectors through forward propagation, it maps them to the emotion category space through a fully connected classification head to generate a preliminary visual modal emotion probability distribution. The speech analysis module inputs the aligned audio feature set from the feature tuple into the speech encoder, extracts the high-level speech feature vector, and uses it through the classification head to obtain the preliminary speech modal emotion probability distribution. The physiological analysis module concatenates the aligned heart rate variability feature set and EEG feature vector in the feature tuple to form a combined physiological feature vector, which is then input into the physiological encoder to obtain the high-level physiological feature vector. After that, the preliminary physiological modality emotion probability distribution is obtained through the classification head.
5. The emotion inference system based on multi-modal emotion monitoring according to claim 1, characterized in that: The probability generation unit includes a matrix construction module, a probability adjustment module, a deviation coefficient calculation module, and a coefficient output module; The matrix construction module receives high-level feature vectors corresponding to different modalities. Then, select Calculate its relationship with Consistency coefficient between After calculating the consistency coefficient between each high-level feature vector and other feature vectors in turn, the corresponding matrix is constructed using all the consistency coefficients. The probability adjustment module calculates the mean consistency value of each mode based on the consistency matrix, and selects the mode. Read the consistency mean of this mode. And the sentiment probability given by the initial classification head ,according to and Calculate the adjusted emotional probability value ,in , Indicates the scaling factor. Indicates the adjusted number The probability value of each emotion. Indicates the first The probability value of each emotion. Indicates the total number of emotion categories; The deviation coefficient calculation module receives the adjusted emotional probability value. Using consistent mean Analyze the deviation coefficient within the current time window ,in , Indicates the penalty factor; The coefficient output module repeats the operation until the adjusted emotional probability value and deviation coefficient corresponding to each modality are calculated.
6. The emotion inference system based on multimodal emotion monitoring according to claim 1, characterized in that: The emotion inference unit includes an evidence fusion module; The evidence fusion module statistically analyzes the modality-adjusted emotional probability value and deviation coefficient for each mode, combines them into a set of evidence for the corresponding mode, and selects the mode. and modality evidence Then, traverse All propositions within, for any two propositions and ,like Then and The product value is added to the corresponding composite proposition probability value, if Then and The product value according to and The original probability values are assigned, and the assigned values are multiplied by the original probability values. The resulting value is used as the new value of the current proposition. After calculating all the proposition combinations, each proposition is summed and normalized to obtain the fused evidence body.
7. The emotion inference system based on multimodal emotion monitoring according to claim 6, characterized in that: The emotion inference unit also includes an evidence output module and a result determination module; The evidence body output module merges the evidence body into a new evidence body, and then merges it with other modal evidence bodies. This process is repeated until all modal evidence bodies are merged to obtain a comprehensive evidence body. The result determination module selects the emotion category with the highest probability value from the comprehensive evidence and uses it as the final emotion prediction result for the current time window.
8. A method for emotion inference based on multimodal emotion monitoring, characterized in that, The emotion inference method is applicable to the emotion inference system based on multimodal emotion monitoring as described in any one of claims 1-7, and includes the following steps: S1. Read the raw multimodal data stream within the current synchronization time window, denoted as a quadruple. The raw multimodal data includes the acquired raw facial video frame sequence, audio time-domain signal, electrocardiogram signal, and electroencephalogram signal. Use the timestamp to synchronize each data within the window. S2. Process the original multimodal data stream to obtain facial image sequences, audio feature sets, heart rate variability feature sets, and EEG feature vectors, and construct them into feature tuples; S3. Input the feature tuple into the value encoder, extract the feature vector, analyze the preliminary emotion probability distribution of different modalities based on the feature vector, calculate the consistency coefficient between different modalities, and adjust the emotion probability value corresponding to each modality based on the consistency coefficient. S4. Construct a comprehensive evidence body based on the modal-adjusted sentiment probability values, and use the comprehensive evidence body to determine the sentiment prediction result for the current time window.
Citation Information
Patent Citations
Emotion monitoring model generation method, emotion monitoring method and emotion intervention method
CN120216948A