Video-based non-contact multi-mode emotion recognition method and system

By using non-contact video acquisition and large language model analysis, the problems of contact sensors and environmental noise in multimodal emotion recognition have been solved, achieving efficient and accurate emotion state recognition.

CN121506499APending Publication Date: 2026-02-10CSIC WUHAN LINCOM ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511855224.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing multimodal emotion recognition methods rely on contact sensors, which are costly, difficult to transfer, and greatly affected by environmental factors. Furthermore, the high cost of multimodal data collection and processing, along with serious inconsistencies in labeling, result in low accuracy in emotion recognition.

Method used

By acquiring video image features in a non-contact manner, parameters such as nystagmus, heart rate, respiratory rate, blood oxygen saturation, and facial expression are extracted, converted into structured text, and input into a large language model for comprehensive analysis to identify emotional states.

Benefits of technology

It achieves efficient and accurate automatic recognition of emotional states, avoiding the complexity of contact devices and environmental noise interference, and improving the interpretability and reliability of the recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506499A_ABST
    Figure CN121506499A_ABST
Patent Text Reader

Abstract

The invention discloses a video-based non-contact multi-modal emotion recognition method and system. The method comprises the steps of facial video acquisition, multi-modal physiological and behavior feature extraction, multi-modal feature textual description and emotion reasoning and output based on a large language model. According to the method, original data collection is carried out in a non-contact video collection mode, and heart rate, respiration rate and blood oxygen saturation parameters of a target are obtained through a remote light volume change tracing method; the method comprises the following steps: obtaining pupil tremor features by segmenting an eye region from a face video and performing pupil and iris segmentation, extracting facial expression features, and obtaining various emotion expression features with practical significance; a large language model is used as an emotion judgment tool, multi-modal feature textualization is input into the large language model, emotional state recognition can be carried out from external and internal expressions at the same time, the problem of single-modal data interference is avoided, and universal emotion recognition is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an emotion recognition method, specifically a video-based non-contact multimodal emotion recognition method and system, belonging to the field of artificial intelligence and affective computing technology. Background Technology

[0002] Emotion recognition has significant application value in fields such as mental health analysis, emotion monitoring, and psychiatric treatment. Generally, common strategies for emotion analysis involve recognizing facial expression features or using contact-based physiological signal detection sensors to identify the target's facial expressions and physiological conditions, in order to analyze the target's current emotional state through external and internal manifestations.

[0003] In reality, facial expressions can be misleading due to the target's disguise. By incorporating the detection of physiological parameters such as heart rate and respiratory rate, the target's intrinsic physiological behavior can reveal their implicit emotional state. However, current physiological signal acquisition mainly relies on contact sensors, requiring the target to wear corresponding devices, which is costly and difficult to transfer. Furthermore, current multimodal sentiment analysis methods primarily rely on learning or basic deep learning models, heavily depending on the initial training data. Different test individuals may possess different underlying physiological states; if the model can only accurately identify individuals with parameters similar to those in the training data, it will struggle to quickly adapt to new, non-standard individuals.

[0004] Furthermore, environmental factors in real-world application scenarios also pose numerous obstacles to feature extraction. For example, dynamic changes in lighting conditions can cause the target and background colors to become confused, leading to false detections and incorrect tracking; cluttered background environments can interfere with the accuracy of feature recognition; rapid movement, irregular motion, and sudden changes in motion parameters of targets in videos significantly increase the difficulty of feature capture. In addition, a large amount of irrelevant information mixed in with video data constitutes noise interference, directly affecting the quality of feature extraction and ultimately leading to a decrease in the accuracy of emotion recognition, making it difficult to meet the application requirements of high precision and high reliability.

[0005] Besides environmental and technological bottlenecks, the collection and processing of multimodal data itself also face significant challenges. On the one hand, the collection of multimodal data requires a large amount of manpower, material resources, and time; on the other hand, accurate sentiment labeling of the data is extremely difficult, and the labeling results are easily affected by subjective factors, resulting in inconsistencies. Such labeling biases will directly interfere with the training effect of the model and reduce the reliability of sentiment recognition.

[0006] Current multimodal emotion recognition methods largely rely on video and audio information. However, audio acquisition is often severely interfered with in noisy environments, limiting their practical application value. Regarding physiological signal acquisition, existing methods primarily rely on contact sensors, such as EEG acquisition devices, electrode caps, and ECG monitoring equipment. These devices need to be worn, causing discomfort and significant inconvenience in practical applications. Multimodal emotion recognition typically integrates physiological signals including ECG, EEG, respiration, and EMG, which can be acquired through wristband, chest-worn, or head-mounted devices. However, wearing multiple devices simultaneously further exacerbates the complexity and inconvenience of use. Overall, traditional emotion recognition methods still fall short in comprehensively utilizing physiological signals, are easily affected by subjective factors or environmental noise, and their robustness in complex scenarios needs improvement. Summary of the Invention

[0007] The purpose of this invention is to provide a video-based non-contact multimodal emotion recognition method and system to solve at least one of the aforementioned technical problems. This method extracts video image features, detects the target's current nystagmus, heart rate, respiratory rate, blood oxygen saturation, and facial expression parameters, and then converts these features into structured text descriptions. These descriptions are then concatenated into a complete multimodal feature text. Finally, the multimodal input text is input into a large language model, and natural language reasoning is used to comprehensively analyze the information from each modality to identify the target's current emotional state. This invention relies solely on video input captured by a color camera to identify the target's emotional state, effectively and accurately achieving automatic recognition of the target's emotional state.

[0008] This invention achieves the above objective through the following technical solution: a video-based non-contact multimodal emotion recognition method, which includes the following steps:

[0009] S1. Facial video capture: Acquire temporal video data of the facial and eye areas of the target object through a video capture device;

[0010] S2. Multimodal physiological and behavioral feature extraction: Process time-series video data and simultaneously extract heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters based on eye morphology, and facial expression feature parameters based on physiological signals.

[0011] S3. Multimodal Feature Textual Description: The extracted heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters, and facial expression feature parameters are transformed into structured natural language descriptions to generate comprehensive text containing multi-dimensional features.

[0012] S4. Emotional reasoning and output based on a large language model: Input comprehensive text containing multi-dimensional features into a pre-trained large language model, and output the judgment result of the current emotional state of the target object through its natural language understanding and reasoning ability.

[0013] As a further technical solution of the present invention, S1 includes:

[0014] S11. Use an RGB color camera to capture the facial video stream of the target object and extract video frames at a preset frame rate.

[0015] S12. For each video frame, a facial key point detection algorithm is used to locate and crop out the facial region sub-image and the eye region sub-image.

[0016] As a further technical solution of the present invention, S2 includes:

[0017] S21. Based on the facial region sub-image, the blood flow signal on the skin surface is extracted by remote photoplethysmography, and the blood flow signal is analyzed in the frequency domain to calculate the heart rate, respiratory rate and blood oxygen saturation values.

[0018] S22. Based on sub-images of the eye region, an image segmentation algorithm is used to identify the pupil region and the iris region, and the ratio of the diameter of the pupil region to the diameter of the iris region is calculated as a pupil-iris ratio parameter representing nystagmus or physiological arousal state.

[0019] S23. Based on facial region sub-images, an expression classification neural network is used to extract feature vectors representing facial emotional expressions and calculate their confidence scores corresponding to preset basic emotion categories.

[0020] As a further technical solution of the present invention, in S21, the extraction of blood flow signals on the skin surface using remote photoplethysmography includes the STMap construction and feature extraction stages;

[0021] STMap is used to transform raw facial video frames into a spatiotemporal graph rich in physiological information.

[0022] In the feature extraction stage, periodic features of remote photoplethysmography (TPM) are extracted using a masked autoencoder. Under label-free conditions, the visual Transformer encoder learns to extract periodic remote TPM signal features from the STMap.

[0023] As a further technical solution of the present invention, in S22, an image segmentation algorithm is used to identify the pupil region and the iris region, and the ratio of the diameter of the pupil region to the diameter of the iris region is calculated, specifically including:

[0024] The eye region sub-image uses ResNet50 as the backbone network to extract multi-level convolutional features of the eye image. Candidate pupil and iris region bounding boxes are generated through the region proposal network. An end-to-end training strategy is adopted to enable the network to learn the visual feature differences between the pupil and iris regions. Finally, the pupil region mask and iris region mask of the left and right eyes are output.

[0025] The masked segmentation results of the pupil region and iris region are used to calculate the ratio of the width of the iris region and the width of the pupil region in the pixel coordinate system by taking the axial difference of the outermost pixels on the left and right sides.

[0026] As a further technical solution of the present invention, in S23, the basic emotion categories include anger, sadness, calmness, happiness, surprise, fear, and disgust.

[0027] As a further technical solution of the present invention, S3 includes:

[0028] S31. Perform time-series statistical analysis on the heart rate, respiratory rate, blood oxygen saturation and pupil-iris ratio parameters extracted within the same time window, and calculate their mean, variance and slope of change trend.

[0029] S32. Based on the preset threshold rules, the above statistical results are converted into qualitative natural language descriptions. At the same time, the average confidence score of the basic emotion category is calculated within the same time window.

[0030] S33. Fill all qualitative natural language descriptions and average data into a predefined natural language template to generate a comprehensive text containing multi-dimensional features.

[0031] As a further technical solution of the present invention, in S33, the natural language template used is: heart rate feature description, respiratory rate feature description, blood oxygen saturation feature description, pupil-iris ratio feature description, and facial expression description.

[0032] As a further technical solution of the present invention, S4 includes:

[0033] S41. Combine the comprehensive text containing multi-dimensional features with the preset instruction prompt template to form the input prompt words of the complete large language model;

[0034] S42. Input the input prompt words into the large language model, obtain the text output generated by it, and parse the quantitative score values ​​representing emotional valence and emotional arousal from the output.

[0035] A video-based non-contact multimodal emotion recognition system, comprising:

[0036] The hardware acquisition unit includes an RGB color camera for acquiring a video stream containing the target's face and eye area at a frame rate of not less than 25fps.

[0037] The processing and control unit includes a non-contact multimodal emotion assessment module and a non-contact unimodal emotion recognition module;

[0038] The non-contact multimodal emotion assessment module includes a multimodal fusion emotion assessment and analysis module and a system human-computer interaction interface software module.

[0039] The multimodal fusion emotion assessment and analysis module includes:

[0040] The physiological parameter detection submodule is used to detect physiological indicators and receives the output from the physiological parameter detection module.

[0041] A physiological emotion detection module that infers emotions based on physiological parameters is used to determine physiological arousal status based on the changing trends of heart rate, respiratory rate, and blood oxygen saturation.

[0042] A behavioral emotion detection module that extracts facial expression features and eye movement features, integrating the fusion analysis of facial expression and eye movement temporal features;

[0043] The multimodal fusion emotion detection module integrates the detection results of various modalities. It is used to perform feature-level fusion of physiological parameters, facial expression confidence, and pupil-iris ratio parameters, and to perform natural language inference through a large language model to output emotion valence and arousal scores.

[0044] The system's human-computer interaction interface software module is used to display the video stream, feature extraction process, intermediate results of emotion analysis of each modality, and final fused emotion output in real time, and provides user configuration and data export functions;

[0045] The non-contact single-modal emotion recognition module includes a physiological parameter detection module, a facial expression-based emotion recognition module, and a behavior-based emotion recognition module.

[0046] The physiological parameter detection module is used to extract physiological signal features from facial video streams. Specifically, it includes: constructing a spatiotemporal map based on remote photoplethysmography, extracting periodic physiological features through a mask autoencoder, and then calculating heart rate, respiratory rate and blood oxygen saturation parameters.

[0047] The facial expression emotion recognition module is used to extract expression features from facial region images. Specifically, it uses an expression classification neural network to extract facial feature vectors and outputs confidence scores corresponding to seven basic emotions: anger, sadness, calmness, happiness, surprise, fear, and disgust.

[0048] The behavior-based emotion recognition module is used to extract eye movement behavior features from eye region images. Specifically, it includes: identifying the pupil and iris regions through image segmentation algorithms, calculating the pupil-iris width ratio, and generating temporal features representing eye tremors based on continuous frame sequences.

[0049] The beneficial effects of this invention are:

[0050] 1) This invention collects raw data through non-contact video acquisition and obtains the target's heart rate, respiratory rate, and blood oxygen saturation parameters through remote photoplethysmography, avoiding the problems of complex use and difficult deployment and migration of contact devices;

[0051] 2) This invention obtains pupil tremor features by segmenting the eye region from facial videos and performing pupil and iris segmentation, and extracts facial expression features, thereby obtaining a variety of meaningful emotional expression features, which improves the interpretability and reliability of the emotion recognition method.

[0052] 3) This invention uses a large language model as an emotion judgment tool. By inputting multimodal features into the large language model in text form, it can simultaneously identify emotional states from both external and internal expressions, avoiding the problem of interference from single-modal data and achieving universal emotion recognition. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the architecture of the evaluation system of the present invention;

[0054] Figure 2 This is a schematic flowchart of the method of the present invention;

[0055] Figure 3 This is a schematic diagram illustrating the application of the evaluation system of this invention;

[0056] Figure 4 This is a schematic diagram of the single-modal emotion recognition process based on RPPG of the present invention;

[0057] Figure 5 This is a schematic diagram of the emotion recognition algorithm based on facial behavior according to the present invention.

[0058] Figure 6 This is a flowchart of the multimodal emotion recognition algorithm of this invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Example 1, as Figure 2 , Figure 4 , Figure 5 and Figure 6 As shown, this embodiment provides a video-based non-contact multimodal emotion recognition method, which specifically includes:

[0061] First: Facial video capture, which uses a video capture device (camera) to obtain time-series video data of the target object's facial and eye areas.

[0062] This step specifically includes:

[0063] 1) Use an RGB color camera to capture the facial video stream of the target object and extract video frames at a preset frame rate; specifically, use an RGB color camera (specifically, the Aoni S500 model) to capture video from the front of the target object. The preferred capture resolution of the camera is 2560×1440 pixels, and the preferred frame rate is 30fps (frames / second). Extract a continuous sequence of image frames from the captured video stream at this frame rate for subsequent processing.

[0064] 2) For each video frame, a facial key point detection algorithm is used to locate and crop out sub-images of the facial region and the eye region, which are then used for subsequent processing;

[0065] 21) For each frame in the image frame sequence, a facial key point detection algorithm (e.g., using the 68-point detection model provided in the dlib open-source library) is used to locate the key point coordinates of the facial contour, left eye contour, and right eye contour;

[0066] 22) Based on the obtained key point coordinates, calculate the minimum bounding rectangle that can completely contain the face, left eye and right eye regions respectively. Based on these rectangles, crop out the corresponding face region sub-image (P1), left eye region sub-image (PE1) and right eye region sub-image (PE2) from the original image frame.

[0067] Second: Multimodal physiological and behavioral feature extraction. The time-series video data is processed to simultaneously extract heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters based on eye morphology, and facial expression feature parameters based on physiological signals.

[0068] This step specifically includes:

[0069] 1) Based on the facial region sub-image (P1), the blood flow signal on the skin surface is extracted by remote photoplethysmography, and the blood flow signal is analyzed in the frequency domain to calculate the heart rate, respiratory rate and blood oxygen saturation value based on the random forest model regression.

[0070] The method of extracting blood flow signals from the skin surface using remote photoplethysmography (TPM) is called remote photoplethysmography (TPM) pulse wave signal feature extraction. The process includes the STMap construction and feature extraction stages. STMap construction is used to transform the original facial video frames into a structured spatio-temporal map (STMap) rich in physiological information to facilitate the subsequent extraction of remote photoplethysmography pulse wave signals by the network. The specific process is as follows: First, the facial key points are detected by a detection algorithm (such as SeetaFace), and then the face is divided into multiple regions such as the forehead and cheeks. The RGB mean of each region is extracted in each frame to form a feature vector. Then, the RGB values ​​of each region in each frame are concatenated to form a three-dimensional tensor, and the STMap is adjusted to a size of 224×224×3 as the network input. Finally, the signal is processed by an orthogonal plane projection algorithm and a chromatic pulse signal extraction algorithm. The processed signal is concatenated with the original RGB to form a 224×224×6 PC-STMap to improve robustness under complex lighting and motion scenes. In the feature extraction stage, periodic features of remote photoplethysmography (TPM) are extracted using a masked autoencoder. Under label-free conditions, the visual Transformer encoder learns to extract periodic remote TPM signal features from the STMap.

[0071] Heart rate calculation sequentially performs signal preprocessing, frequency domain transformation, and peak detection: First, detrending operations are used to remove the DC component and extremely low-frequency drift of the remote photoplethysmography (PPG) signal to avoid spectral leakage. Then, Butterworth zero-phase bandpass filtering (passband 0.7–3Hz, corresponding to 42–180 bpm) is used to filter valid signals. Subsequently, the Welch method is used to transform the preprocessed signal from the time domain to the frequency domain and generate a power spectrum. Finally, the characteristic frequency corresponding to the maximum amplitude in the power spectrum is extracted and multiplied by 60 to obtain the heart rate parameter (unit: bpm).

[0072] The respiratory rate calculation is divided into two steps: respiratory signal feature extraction and respiratory rate conversion. First, the amplitude modulation component in the remote photoplethysmography pulse wave signal is extracted as the respiratory wave signal through continuous wavelet transform. Then, the autocorrelation function is used to analyze the autocorrelation characteristics of the respiratory wave signal, extract the characteristic frequency corresponding to the respiratory rhythm, and convert it into a respiratory rate parameter (unit: breaths / minute).

[0073] Blood oxygen saturation calculation mainly uses the random forest regression method. First, multiple classification regression trees are constructed through bootstrapping. When splitting nodes, a subset of features is randomly selected for optimal splitting. Finally, the output of all trees is averaged to obtain the blood oxygen saturation parameter (percentage value).

[0074] 2) Based on the sub-images of the eye region (PE1 and PE2), the image segmentation algorithm is used to accurately identify the pupil region and the iris region, and the ratio of the diameter of the pupil region to the diameter of the iris region is calculated as the pupil-iris ratio parameter representing the eye tremor or physiological arousal state.

[0075] For the eye region image of the i-th frame and The ResNet50 was used as the backbone network to extract multi-level convolutional features from eye images; candidate pupil and iris region bounding boxes were generated through a Region Proposal Network (RPN); an end-to-end training strategy was adopted to enable the network to learn the visual feature differences between the pupil and iris regions; and finally, the pupil region mask and iris region mask of the left and right eyes were output.

[0076] For the masked segmentation results of the pupil and iris regions in the i-th frame, the x-axis difference between the left and right edge pixels is taken as the region width, and the pupil region width is denoted as... The width of the iris region is denoted as Therefore, the ratio of the width of the iris region to the width of the pupil region in the pixel coordinate system. It can be calculated using the following formula:

[0077] ;

[0078] For N consecutive video frames, generate a temporal feature sequence { , , ..., }, as a quantitative parameter index for nystagmus.

[0079] 3) Based on the facial region sub-image (P1), an expression classification neural network is used to extract feature vectors representing facial emotional expressions and calculate their confidence scores corresponding to preset basic emotion categories, including anger, sadness, calmness, happiness, surprise, fear, and disgust.

[0080] The facial feature vector extracted from the i-th frame is normalized using the Softmax function to obtain the feature confidence probabilities of seven facial expressions (anger, sadness, calmness, happiness, surprise, fear, and disgust). As the facial expression feature parameter for the current frame, for a video with N frames, each frame's... Composition parameter sequence , as a parameter for facial expression features.

[0081] Third: Multimodal feature textual description, which transforms the extracted heart rate value, respiratory rate value, blood oxygen saturation value, pupil-iris ratio parameter and facial expression feature parameter into a structured natural language description, generating a comprehensive text containing multi-dimensional features.

[0082] This step specifically includes:

[0083] 1) Perform time-series statistical analysis on the heart rate, respiratory rate, blood oxygen saturation, and pupil-iris ratio parameters extracted within the same time window, and calculate their mean, variance, and slope of change trend.

[0084] Calculate the average values ​​of heart rate, respiratory rate, blood oxygen saturation, and pupil-iris ratio over a time window (e.g., 3 seconds). ),variance( ) and trends ( ).

[0085] Taking heart rate as an example, let the parameter sequence of heart rate within an N-frame time window be... The formula for calculating the average value of the parameter is:

[0086] ;

[0087] The formula for calculating the variance of the parameters is:

[0088] ;

[0089] The trend of parameter change is measured by the Pearson correlation coefficient between the parameter and the time series, and the calculation formula is:

[0090] ;

[0091] Based on the above calculation method, the distribution characteristics of parameters under the three modes can be obtained, which can be used for the next step of parameter feature evaluation and text description generation.

[0092] 2) Based on the preset threshold rules, the above statistical results are converted into qualitative natural language descriptions; at the same time, the average confidence score of the basic emotion category is calculated within the same time window.

[0093] Assume the resting heart rate is respiratory rate The ratio of pupil to iris diameter is In one example, the judgment thresholds and text expressions for the above three parameters are set as shown in the table below:

[0094]

[0095] Simultaneously, the average confidence score of the seven basic facial expressions (anger, sadness, calmness, happiness, surprise, fear, and disgust) for each frame within the video segment is calculated to obtain a comprehensive confidence vector. Facial expressions are then used to generate facial expression description text based on their corresponding confidence scores. For expressions with an average confidence score of 0, the description is omitted. The description text template is as follows:

[0096] "Facial expression showed happiness ( ),calm( ),surprise( ),angry( ),sad( ),fear( ),disgust( )".

[0097] 3) Fill all the qualitative natural language descriptions and average data into the predefined natural language template to generate comprehensive feature text.

[0098] A predefined natural language template is used to combine the textual descriptions of the above multimodal features into a coherent, comprehensive feature text. The structure of the natural language template is as follows:

[0099] "[Heart rate characteristics description], [Respiratory rate characteristics description], [Blood oxygen saturation characteristics description], [Pupil-iris ratio characteristics description]. Facial expression is described as: [Facial expression description]."

[0100] Based on the specific parameter values ​​obtained in 1) and the descriptive text obtained in 2), fill the template to generate the final multimodal feature text.

[0101] The generated text example is as follows:

[0102] "The current test target has a high heart rate, averaging 90 bpm, with significant fluctuations and a gradually increasing trend; a high respiratory rate, with significant fluctuations and a gradually increasing trend; low blood oxygen saturation; and significant pupillary nystagmus. Facial expressions are: happiness (82.22% confidence level), calmness (10.23% confidence level), and surprise (7.55% confidence level)."

[0103] Fourth: Emotion reasoning and output based on large language models. Comprehensive text containing multi-dimensional features is input into a pre-trained large language model. Through its natural language understanding and reasoning capabilities, the model outputs a judgment result on the current emotional state of the target object.

[0104] This step specifically includes:

[0105] 1) The comprehensive feature text is combined with the preset instruction prompt template to form the complete input prompt words for the large language model. The instruction prompt template is used to guide the large language model to perform the emotion quantification analysis task. Its core instruction can be designed as: "Please predict the subject's emotional state based on the above physiological and facial features. Please return only two values: the first represents the emotional valence (1-9 points), and the second represents the emotional arousal (1-9 points)".

[0106] 2) Input prompts from the complete large language model are fed into the pre-trained large language model. Based on its powerful natural language understanding and logical reasoning capabilities, this model analyzes the multimodal text and generates a text response containing sentiment quantification metrics. Finally, the specific scores for sentiment valence and arousal are extracted from the large language model's text output. It should be noted that an open-source model with fine-tuned instructions, such as the Qwen2.5-7B-Instruct model, can be used to achieve the desired sentiment inference effect.

[0107] This embodiment can be summarized as follows: 1) Video capture of the target's face is performed, and frame sequences of the target's face and eye area are obtained through a camera; 2) The acquired temporal video data is processed, and heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters based on physiological signals, and facial expression feature parameters are extracted simultaneously; 3) The extracted heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters, and facial expression feature parameters are converted into structured natural language descriptions to generate a comprehensive text containing multi-dimensional features; 4) The obtained comprehensive text is input into a pre-trained large language model, and through its natural language understanding and reasoning capabilities, the model outputs a judgment result on the target's current emotional state.

[0108] Example 2, as Figure 1 and Figure 3 As shown, this embodiment provides a video-based non-contact multimodal emotion recognition system. This system is used to implement the non-contact multimodal emotion recognition method in Embodiment 1. The system specifically includes:

[0109] The hardware acquisition unit includes an RGB color camera for acquiring a video stream containing the target's face and eye area at a frame rate of not less than 25fps.

[0110] The processing and control unit includes a non-contact multimodal emotion assessment module and a non-contact unimodal emotion recognition module;

[0111] The non-contact multimodal emotion assessment module includes a multimodal fusion emotion assessment and analysis module and a system human-computer interaction interface software module.

[0112] The multimodal fusion emotion assessment and analysis module includes:

[0113] The physiological parameter detection submodule is used to detect physiological indicators and receives the output from the physiological parameter detection module.

[0114] A physiological emotion detection module that infers emotions based on physiological parameters is used to determine physiological arousal status based on the changing trends of heart rate, respiratory rate, and blood oxygen saturation.

[0115] A behavioral emotion detection module that extracts facial expression features and eye movement features, integrating the fusion analysis of facial expression and eye movement temporal features;

[0116] The multimodal fusion emotion detection module integrates the detection results of various modalities. It is used to perform feature-level fusion of physiological parameters, facial expression confidence, and pupil-iris ratio parameters, and to perform natural language inference through a large language model to output emotion valence and arousal scores.

[0117] The system's human-computer interaction interface software module is used to display the video stream, feature extraction process, intermediate results of emotion analysis of each modality, and final fused emotion output in real time, and provides user configuration and data export functions;

[0118] The non-contact single-modal emotion recognition module includes a physiological parameter detection module, a facial expression-based emotion recognition module, and a behavior-based emotion recognition module.

[0119] The physiological parameter detection module is used to extract physiological signal features from facial video streams. Specifically, it includes: constructing a spatiotemporal map based on remote photoplethysmography, extracting periodic physiological features through a mask autoencoder, and then calculating heart rate, respiratory rate and blood oxygen saturation parameters.

[0120] The facial expression emotion recognition module is used to extract expression features from facial region images. Specifically, it uses an expression classification neural network to extract facial feature vectors and outputs confidence scores corresponding to seven basic emotions: anger, sadness, calmness, happiness, surprise, fear, and disgust.

[0121] The behavior-based emotion recognition module is used to extract eye movement behavior features from eye region images. Specifically, it includes: identifying the pupil and iris regions through image segmentation algorithms, calculating the pupil-iris width ratio, and generating temporal features representing eye tremors based on continuous frame sequences.

[0122] Working principle and process: First, behavior-based emotion recognition technology captures and analyzes behavioral features such as eye movements and facial expressions in videos of subjects, using computer vision technology to identify subtle changes in an individual's emotional state. Second, video-based physiological parameter detection technology and physiological performance-based emotion recognition technology mainly extract key physiological parameters, including heart rate, respiratory rate, and blood oxygen saturation, through non-contact physiological signal detection (such as IPPG signals) in videos. The system analyzes the fluctuations of these physiological indicators to identify emotional states from a physiological perspective. Physiological signals can reveal changes in cardiovascular activity in individuals under emotional states such as tension and relaxation, thus providing reliable emotional feature output. Finally, multimodal fusion emotion recognition technology combines behavioral and physiological performance features, making full use of the understanding and generalization capabilities of large language models to integrate and process input information from different modalities.

[0123] By fusing and analyzing multimodal data, the system can provide a more comprehensive and accurate assessment of a subject's emotional state. Compared to single-modal analysis, multimodal fusion not only captures changes in emotional state across different dimensions but also improves the accuracy and robustness of recognition. This technology maintains high recognition performance in complex emotional scenarios, effectively addressing the complexity and diversity of various emotional expressions in reality.

[0124] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0125] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A video-based non-contact multimodal emotion recognition method, characterized in that, The non-contact multimodal emotion recognition method includes the following steps: S1. Facial video capture: Acquire temporal video data of the facial and eye areas of the target object through a video capture device; S2. Multimodal physiological and behavioral feature extraction: The time-series video data is processed to simultaneously extract heart rate, respiratory rate, blood oxygen saturation, pupil-iris ratio parameters based on eye morphology, and facial expression feature parameters based on physiological signals. S3. Multimodal feature textual description: The extracted heart rate value, respiratory rate value, blood oxygen saturation value, pupil-iris ratio parameter and facial expression feature parameter are transformed into structured natural language description, generating comprehensive text containing multi-dimensional features; S4. Emotional reasoning and output based on a large language model: Input comprehensive text containing multi-dimensional features into a pre-trained large language model, and output the judgment result of the current emotional state of the target object through its natural language understanding and reasoning ability.

2. The non-contact multimodal emotion recognition method according to claim 1, characterized in that, S1 includes: S11. Use an RGB color camera to capture the facial video stream of the target object and extract video frames at a preset frame rate. S12. For each video frame, a facial key point detection algorithm is used to locate and crop out the facial region sub-image and the eye region sub-image.

3. The non-contact multimodal emotion recognition method according to claim 2, characterized in that, S2 includes: S21. Based on the facial region sub-image, the blood flow signal on the skin surface is extracted using remote photoplethysmography, and the blood flow signal is analyzed in the frequency domain to calculate the heart rate, respiratory rate and blood oxygen saturation values. S22. Based on the sub-image of the eye region, an image segmentation algorithm is used to identify the pupil region and the iris region, and the ratio of the diameter of the pupil region to the diameter of the iris region is calculated as a pupil-iris ratio parameter representing the nystagmus or physiological arousal state. S23. Based on facial region sub-images, an expression classification neural network is used to extract feature vectors representing facial emotional expressions and calculate their confidence scores corresponding to preset basic emotion categories.

4. The non-contact multimodal emotion recognition method according to claim 3, characterized in that: In step S21, the extraction of blood flow signals from the skin surface using remote photoplethysmography includes the STMap construction and feature extraction stages. The STMap construction is used to transform raw facial video frames into a structured spatiotemporal graph rich in physiological information. The feature extraction stage extracts the periodic features of remote photoplethysmography (TPM) waves through a masked autoencoder, enabling the visual Transformer encoder to learn to extract periodic remote TPM signal features from the STMap under label-free conditions.

5. The non-contact multimodal emotion recognition method according to claim 1, characterized in that: In step S22, an image segmentation algorithm is used to identify the pupil region and the iris region, and the ratio of the pupil region diameter to the iris region diameter is calculated. Specifically, this includes: The eye region sub-image uses ResNet50 as the backbone network to extract multi-level convolutional features of the eye image. Candidate pupil and iris region bounding boxes are generated through the region proposal network. An end-to-end training strategy is adopted to enable the network to learn the visual feature differences between the pupil and iris regions. Finally, the pupil region mask and iris region mask of the left and right eyes are output. The masked segmentation results of the pupil region and iris region are used to calculate the ratio of the width of the iris region and the width of the pupil region in the pixel coordinate system by taking the axial difference of the outermost pixels on the left and right sides.

6. The non-contact multimodal emotion recognition method according to claim 3, characterized in that: In S23, the basic emotion categories include anger, sadness, calmness, happiness, surprise, fear, and disgust.

7. The non-contact multimodal emotion recognition method according to claim 1, characterized in that, S3 includes: S31. Perform time-series statistical analysis on the heart rate, respiratory rate, blood oxygen saturation and pupil-iris ratio parameters extracted within the same time window, and calculate their mean, variance and slope of change trend. S32. Based on the preset threshold rules, the above statistical results are converted into qualitative natural language descriptions. At the same time, the average value of the confidence scores of the basic emotion categories within the same time window is calculated. S33. Fill all qualitative natural language descriptions and average data into a predefined natural language template to generate a comprehensive text containing multi-dimensional features.

8. The non-contact multimodal emotion recognition method according to claim 7, characterized in that: In S33, the natural language templates used are: heart rate feature description, respiratory rate feature description, blood oxygen saturation feature description, pupil-iris ratio feature description, and facial expression description.

9. The non-contact multimodal emotion recognition method according to claim 1, characterized in that, S4 includes: S41. Combine the comprehensive text containing multi-dimensional features with the preset instruction prompt template to form the input prompt words of the complete large language model; S42. Input the input prompt words into the large language model, obtain the text output generated by it, and parse the quantitative score value representing the emotional valence and emotional arousal from the output.

10. A video-based non-contact multimodal emotion recognition system, used to implement the non-contact multimodal emotion recognition method according to any one of claims 1 to 9; characterized in that, The non-contact multimodal emotion recognition system includes: The hardware acquisition unit includes an RGB color camera for acquiring a video stream containing the target's face and eye area at a frame rate of not less than 25fps. The processing and control unit includes a non-contact multimodal emotion assessment module and a non-contact unimodal emotion recognition module; The non-contact multimodal emotion assessment module includes a multimodal fusion emotion assessment and analysis module and a system human-computer interaction interface software module. The multimodal fusion emotion assessment and analysis module includes: The physiological parameter detection submodule is used to detect physiological indicators and receives the output from the physiological parameter detection module. A physiological emotion detection module that infers emotions based on physiological parameters is used to determine physiological arousal status based on the changing trends of heart rate, respiratory rate, and blood oxygen saturation. A behavioral emotion detection module that extracts facial expression features and eye movement features, integrating the fusion analysis of facial expression and eye movement temporal features; The multimodal fusion emotion detection module integrates the detection results of various modalities. It is used to perform feature-level fusion of physiological parameters, facial expression confidence, and pupil-iris ratio parameters, and to perform natural language inference through a large language model to output emotion valence and arousal scores. The system's human-computer interaction interface software module is used to display the video stream, feature extraction process, intermediate results of each modality of emotion analysis, and final fused emotion output in real time, and provides user configuration and data export functions. The non-contact single-modal emotion recognition module includes a physiological parameter detection module, a facial expression-based emotion recognition module, and a behavior-based emotion recognition module. The physiological parameter detection module is used to extract physiological signal features from the facial video stream, specifically including: constructing a spatiotemporal map based on remote photoplethysmography, extracting periodic physiological features through a mask autoencoder, and then calculating heart rate, respiratory rate and blood oxygen saturation parameters. The facial expression emotion recognition module is used to extract expression features from facial region images, specifically including: using an expression classification neural network to extract facial feature vectors and outputting confidence scores corresponding to seven basic emotions: anger, sadness, calmness, happiness, surprise, fear, and disgust. The emotion recognition module based on behavioral performance is used to extract eye movement behavior features from eye region images. Specifically, it includes: identifying the pupil and iris regions through image segmentation algorithms, calculating the pupil-iris width ratio, and generating temporal features representing eye tremors based on continuous frame sequences.